Image classification model training method and device, image classification method and device and electronic equipment

By dividing the VIT model area and extracting features, and adjusting the model parameters in combination with category and position detection networks, the problem of insufficient robustness of the image classification model in the prior art is solved, and higher image classification accuracy is achieved.

CN119992208APending Publication Date: 2025-05-13INTELLINDUST INFORMATION TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510119198.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, the image classification model based on the VIT model has shortcomings in terms of robustness, making it difficult to obtain high accuracy in complex scenarios.

Method used

By dividing the sample image into multiple region sequences, and using the conversion encoder in the image classification model of the initial structure to extract the region sequences, combining the category detection network and the position detection network, and adjusting the model parameters to achieve the preset convergence conditions, a trained image classification model is obtained.

Benefits of technology

It improves the robustness of the image classification model, allowing it to more accurately classify images in complex scenes, and compensates for the false detection problem when classification based on object outlines only.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992208A_ABST
    Figure CN119992208A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image classification model training method and device, an image classification method and device and electronic equipment, and relates to the technical field of machine vision. The method comprises the following steps: acquiring a sample image, and dividing the sample image into sample image areas according to a preset division mode; obtaining a region sequence formed by the sample image regions; performing feature extraction on the region sequence by using a conversion encoder to obtain sample image features; performing category detection on the sample image features by using a category detection network to obtain a category detection result; performing position detection on the sample image features by using a position detection network to obtain a position detection result; and based on the difference between the category label and the category detection result and the difference between the position label and the position detection result, adjusting model parameters of the image classification model of the initial structure until a preset convergence condition is reached, and obtaining a trained image classification model. And an image classification model with higher robustness can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine vision technology, and in particular to an image classification model training method, an image classification method, a device and an electronic device. Background Art

[0002] In the field of machine vision technology, image classification can be performed based on a VIT (Vision Transformer) model that includes a Transformer (conversion network) encoder and a category detection network. Specifically, the image is input into the VIT model, and the features of the image are extracted through the Transformer encoder; the extracted features are subjected to category detection using the category detection network to obtain the category detection result of the image. The category detection result can indicate: the probability that the object in the image is of each preset category. Subsequently, the category of the object can be determined based on the category detection result, that is, the image classification result of the image is obtained. For example, the preset category corresponding to the maximum value of each probability is determined as the category of the object.

[0003] Currently, there is an urgent need for an image classification model training method to obtain a more robust VIT model. Summary of the invention

[0004] The purpose of the embodiments of the present invention is to provide an image classification model training method, an image classification method, an apparatus and an electronic device to obtain an image classification model with higher robustness. The specific technical solution is as follows:

[0005] In a first aspect of an embodiment of the present invention, a method for training an image classification model is provided, the method comprising: obtaining a sample image, and dividing the sample image according to a preset division method to obtain sample image regions; obtaining a region sequence composed of the divided sample image regions; using a conversion encoder in an image classification model of an initial structure to perform feature extraction on the obtained region sequence to obtain sample image features; wherein the image classification model further comprises a category detection network and a position detection network; using the category detection network to perform category detection on the sample image features to obtain category detection results; wherein the category detection results represent: the probability that the sample object in the sample image is of each preset category; using the position detection network to extract features from the obtained region sequence to obtain sample image features; wherein the image classification model further comprises a category detection network and a position detection network; wherein the category detection results represent: the probability that the sample object in the sample image is of each preset category; wherein the position detection network is used to extract features from the obtained region sequence to obtain sample image features ... The network performs position detection on the sample image features to obtain a position detection result; wherein the position detection result represents: the predicted probability that each sample image area is located at each preset position in the sample image; each preset position is determined by dividing the sample image according to the preset division method; based on the difference between the category label and the category detection result, and the difference between the position label and the position detection result, the model parameters of the image classification model of the initial structure are adjusted until the preset convergence condition is reached to obtain a trained image classification model; wherein the category label represents the preset category to which the sample object belongs; and the position label represents the preset position of each sample image area in the sample image.

[0006] Optionally, the region sequence formed by each sample image region obtained by division may include: arranging the sample image regions in a preset order to obtain a region sequence; or arranging the sample image regions in a preset order and shuffling the arrangement result using a shuffling algorithm; combining the arrangement result and the shuffling result to obtain a region sequence; or arranging the sample image regions in a preset order; and shuffling the arrangement result using a shuffling algorithm to obtain a region sequence.

[0007] Optionally, based on the difference between the category label and the category detection result, and the difference between the position label and the position detection result, the model parameters of the image classification model of the initial structure are adjusted until a preset convergence condition is reached to obtain a trained image classification model, including: using a preset loss function to calculate a position loss value representing the difference between the position label and the position detection result, and a category loss value representing the difference between the category label and the category detection result; calculating the weighted sum of the position loss value and the category loss value according to preset position loss weights and category loss weights to obtain a total loss value; based on the total loss value, the model parameters of the image classification model of the initial structure are adjusted until a preset convergence condition is reached to obtain a trained image classification model.

[0008] Optionally, the image classification model is a visual transformer (VIT) model; the conversion encoder is a conversion network (Transformer) encoder; the category detection network includes a fully connected layer; and the position detection network is a multi-layer perceptron (MLP).

[0009] In a second aspect of an embodiment of the present invention, an image classification method is provided, the method comprising: obtaining an image to be classified, and dividing the image to be classified according to a preset division method to obtain image regions to be classified; obtaining a region sequence composed of the divided image regions to be classified; using a conversion encoder in a trained image classification model to perform feature extraction on the obtained region sequence to obtain image features to be classified; wherein the image classification model is obtained based on the image classification model training method described in any one of the first aspects above; the image classification model also includes a category detection network; using the category detection network to perform category detection on the image features to be classified to obtain category detection results; wherein the category detection results represent: the probability that the object in the image to be classified is of each preset category.

[0010] In a third aspect of an embodiment of the present invention, an image classification model training device is provided, the device comprising:

[0011] The image acquisition module is used to acquire a sample image and divide the sample image according to a preset division method to obtain each sample image region; the sequence acquisition module is used to acquire a region sequence composed of each sample image region obtained by division; the feature extraction module is used to extract features from the acquired region sequence using the conversion encoder in the image classification model of the initial structure to obtain sample image features; wherein the image classification model also includes a category detection network and a position detection network; the category detection module is used to perform category detection on the sample image features using the category detection network to obtain a category detection result; wherein the category detection result indicates: the probability that the sample object in the sample image is of each preset category; the position detection ... wherein the category detection result indicates: the probability that the sample object in the sample image is of each preset category; wherein the position detection module is used to extract features from the acquired region sequence using the conversion encoder in the image classification model of the initial structure to obtain sample image features; wherein the image classification model also includes a category detection network and a position detection network; wherein the image classification model also includes a category detection network and a position detection network; wherein the image classification model also includes a category detection network and a position detection network; wherein the image classification model also includes a category detection network and a position detection network; wherein the image classification model also includes a category detection network and a position detection network; wherein the image classification model also includes a category detection network and a position detection network; wherein the image classification model also includes a category detection network and a position detection network; wherein the image classification model also includes a category detection network and a position detection network; wherein the image classification model also includes a category detection network and a position detection network; wherein The detection network performs position detection on the sample image features to obtain a position detection result; wherein the position detection result represents: the predicted probability that each sample image area is located at each preset position in the sample image; each preset position is determined by dividing the sample image according to the preset division method; a training module is used to adjust the model parameters of the image classification model of the initial structure based on the difference between the category label and the category detection result, and the difference between the position label and the position detection result, until a preset convergence condition is reached to obtain a trained image classification model; wherein the category label represents the preset category to which the sample object belongs; the position label represents the preset position of each sample image area in the sample image.

[0012] Optionally, the sequence acquisition module is specifically used to: arrange the sample image areas in a preset order to obtain a region sequence; or, arrange the sample image areas in a preset order, and shuffle the arrangement results using a shuffling algorithm; combine the arrangement results and the shuffling results to obtain a region sequence; or, arrange the sample image areas in a preset order; and shuffle the arrangement results using a shuffling algorithm to obtain a region sequence.

[0013] Optionally, the training module is specifically used to: use a preset loss function to calculate a position loss value representing the difference between a position label and the position detection result, and a category loss value representing the difference between a category label and the category detection result; calculate a weighted sum of the position loss value and the category loss value according to preset position loss weights and category loss weights to obtain a total loss value; and adjust model parameters of the image classification model of the initial structure based on the total loss value until a preset convergence condition is reached to obtain a trained image classification model.

[0014] Optionally, the image classification model is a visual transformer (VIT) model; the conversion encoder is a conversion network (Transformer) encoder; the category detection network includes a fully connected layer; and the position detection network is a multi-layer perceptron (MLP).

[0015] In a fourth aspect of the embodiments of the present invention, an image classification device is provided, the device comprising:

[0016] A first acquisition module is used to acquire an image to be classified, and divide the image to be classified according to a preset division method to obtain image regions to be classified; a second acquisition module is used to acquire a region sequence composed of the image regions to be classified obtained by division; an extraction module is used to perform feature extraction on the acquired region sequence using a conversion encoder in a trained image classification model to obtain features of the image to be classified; wherein the image classification model is obtained based on any of the image classification model training methods described in the first aspect above; the image classification model also includes a category detection network; a detection module uses the category detection network to perform category detection on the image features to be classified to obtain category detection results; wherein the category detection results represent: the probability that the object in the image to be classified is of each preset category.

[0017] In a fifth aspect of an embodiment of the present invention, an electronic device is provided, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; the memory is used to store computer programs; and the processor is used to implement the image classification model training method described in any one of the first aspects above, or implement the image classification method described in the second aspect above, when executing the program stored in the memory.

[0018] In a sixth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the image classification model training method described in any one of the first aspects above is implemented, or the image classification method described in the second aspect above is implemented.

[0019] An embodiment of the present invention also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any of the image classification model training methods described in the first aspect above, or to execute the image classification method described in the second aspect above.

[0020] The image classification model training method provided by the embodiment of the present invention comprises the following steps: an electronic device obtains a sample image, and divides the sample image according to a preset division method to obtain sample image regions; obtains a region sequence composed of the divided sample image regions; uses a conversion encoder in an image classification model of an initial structure to extract features from the obtained region sequence to obtain sample image features; the image classification model also comprises a category detection network and a position detection network; uses the category detection network to perform category detection on the sample image features to obtain category detection results indicating the probability that the sample object in the sample image is of each preset category; uses the position detection network to perform position detection on the sample image features to obtain position detection results indicating the predicted probability that each sample image region is located at each preset position in the sample image; each preset position is determined by dividing the sample image according to a preset division method; based on the difference between the category label and the category detection result, and the difference between the position label and the position detection result, the model parameters of the image classification model of the initial structure are adjusted until a preset convergence condition is reached to obtain a trained image classification model; the category label indicates the preset category to which the sample object belongs; and the position label indicates the preset position where each sample image region is located in the sample image.

[0021] Based on the above processing, when training the image classification model, after the electronic device obtains the sample image, it will divide the sample image into multiple sample image areas according to the preset division method, and then obtain the area sequence composed of the divided sample image areas. The image classification model includes a conversion encoder, a category detection network and a position detection network. After the sample image features of the area sequence are extracted by the conversion encoder, the sample image features are subjected to category detection by the category detection network to obtain the category detection result, that is, the probability that the sample object in the sample image belongs to each preset category; according to the difference between the category detection result and the category label representing the preset category to which the sample object belongs, the category loss value can be obtained. The sample image features are subjected to position detection by the position detection network to obtain the position detection result, that is, the predicted probability that each sample image area in the area sequence is located at each preset position in the sample image; according to the difference between the position detection result and the position label, the position loss value can be obtained. The position label indicates the preset position where each sample image area is located in the sample image. Subsequently, the image classification model is trained by combining the category loss value and the position loss value, that is, the image classification model is trained based on the positional relationship between the image regions divided according to the preset division method and the preset category to which the object in the image belongs. Then, the trained image classification model can learn the local information of each image region, the global information of the image, and the positional relationship between the image regions. Subsequently, when image classification is performed based on the trained image classification model, the features that can be extracted by the trained image classification model can indicate: the local features of each image region, the global features of the object in the image, and the positional relationship between the image regions. Image classification is performed based on global features, that is, image classification is performed according to the contour of the object in the image; image classification is performed based on the local features of each image region and the positional relationship between the image regions, that is, image classification is performed according to the details of the object. Image classification is performed by combining the contour of the object and the details of the object, that is, combining more abundant information indicating the category to which the object in the image belongs, and obtaining the category detection result together, which can make up for the defect that image classification is performed based only on the contour of the object, resulting in the object being misdetected as belonging to the preset category when the contour of the object is similar to the contour corresponding to a preset category. That is, the trained image classification model obtained by the image classification model training method provided by the present invention has higher robustness.

[0022] Of course, it is not necessary to achieve all of the advantages described above at the same time to implement any product or method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.

[0024] Figure 1 A first flow chart of the image classification model training method provided by an embodiment of the present invention;

[0025] Figure 2 A schematic diagram of training a VIT model provided by an embodiment of the present invention;

[0026] Figure 3 A second flow chart of the image classification model training method provided by an embodiment of the present invention;

[0027] Figure 4 A schematic diagram of training a VIT model based on a training method;

[0028] Figure 5 A flow chart of an image classification method provided by an embodiment of the present invention;

[0029] Figure 6 A structural diagram of an image classification model training device provided by an embodiment of the present invention;

[0030] Figure 7 A structural diagram of an image classification device provided by an embodiment of the present invention;

[0031] Figure 8 A structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field based on the present invention belong to the scope of protection of the present invention.

[0033] In the field of machine vision technology, based on the VIT model including the Transformer encoder and the category detection network, the category detection result of the image to be classified can be obtained. The category of the object is determined based on the category detection result, that is, the image classification result of the image is obtained.

[0034] In order to obtain a more robust image classification model, the present invention provides an image classification model training method, which is applied to an electronic device. The electronic device may be a server. After acquiring a sample image, the electronic device may divide the sample image according to a preset division method to obtain each sample image region; obtain a region sequence composed of each divided sample image region; use the conversion encoder in the image classification model of the initial structure to extract features of the obtained region sequence to obtain sample image features; the image classification model also includes a category detection network and a position detection network. The category detection network is used to perform category detection on the sample image features to obtain a type detection result indicating the probability that the sample object in the sample image is each preset category. The position detection network is used to perform position detection on the sample image features to obtain a position detection result indicating the predicted probability that each sample image region is located at each preset position in the sample image; each preset position is determined by dividing the sample image according to the preset division method. Based on the difference between the category label indicating the preset category to which the sample object belongs and the category detection result, and the difference between the position label indicating the preset position where each sample image region is located in the sample image and the position detection result, the model parameters of the image classification model of the initial structure are adjusted until the preset convergence condition is reached to obtain a trained image classification model. The trained image classification model obtained in this way can learn the local information of each image area, the global information of the image, and the positional relationship between each image area. Subsequently, when image classification is performed based on the trained image classification model, the features that can be extracted by the trained image classification model can indicate: the local features of each image area, the global features of the object in the image, and the positional relationship between each image area. Image classification is performed based on global features, that is, image classification is performed according to the contour of the object in the image; image classification is performed based on the local features of each image area and the positional relationship between each image area, that is, image classification is performed according to the details of the object. Image classification is performed by integrating the contour of the object and the details of the object, that is, more abundant information indicating the category to which the object in the image belongs is integrated to obtain a category detection result, which can make up for the defect that image classification is performed only based on the contour of the object, resulting in the object being misdetected as belonging to the preset category when the contour of the object is similar to the contour corresponding to a preset category. That is, the trained image classification model obtained based on the image classification model training method provided by the present invention has higher robustness.

[0035] See also Figure 1 , Figure 1 A first flow chart of the image classification model training method provided by an embodiment of the present invention. The method may include the following steps:

[0036] S101: Acquire a sample image, and divide the sample image according to a preset division method to obtain various sample image regions.

[0037] S102: Obtain a region sequence consisting of the divided sample image regions.

[0038] S103: Using the conversion encoder in the image classification model of the initial structure to extract features from the acquired region sequence, to obtain sample image features.

[0039] Among them, the image classification model also includes a category detection network and a position detection network.

[0040] S104: Perform category detection on sample image features using a category detection network to obtain a category detection result.

[0041] The category detection result indicates the probability that the sample object in the sample image belongs to each preset category.

[0042] S105: Perform position detection on sample image features using a position detection network to obtain a position detection result.

[0043] The position detection result indicates: the predicted probability that each sample image region is located at each preset position in the sample image; each preset position is determined by dividing the sample image according to a preset division method.

[0044] S106: Based on the difference between the category label and the category detection result, and the difference between the position label and the position detection result, the model parameters of the image classification model of the initial structure are adjusted until a preset convergence condition is reached to obtain a trained image classification model.

[0045] The category label indicates the preset category to which the sample object belongs; and the position label indicates the preset position where each sample image region is located in the sample image.

[0046] Based on the image classification model training method provided by the embodiment of the present invention, when training the image classification model, after acquiring the sample image, the electronic device will divide the sample image into multiple sample image areas according to a preset division method, and then obtain a region sequence composed of each sample image area obtained by the division. The image classification model includes a conversion encoder, a category detection network and a position detection network. After the sample image features of the region sequence are extracted by the conversion encoder, the sample image features are subjected to category detection by the category detection network to obtain a category detection result, that is, the probability that the sample object in the sample image is each preset category is obtained; according to the difference between the category detection result and the category label representing the preset category to which the sample object belongs, the category loss value can be obtained. The sample image features are subjected to position detection by the position detection network to obtain a position detection result, that is, the predicted probability that each sample image area in the region sequence is located at each preset position in the sample image; according to the difference between the position detection result and the position label, the position loss value can be obtained. The position label indicates the preset position where each sample image area is located in the sample image. Subsequently, the image classification model is trained by combining the category loss value and the position loss value, that is, the image classification model is trained based on the positional relationship between the image regions divided according to the preset division method and the preset category to which the object in the image belongs. Then, the trained image classification model can learn the local information of each image region, the global information of the image, and the positional relationship between the image regions. Subsequently, when image classification is performed based on the trained image classification model, the features that can be extracted by the trained image classification model can indicate: the local features of each image region, the global features of the object in the image, and the positional relationship between the image regions. Image classification is performed based on global features, that is, image classification is performed according to the contour of the object in the image; image classification is performed based on the local features of each image region and the positional relationship between the image regions, that is, image classification is performed according to the details of the object. Image classification is performed by combining the contour of the object and the details of the object, that is, combining more abundant information indicating the category to which the object in the image belongs, and obtaining the category detection result together, which can make up for the defect that image classification is performed based only on the contour of the object, resulting in the object being misdetected as belonging to the preset category when the contour of the object is similar to the contour corresponding to a preset category. That is, the trained image classification model obtained by the image classification model training method provided by the present invention has higher robustness.

[0047] For step S101, the sample image may be an image in a training sample set of a public image classification model, such as a Cifar 100 training set, a Texture dataset, a SVHN (Street View House Number) dataset, etc., which is not limited in the present invention.

[0048] It is also understandable that the sample image can be selected from one training sample set or from multiple training sample sets, and the present invention is not limited thereto. For each sample image obtained, the electronic device can train the image classification model based on the sample image according to the image classification model training method provided by the present invention.

[0049] After acquiring the sample image, the electronic device may divide the sample image according to a preset division method to obtain multiple sample image regions. For example, the division method may indicate: according to a 3×3 division method, the image is divided into 9 non-overlapping image regions of the same size. For example, see Figure 2 , Figure 2 A schematic diagram of a VIT model training method according to an embodiment of the present invention. Figure 2 In the example, the sample image is divided into 9 sample image regions according to a 3×3 division method, which are respectively denoted as “1”, “2”, “3”, …, and “9”.

[0050] Alternatively, the division method may also indicate that the size of an image region is 16×16, and the image is divided into a plurality of non-overlapping image regions. The division methods in actual scenes are far more than this, and the present invention is not limited to this.

[0051] With respect to step S102 and step S103, after obtaining a plurality of sample image regions, the electronic device may sort the plurality of sample image regions. The sorted sample image regions are a region sequence composed of the sample image regions. The manner in which the electronic device sorts the plurality of sample image regions may be referred to in detail in the subsequent embodiments.

[0052] The image classification model consists of a transformation encoder, a category detection network, and a location detection network.

[0053] Based on the region sequence, input data of a conversion encoder can be obtained. Exemplarily, the conversion encoder can be a Transformer encoder.

[0054] Specifically, for each sample image area, the electronic device may use a linear projection method to map the pixel matrix of the sample image area into a vector, that is, to obtain a projection feature corresponding to the sample image area.

[0055] Then, according to the feature dimension corresponding to the projection feature, that is, the number of elements in the mapped vector (hereinafter referred to as the number of elements), and the preset position where the sample image area is actually located in the sample image area, the electronic device can generate a position embedding feature indicating the preset position where the sample image area is located in the sample image. For example, a position number can be set for each preset position, and the electronic device can generate a vector carrying the position number as a position embedding feature of the sample image area located at the preset position; or, the electronic device can also encode the position number of each preset position according to a preset position encoding formula, and then generate a vector carrying the encoding result as a position embedding feature of the sample image area located at the preset position. The present invention is not limited to this.

[0056] For example, the number of elements is 4, and the electronic device divides the sample image into 3 sample image areas from left to right. For the leftmost sample image area, the electronic device can generate a vector (1, 0, 0, 0) as the position embedding feature corresponding to the leftmost sample image area. Similarly, a vector (0, 1, 0, 0) is generated as the position embedding feature corresponding to the middle sample image area; and a vector (0, 0, 1, 0) is generated as the position embedding feature corresponding to the right sample image area.

[0057] Alternatively, the number of elements is 2. Under the above division method, the electronic device can generate a vector (1, 0) as the position embedding feature corresponding to the leftmost sample image area; generate a vector (0, 1) as the position embedding feature corresponding to the middle sample image area; and generate a vector (1, 1) as the position embedding feature corresponding to the right sample image area.

[0058] Or, if the number of elements is 2, as above Figure 2 In the division method shown, when the sample image is divided into 9 sample image areas, the electronic device can generate a vector (1, 1) as a positional embedding feature corresponding to sample image area "1"; generate a vector (1, 2) as a positional embedding feature corresponding to sample image area "2"; generate a vector (1, 3) as a positional embedding feature corresponding to sample image area "3"; generate a vector (2, 1) as a positional embedding feature corresponding to sample image area "4"; and so on, generate a vector (3, 3) as a positional embedding feature corresponding to sample image area "9".

[0059] Obviously, in actual scenarios, there are far more ways to generate position embedding features corresponding to each sample image region than this. As long as a vector indicating the actual position of a sample image region in the sample image is generated according to a set generation method, it can be used as the position embedding feature corresponding to the sample image region. Therefore, the present invention does not limit the specific method of generating the position embedding feature corresponding to the sample image region.

[0060] Then, the electronic device can calculate the sum of the projection feature and the position embedding feature corresponding to the sample image area, that is, obtain the input data corresponding to the sample image area. For example, based on the example in which the number of elements is 4, if the projection feature corresponding to the leftmost sample image area is (1, 2, 3, 4), then the input data corresponding to the sample image area is (2, 2, 3, 4).

[0061] In addition, the electronic device can also generate a learnable class marker feature and a position embedding feature corresponding to the class marker feature. Then, the sum of the class marker feature and the corresponding position embedding feature is calculated to obtain the input data corresponding to the learnable class marker feature (hereinafter referred to as learnable data). The learnable data and the input data corresponding to each sample image region in the region sequence are concatenated to obtain the input data of the conversion encoder in the image classification model.

[0062] The learnable class marking feature indicates the initial probability that the sample object is each preset category; the corresponding position embedding feature indicates the virtual position corresponding to the learnable class marking feature. For example, the electronic device can generate: a vector containing the number of elements, and each element is a preset value, as a learnable class marking feature; or, the electronic device can also randomly determine the number of elements from a preset numerical range as each element in the learnable class marking feature. Similarly, the electronic device can also generate: a vector containing the number of elements, and each element is an initial value, as a corresponding position embedding feature; or, the electronic device can also randomly determine the number of elements from a preset numerical range as each element in the corresponding position embedding feature. The present invention is not limited to this.

[0063] Subsequently, the conversion encoder will update the class marking features based on the extraction results during the process of feature extraction for the region sequence. Since the present invention subsequently adjusts the model parameters of the image classification model based on the position loss value, the image classification model can learn the local information of each image region, the global information of the image, and the positional relationship between the image regions (hereinafter collectively referred to as image information). After adjusting the model parameters of the conversion encoder in the image classification model in the above manner, the conversion encoder will also perform feature extraction based on the above image information during the process of feature extraction for the region sequence, that is, update the class marking features based on the above image information, and the final class marking features can learn the contour information of the objects in the image, as well as the detail information of the objects.

[0064] Exemplarily, a sample image can be represented as: . Wherein, R represents the pixel matrix of the sample image; the pixel matrix includes H rows of pixels, each row includes W columns of pixels; C represents the number of image channels of the sample image.

[0065] The divided sample image area can be expressed as: .in, , represents the number of sample image areas; the pixel matrix of each divided sample image area includes P rows of pixels, and each row includes P columns of pixels, that is, the resolution of a sample image area is P×P.

[0066] For the sample image area Perform linear projection to obtain the corresponding projection features ; For the sample image area Perform linear projection to obtain the corresponding projection features , and so on, for the sample image area Perform linear projection to obtain the corresponding projection features That is, the projection features corresponding to each sample image area are obtained, which can be recorded as , where D represents the feature dimension of each projection feature.

[0067] The learnable class labeling features can be written as , embed the learnable class label features into the projection features corresponding to each sample image area, and the embedding result can be recorded as: .

[0068] For each sample image region, the electronic device may generate a position embedding feature indicating the preset position of the sample image region in the sample image according to the generation method of the position embedding feature described in the above embodiment. And the electronic device may also obtain the position embedding feature corresponding to the class marker feature according to the generation method of the position embedding feature corresponding to the learnable class marker feature described above. Then the position embedding features corresponding to the class mark features and the position embedding features corresponding to each sample image area are spliced, and the splicing result can be recorded as: .in, Represents the sample image area The corresponding position embedding features; Represents the sample image area The corresponding position embedding features, and so on, Represents the sample image area The corresponding position embedding features.

[0069] Then, the sum of the concatenation result P and the embedding result Z is calculated, that is, (Z+P), and the input data V of the conversion encoder can be obtained. The input data V can be written as: ,in, ; ; , and so on, .

[0070] Then, the input data V is subjected to feature extraction processing using the conversion encoder to obtain high-dimensional sample image features.

[0071] With respect to step S104, after acquiring the high-dimensional sample image features, the electronic device may input the sample image features into a category detection network to obtain a category detection result output by the category detection network. If the number of output channels of the category detection network may be the number of categories of a preset category, the category detection network may map the high-dimensional sample image features to the dimension of the number of categories, that is, perform category detection on the sample image features to obtain a category detection result. Figure 2 As shown, the classification output of the VIT model, i.e., the category detection result. For example, the category detection network may be a prediction head including a fully connected layer, and the prediction head may be an MLP (Multi-Layer Perceptron).

[0072] The category detection result can represent the probability that the sample object belongs to each preset category. For example, the category detection result can be represented as a vector (hereinafter referred to as the category vector), and the elements in the category vector (hereinafter referred to as the category elements) correspond to the preset categories one by one, such as the first category element corresponds to the first preset category, the second category element corresponds to the second preset category, and so on. If the preset categories include: category A, category B, the category vector is (0.3, 0.7). That is, the probability that the sample object belongs to category A is 0.3; the probability that it belongs to category B is 0.7.

[0073] With respect to step S105, when the sample image is divided according to a preset division method to obtain each sample image region, the position of a sample image region in the sample image is a preset position. Figure 2 As shown, when the sample image is divided into 9 sample image areas according to the 3×3 division method, the position of sample image area "1" in the sample image can be recorded as preset position 1; the position of sample image area "2" in the sample image can be recorded as preset position 2; the position of sample image area "3" in the sample image can be recorded as preset position 3; and so on, the position of sample image area "9" in the sample image can be recorded as preset position 9. In this way, the sample image is divided according to the preset division method, and 9 preset positions in the sample image are determined.

[0074] After obtaining the high-dimensional sample image features, the electronic device can input the sample image features into the position detection network to obtain the position detection result output by the position detection network. For example, the number of output channels of the position detection network can be the number of regions of the sample image region, and the position detection network can map the high-dimensional sample image features to the dimension of the number of regions, that is, perform position detection on the sample image features to obtain the position detection result. For example, the position detection network can be an MLP.

[0075] The position detection result may include the position detection result of each sample image region (hereinafter referred to as region detection result). For each sample image region, the region detection result of the sample image region may represent: the predicted probability that the sample image region is located at each preset position in the sample image.

[0076] Exemplarily, the position detection result may include: a region detection result for each sample image region, that is, the position detection network may output region detection results as many as the regions, and one region detection result corresponds to one sample image region. A region detection result may be represented as a vector (hereinafter referred to as the position vector). The elements in the position vector (hereinafter referred to as the position elements) correspond to preset positions in the sample image, such as the first position element corresponds to preset position 1 in the sample image, the second position element corresponds to preset position 2 in the sample image, and so on. A position element in the position vector of a sample image region represents: the probability that the sample image region is located in the sample image at the preset position corresponding to the position element. For example Figure 2 As shown, the preset positions include: preset position 1 to preset position 9, and the position vector corresponding to the sample image area 1 is (0.9, 0.1, 0, 0, 0, 0, 0, 0, 0). That is, the probability that the sample image area 1 is located at the preset position 1 is 0.9; the probability that it is located at the preset position 2 is 0.1; and the probability that it is located at the preset position 3 to the preset position 9 is 0.

[0077] It is understandable that the present invention does not limit the execution order of the above-mentioned step S104 and step S105. For example, step S104 may be executed first, and then step S105; or, step S105 may be executed first, and then step S104; or, step S104 and step S105 may be executed in parallel, which are all reasonable.

[0078] For step S106, the electronic device may also obtain a category label indicating the preset category to which the sample object in the sample image belongs. Exemplarily, the category label may be represented as a vector (hereinafter referred to as the first label vector), and the elements in the first label vector (hereinafter referred to as the first label element) correspond one-to-one to the preset categories, such as the first first label element corresponds to the first preset category, the second first label element corresponds to the second preset category, and so on. A first label element indicates the probability that the sample object belongs to the preset category corresponding to the first label element. If the preset categories include: category A, category B, and the sample object belongs to category B, then the category label may be (0, 1).

[0079] Then, the electronic device can obtain the difference between the category label and the category detection result (hereinafter referred to as the category difference). For example, when the category label and the category detection result are both represented as vectors, the electronic device can calculate the vector distance between the category label and the category detection result, that is, obtain the category difference.

[0080] The electronic device may also obtain a position label indicating the preset position of each sample image region in the sample image. Exemplarily, the electronic device may obtain a position label (hereinafter referred to as a region position label) for each sample image region, and a region position label may be represented as a vector (hereinafter referred to as a second label vector). The elements in the second label vector (hereinafter referred to as second label elements) correspond to preset positions in the sample image, such as the first second label element corresponds to preset position 1 in the sample image, the second second label element corresponds to preset position 2 in the sample image, and so on. A second label element in the region position label of a sample image region indicates: the probability that the sample image region is located at the preset position corresponding to the second label element in the sample image. As in the above example, in the case where the preset positions include: preset positions 1 to preset positions 9, the sample image region 1 belongs to the preset position 1, and the position label corresponding to the sample image region 1 may be (1, 0, 0, 0, 0, 0, 0, 0). Alternatively, the electronic device may also obtain a matrix composed of the region position labels, and a row vector in the matrix represents a region position label. Subsequently, the electronic device may search for the region position label of each sample image region from the matrix.

[0081] Furthermore, the electronic device can obtain the difference between the position label and the position detection result (hereinafter referred to as the total position difference). Exemplarily, when the position label and the position detection result are both represented as vectors, for each sample image area, the electronic device can calculate: the vector distance between the regional position label of the sample image area and the position detection result of the sample image area as the position difference corresponding to the sample image area; calculate the sum of the position differences corresponding to each sample image area to obtain the total position difference.

[0082] Then, the electronic device can adjust the model parameters of the image classification model of the initial structure based on the category difference and the total position difference until the preset convergence condition is reached to obtain a trained image classification model, that is, the parameters of the feature encoder, the category detection network, and the position detection network are adjusted.

[0083] For example, the electronic device may adjust the model parameters based on the category difference according to the preset adjustment method, and then adjust the model parameters based on the total position difference. Alternatively, the electronic device may also fuse the category difference and the total position difference, and then adjust the model parameters based on the fusion result according to the preset adjustment method. For the specific fusion method, please refer to the detailed description of the subsequent embodiments. The preset adjustment method may be a gradient descent method, a grid search method, etc., which is not limited by the present invention.

[0084] The preset convergence condition may include any of the following: the number of training times reaches the preset number of training times, and the difference between the loss function value calculated this time and the loss function value calculated last time is less than a preset difference threshold.

[0085] In some embodiments, the image classification model is a VIT model; the conversion encoder is a Transformer encoder; the category detection network includes a fully connected layer; and the position detection network is an MLP.

[0086] The image classification model can be a basic VIT model, or other versions of the VIT model improved on the basis of the basic VIT model, and the present invention is not limited thereto. For example, the image classification model can be a VIT-B / 16 model, a VIT-L / 32 model, a DeiT-Small (Data-efficient Image Transformer-Small, a small and efficient image data converter), and the like.

[0087] The Transformer encoder includes multiple modules, each of which includes LN (Layer Normalization), MSA (Multi-Head Self-Attention), and MLP. If the Transformer encoder includes N modules, the output sample image features of the Transformer encoder can be expressed as:

[0088] ;

[0089] , ;

[0090] st

[0091] , ;

[0092] in, is the class prediction result corresponding to the aforementioned learnable class label feature, that is, the aforementioned final class label feature; MLP(·) represents the output data of MLP, and LN(·) represents the output data of LN; Represents the output of the Nth module in the Transformer encoder ; Represents learnable data that has learned the contour information of the object in the aforementioned image, as well as the detailed information of the image; represents the output data of the i-th module; MSA(·) represents the output data of MSA.

[0093] The number of output channels of the fully connected layer is the number of categories. Through the fully connected layer, the high-dimensional sample image features can be mapped to the dimension of the number of categories, that is, the category detection result can be obtained.

[0094] The position detection network may include one layer of MLP or multiple layers of MLP. When the position detection network includes one layer of MLP, the number of output channels of the MLP is the number of regions; when the position detection network includes multiple layers of MLP, the number of output channels of the last layer of MLP is the number of regions, and the number of input channels of the other layers of MLP is the same as the number of output channels.

[0095] It can be understood that the more layers of MLP that can be included in the position detection network, the longer it takes for the position detection network to process the sample image features, the higher the accuracy of the position detection result, and the better the subsequent training effect of the image classification model based on the more accurate position detection result, that is, the higher the robustness of the trained image classification model. Therefore, in actual scenarios, the number of layers of MLP included in the position detection network can be determined according to business needs and the hardware performance of electronic devices. For example, when it is required to reduce the time spent on obtaining the trained image classification model and improve the efficiency of obtaining the image classification model, the position detection network can include a smaller number of layers of MLP, such as one layer of MLP; when the robustness of the trained image classification model is required to be higher, the position detection network can include a larger number of layers of MLP, such as three layers of MLP.

[0096] In some embodiments, the above step S202 may include the following steps: arranging the sample image areas in a preset order to obtain a region sequence; or arranging the sample image areas in a preset order, and shuffling the arrangement results using a shuffling algorithm; combining the arrangement results and the shuffling results to obtain a region sequence; or arranging the sample image areas in a preset order; shuffling the arrangement results using a shuffling algorithm to obtain a region sequence.

[0097] The sample image regions are arranged in a preset order, and the arrangement result obtained can be called an initial sequence; the arrangement result is shuffled using a shuffling algorithm, and the shuffled result obtained can be called a shuffled sequence.

[0098] The preset order can be set by the technician according to the requirements. Taking the above-mentioned division of the sample image in the 3×3 division method as an example, the preset order can be: from top to bottom, from left to right. For example, Figure 2 As shown in FIG. 1 , the 9 sample image regions obtained by division are the division results; the sample image regions are arranged in order from top to bottom and from left to right, and the Figure 2 The initial sequence in can be recorded as "123456789".

[0099] The initial sequence is shuffled using a shuffle algorithm, that is, the order of each sample image region in the initial sequence is adjusted to obtain a shuffle sequence. For example, the shuffle algorithm may be a Knuth-Durstenfeld Shuffle algorithm, a Fisher-Yates algorithm, etc., and the present invention is not limited thereto.

[0100] For example, after the initial sequence is shuffled by the shuffling algorithm, the order of the sample image regions in the shuffled sequence is: 973546281, that is, the shuffled sequence can be recorded as "973546281". The order of the sample image regions in the shuffled sequence can also be recorded as the position of each network prediction block, such as the first network prediction block in the shuffled sequence (i.e., sample image region "9") is actually located at the preset position 9 in the sample image.

[0101] When the sample image regions are arranged in a preset order to obtain a region sequence, the region sequence only includes the initial sequence. That is, the image classification model is always trained based on the image regions arranged in the same order, which can reduce the difficulty of the image classification model to learn the positional relationship between image regions and improve the model training efficiency.

[0102] When the shuffle algorithm is used to shuffle the arrangement results to obtain the region sequence, the region sequence only includes the shuffle sequence. That is, the image classification model is trained using the image regions arranged in a random order, which increases the training difficulty of the image classification model, improves the accuracy of the positional relationship between the image regions learned by the image classification model, and thus improves the accuracy of the trained image classification model.

[0103] For example, Figure 2 As shown, after obtaining the shuffle sequence, the shuffle sequence can be processed based on the image classification model to obtain a position output indicating the actual position of each sample image area in the sample image, that is, the predicted probability of each sample image area being located at each preset position in the sample image. As shown in the position output, the order of the predicted probabilities corresponding to each sample image area is consistent with the order of each sample image area in the shuffle sequence. That is, the first predicted probability corresponds to the sample image area "9"; the second predicted probability corresponds to the sample image area "7", and so on. Subsequently, for each sample image area, the position loss value corresponding to the sample image area can be calculated based on the difference between the position output of the sample image area and the position label; based on the position loss values ​​corresponding to each sample image area, the total position loss value is calculated.

[0104] In the case where the region sequence is obtained by combining the arrangement result and the shuffle result, that is, for each sample image, the electronic device will obtain the initial sequence and the shuffle sequence corresponding to the sample image. Then, the image classification model is trained using the initial sequence and the shuffle sequence respectively. For example, the image classification model can be trained using the initial sequence first, and then the image classification model can be trained using the shuffle sequence. In this way, since the shuffle sequence is a difficult example corresponding to the initial sequence, the image classification model is trained based on the initial sequence and the shuffle sequence together, which can improve the accuracy of the positional relationship between the image regions learned by the image classification model, thereby improving the accuracy of the trained image classification model.

[0105] In some embodiments, Figure 1 Based on Figure 3 , the above step S106 may include the following steps:

[0106] S1061: Using a preset loss function, calculate a position loss value representing the difference between the position label and the position detection result, and a category loss value representing the difference between the category label and the category detection result.

[0107] S1062: According to the preset position loss weight and category loss weight, calculate the weighted sum of the position loss value and the category loss value to obtain the total loss value.

[0108] S1063: Adjust the model parameters of the image classification model of the initial structure based on the total loss value until a preset convergence condition is reached to obtain a trained image classification model.

[0109] The preset loss function may be a CE (Cross Entropy) loss function, an MSE (Mean Squared Error) loss function, etc., and the present invention is not limited to this.

[0110] The position loss weight and the category loss weight can be data measured by technicians through experiments. For example, technicians can set the category loss weight to a constant 1 and set the position loss weight to a hyperparameter. , test the hyperparameters experimentally The relationship between the model training effect and the model training effect, and determine the hyperparameters when the model training effect best meets the business needs of the actual scenario. The value of is the position loss weight. For example, the position loss weight and the category loss weight can both be 1.

[0111] For example, taking the preset loss function as the CE loss function as an example, the total loss value can be calculated based on the following formula:

[0112] ;

[0113] in, represents the total loss value; Represents the classification loss value; Indicates the category detection result; Represents the category label; represents the position loss weight; Represents the position loss value; Indicates the position detection result; Represents a location tag.

[0114] Based on the above processing, the electronic device can fuse the category difference and the total position difference, and then adjust the model parameters of the image classification model based on the total difference obtained by the fusion. When the image classification model learns the information of the category of the object in the sample image, it can also learn the local information of each sample image area, the global information of the sample image, and the positional relationship between the sample image areas. Then, when the image classification is subsequently performed based on the image classification model trained in this way, the features that can be extracted by the trained image classification model can indicate: the local features of each image area, the global features of the object in the image, and the positional relationship between the image areas. Image classification is performed based on global features, that is, image classification is performed according to the contour of the object in the image; image classification is performed based on the local features of each image area and the positional relationship between the image areas, that is, image classification is performed according to the details of the object. Image classification is performed by integrating the contour of the object and the details of the object, that is, the image classification is performed by integrating the contour of the object and the details of the object, that is, the category detection result is obtained by integrating more abundant information indicating the category to which the object in the image belongs, which can make up for the defect that the image classification is performed only based on the contour of the object, resulting in the object being misdetected as belonging to the preset category when the contour of the object is similar to the contour corresponding to a preset category. That is, the trained image classification model obtained by the image classification model training method provided by the present invention has higher robustness.

[0115] After the trained image classification model is obtained, the performance of the trained image classification model obtained based on the present invention can be tested through experiments.

[0116] In the experiment, the trained image classification model is obtained by training DeiT-Small based on the image classification model training method provided by the present invention. The electronic device obtains sample images from the Cifar100 training set.

[0117] The input resolution of DeiT-Small is 224×224, that is, images with a resolution of 224×224 can be classified. Based on DeiT-Small, the size of each image block (i.e., the sample image area in the aforementioned embodiment) is 16×16, and since 224 divided by 16 is 14, the sample image is divided into 14×14 image blocks.

[0118] After the encoder included in DeiT-Small obtains the sample image features, the sample image features are input into a prediction head including a fully connected layer (i.e., the category detection network in the aforementioned embodiment). The fully connected layer in the prediction head maps the sample image features to the dimension of the number of categories, that is, the category detection network is used to perform category detection on the sample image features to obtain the category detection results.

[0119] The sample image features are input into the three-layer MLP head (i.e., the position detection network in the aforementioned embodiment), and the input dimensions of the first two layers of the three-layer MLP head are equal to the output dimensions; the output dimension of the third layer of the MLP head is equal to the total number of image blocks (i.e., the number of regions in the aforementioned embodiment), and the output data of the third layer of the MLP head is the position detection result in the aforementioned embodiment.

[0120] Based on the difference between the category label and the category detection result, and the difference between the position label and the position detection result, the DeiT-Small of the initial structure is trained to obtain the trained DeiT-Small.

[0121] Then, the electronic device determines the verification image from the verification set, and classifies the verification image through the trained DeiT-Small. Since the actual scene is extremely complex, the preset categories cannot cover all the categories that may exist in the actual scene, that is, there are unknown categories other than the preset categories. Therefore, in order to test the classification effect of the trained DeiT-Small on images including objects of unknown categories, a data set in which the categories to which the objects in the included images belong are different from the categories to which the objects in the images included in the training set belong can be used as a verification set. For example, the verification set can be: Cifar100 verification set, Texture data set, SVHN data set, Places 365 data set, LSUN_C (Large-scale Scene Understanding_Classification, large-scale scene understanding and classification) data set, LSUN_Resize (large-scale scene understanding and adjustment) data set, iSUN (a saliency data set), etc.

[0122] The verification results are shown in the following table (1):

[0123] Table (1)

[0124] AUROC FPR95 ACC VIT 91.67 33.20 90.47 Ours 95.54 16.40 90.57

[0125] Among them, the larger the AUROC (Area Under the Receiver Operating Characteristic Curve), the higher the accuracy of the model; the smaller the FPR95 (False Positive Rateat 95% True Positive Rate), the higher the accuracy of the model; AUROC and FPR95 are used to evaluate: the accuracy of the model when the model classifies images including objects of unknown categories; the increase of AUROC and the decrease of FPR95 can indicate that the robustness of Ours (the trained DeiT-Small obtained based on the image classification model training method provided by the present invention) is improved. ACC (Accuracy) represents the overall classification effect of the model.

[0126] In one training method, for example, see Figure 4 , Figure 4 The figure is a schematic diagram of training a VIT model based on a training method. Based on this training method, the sample image can be divided into 9 image regions according to a 3×3 division method, which are respectively recorded as "1", "2", "3", ..., "9". Then, the 9 divided image regions are arranged in sequence from top to bottom and from left to right to form a region sequence; the region sequence is subjected to image classification processing using the VIT model to obtain the classification output of the VIT model, that is, the probability that the object in the image belongs to each preset category. Then, the model parameters of the VIT model can be adjusted only based on the difference between the classification output of the VIT model and the label indicating the preset category to which the object actually belongs, until the VIT model converges, and a VIT model trained based on this training method is obtained. In this training method, the VIT model does not include a position detection network, does not output a position detection result, and does not adjust the model parameters of the VIT model based on the difference between the position detection result and the position label.

[0127] The conversion encoder in the VIT model trained based on this training method will only: classify images based on global features, that is, classify images only according to the contours of objects in the image. When an object of an unknown category has a similar contour to an object of a preset category, using the VIT model obtained based on this training method to classify the image to be classified, that is, classifying the image based on the contours of the object of the unknown category, may result in false positives. That is, the robustness of the VIT model obtained based on this training method is not high.

[0128] The trained image classification model obtained based on the image classification model training method provided by the present invention can learn the local information of each image area, the global information of the image, and the positional relationship between each image area (that is, the aforementioned image information). Subsequently, when image classification is performed based on the trained image classification model, the trained image classification model can extract image features based on the learned image information, and can extract image features indicating the local features of each image area, the global features of the object in the image, and the positional relationship between each image area, that is, extract more general image features. Compared with only extracting image features indicating the outline of the object, the information of the image described by the image features extracted in this way is more comprehensive and more accurate. That is, the position learning process of the image block is added in the VIT model training process, so that the VIT model considers the position prediction of the image block when classifying the target, which is conducive to the VIT model learning the positional relationship between the image blocks, so that the VIT model has a deeper understanding of the image, thereby improving the performance of the trained VIT model.

[0129] Image classification is performed based on global features, that is, image classification is performed according to the contours of objects in the image; image classification is performed based on the local features of each image area and the positional relationship between each image area, that is, image classification is performed according to the details of the object. Image classification is performed by integrating the contours of the object and the details of the object, that is, the image classification is performed by integrating the contours of the object and the details of the object, that is, more abundant information indicating the category to which the object in the image belongs is integrated to obtain a category detection result, which can make up for the defect of image classification based only on the contour of the object, resulting in the object being misdetected as belonging to a preset category when the contour of the object is similar to the contour corresponding to a preset category, and improve the accuracy of the obtained image classification result. This will improve the user experience and enhance the user satisfaction with the trained image classification model obtained by using the training method of the image classification model provided by the present invention.

[0130] Compared with this training method, the training method of the image classification model provided by the present invention can obtain a trained image classification model with higher robustness, that is, the robustness of the model to position samples is greatly improved. And because the network structure of the VIT model does not need to be adjusted in the present invention, it can be applied to the VIT models of various networks. In the training process of the VIT model, as long as the auxiliary training branch is added, the VIT model can be trained according to the training method of the image classification model provided by the present invention. The method is more versatile, and does not increase unnecessary network parameters and computational complexity, which can avoid reducing the efficiency of model training.

[0131] The present invention also provides an image classification method, which is applied to an electronic device. The electronic device may be the aforementioned electronic device for executing the image classification model training method, or may be another device. To distinguish it from the aforementioned electronic device for executing the image classification model training method, the electronic device for executing the image classification method is hereinafter referred to as a classification device. After the electronic device obtains a trained image classification model according to any of the image classification model training methods in the above embodiments, the trained image classification model may be deployed in the classification model, and then image classification may be performed based on the locally deployed trained image classification model.

[0132] See also Figure 5 , Figure 5 A flow chart of an image classification method provided by an embodiment of the present invention. The method may include the following steps:

[0133] S501: Acquire an image to be classified, and divide the image to be classified according to a preset division method to obtain regions of the image to be classified.

[0134] S502: Obtain a region sequence consisting of the divided image regions to be classified.

[0135] S503: Using the conversion encoder in the trained image classification model to extract features from the acquired region sequence to obtain features of the image to be classified.

[0136] Among them, the image classification model is obtained based on any image classification model training method in the aforementioned embodiments; the image classification model also includes a category detection network.

[0137] S504: Perform category detection on the features of the image to be classified using a category detection network to obtain a category detection result.

[0138] The category detection result indicates the probability that the object in the image to be classified belongs to each preset category.

[0139] Based on the image classification method provided by the present invention, since the image classification model is obtained by joint training based on the positional relationship between each image area divided according to a preset division method and the preset category to which the object in the image belongs, the trained image classification model can learn the local information of each image area, the global information of the image, and the positional relationship between each image area. Therefore, when image classification is performed based on the trained image classification model, the features extracted by the trained image classification model can indicate: the local features of each image area, the global features of the object in the image, and the positional relationship between each image area. Image classification is performed based on global features, that is, image classification is performed according to the contour of the object in the image; image classification is performed based on the local features of each image area and the positional relationship between each image area, that is, image classification is performed according to the details of the object. Image classification is performed based on the contour of the object and the details of the object, that is, more abundant information indicating the category to which the object in the image belongs is combined to obtain a category detection result, which can make up for the defect of performing image classification based only on the contour of the object, which leads to the object being misdetected as belonging to the preset category when the contour of the object is similar to the contour corresponding to a preset category, and improves the accuracy of the obtained image classification result.

[0140] With respect to step S501 and step S502, the image to be classified is an image whose category of the object needs to be detected, such as an image input by a user, or a video frame sampled from a video stream by a classification device, etc., and the present invention is not limited to this.

[0141] After acquiring the image to be classified, the classification device can divide the image to be classified into multiple image regions to be classified according to the preset division method used when training the image classification model. If the sample image is divided into 9 non-overlapping and equal-sized sample image regions according to the 3×3 division method when training the image classification model, the classification device will also divide the image to be classified into 9 non-overlapping and equal-sized image regions to be classified according to the 3×3 division method.

[0142] The preset order can be set by the technician according to the requirements. For example, according to the 3×3 division method, the image to be classified is divided into 9 image areas to be classified, which are respectively recorded as "1", "2", "3", ..., "9"; the preset order can be: from top to bottom, from left to right.

[0143] The classification device may arrange the images to be classified in a preset order to obtain a region sequence. In this case, the region sequence may be recorded as "123456789". Alternatively, the classification device may also shuffle the above arrangement results using a shuffling algorithm and use the shuffled results as the region sequence. The present invention is not limited to this.

[0144] For step S503, based on the region sequence, input data of the conversion encoder can be obtained. Exemplarily, the conversion encoder can be a Transformer encoder.

[0145] According to the feature dimension corresponding to the projection feature, that is, the number of elements in the mapped vector, and the preset position where the image region to be classified is actually located in the image region to be classified, a position embedding feature indicating the preset position where the image region to be classified is located in the image to be classified is generated. Then, the sum of the projection feature and the position embedding feature corresponding to the image region to be classified is calculated, and the input data corresponding to the image region to be classified is obtained.

[0146] For each image region to be classified in the region sequence, the classification device can perform linear projection on the image region to be classified to obtain the projection feature corresponding to the image region to be classified; according to the feature dimension corresponding to the projection feature and the preset position where the image region to be classified is actually located in the image to be classified, the position embedding feature corresponding to the image region to be classified is generated. The sum of the projection feature and the position embedding feature corresponding to the image region to be classified is calculated to obtain the input data corresponding to the image region to be classified.

[0147] In addition, the classification device can also generate learnable class marker features and position embedding features corresponding to the class marker features. Then, the sum of the class marker features and the corresponding position embedding features is calculated to obtain the input data corresponding to the learnable class marker features. The input data corresponding to the learnable class marker features and the input data corresponding to each image region to be classified in the region sequence are concatenated to obtain the input data of the conversion encoder in the image classification model (hereinafter referred to as the data to be extracted).

[0148] The learnable class marker feature indicates the initial probability that the object to be classified is each preset category; the corresponding position embedding feature indicates the virtual position corresponding to the learnable class marker feature. For example, the classification device can generate a vector containing the number of elements, each of which is a preset value, as a learnable class marker feature; and generate a vector containing the number of elements, each of which is an initial value, as a corresponding position embedding feature. In actual scenarios, there are many ways to generate learnable class marker features and corresponding position embedding features, which are not limited by the present invention.

[0149] The feature extraction of the data to be extracted is performed by converting the encoder to obtain the features of the image to be classified.

[0150] For step S504, exemplarily, the category detection network may be a prediction head including a fully connected layer, and the prediction head may be an MLP. After acquiring the features of the image to be classified, the classification device may input the features of the image to be classified into the category detection network to obtain the category detection result output by the category detection network. For example, the number of output channels of the category detection network may be the number of categories of a preset category, and the category detection network may map the features of the image to be classified to the dimension of the number of categories, that is, perform category detection processing on the features of the image to be classified, and obtain the output of the category detection network: the probability that the object to be classified in the image to be classified is each preset category, and obtain the category detection result.

[0151] Subsequently, the classification device may post-process the category detection result to obtain the category to which the object to be classified belongs. If the difference between the probabilities of any two preset categories is less than the difference threshold, a detection result indicating that the object to be classified is not a preset category is obtained, such as outputting "unknown category". The difference threshold may be an empirical value, such as 0.1.

[0152] If the probability that the object to be classified is a preset category is greater than a probability threshold, it can be determined that the object to be classified is the preset category. The probability threshold can be an empirical value, such as 0.95.

[0153] Based on the same inventive concept as the above-mentioned image classification model training method, the present invention also provides an image classification model training device, see Figure 6 , Figure 6 A structural diagram of an image classification model training device provided by an embodiment of the present invention. The device comprises:

[0154] The image acquisition module 601 is used to acquire a sample image and divide the sample image according to a preset division method to obtain each sample image area;

[0155] A sequence acquisition module 602 is used to acquire a region sequence composed of each sample image region obtained by division;

[0156] A feature extraction module 603 is used to extract features from the acquired region sequence using a conversion encoder in an image classification model of an initial structure to obtain sample image features; wherein the image classification model further includes a category detection network and a position detection network;

[0157] The category detection module 604 is used to perform category detection on the sample image features using the category detection network to obtain a category detection result; wherein the category detection result indicates: the probability that the sample object in the sample image belongs to each preset category;

[0158] The position detection module 605 is used to perform position detection on the sample image features using the position detection network to obtain a position detection result; wherein the position detection result represents: the predicted probability that each sample image region is located at each preset position in the sample image; each preset position is determined by dividing the sample image according to the preset division method;

[0159] The training module 606 is used to adjust the model parameters of the image classification model of the initial structure based on the difference between the category label and the category detection result, and the difference between the position label and the position detection result, until the preset convergence condition is reached to obtain a trained image classification model; wherein the category label indicates the preset category to which the sample object belongs; and the position label indicates the preset position of each sample image area in the sample image.

[0160] Optionally, the sequence acquisition module 602 is specifically used to: arrange the sample image areas in a preset order to obtain a region sequence; or, arrange the sample image areas in a preset order, and shuffle the arrangement results using a shuffling algorithm; combine the arrangement results and the shuffling results to obtain a region sequence; or, arrange the sample image areas in a preset order; and shuffle the arrangement results using a shuffling algorithm to obtain a region sequence.

[0161] Optionally, the training module 606 is specifically used to: use a preset loss function to calculate a position loss value representing the difference between the position label and the position detection result, and a category loss value representing the difference between the category label and the category detection result; calculate a weighted sum of the position loss value and the category loss value according to preset position loss weights and category loss weights to obtain a total loss value; and adjust model parameters of the image classification model of the initial structure based on the total loss value until a preset convergence condition is reached to obtain a trained image classification model.

[0162] Optionally, the image classification model is a visual transformer (VIT) model; the conversion encoder is a conversion network (Transformer) encoder; the category detection network includes a fully connected layer; and the position detection network is a multi-layer perceptron (MLP).

[0163] Based on the image classification model training device provided by the embodiment of the present invention, when training the image classification model, after the electronic device obtains the sample image, it will divide the sample image into multiple sample image areas according to the preset division method, and then obtain the area sequence composed of the divided sample image areas. The image classification model includes a conversion encoder, a category detection network and a position detection network. After the sample image features of the area sequence are extracted by the conversion encoder, the sample image features are subjected to category detection by the category detection network to obtain the category detection result, that is, the probability that the sample object in the sample image is each preset category is obtained; according to the difference between the category detection result and the category label representing the preset category to which the sample object belongs, the category loss value can be obtained. The sample image features are subjected to position detection by the position detection network to obtain the position detection result, that is, the predicted probability that each sample image area in the area sequence is located at each preset position in the sample image; according to the difference between the position detection result and the position label, the position loss value can be obtained. The position label indicates the preset position where each sample image area is located in the sample image. Subsequently, the image classification model is trained by combining the category loss value and the position loss value, that is, the image classification model is trained based on the positional relationship between the image regions divided according to the preset division method and the preset category to which the object in the image belongs. Then, the trained image classification model can learn the local information of each image region, the global information of the image, and the positional relationship between the image regions. Subsequently, when image classification is performed based on the trained image classification model, the features that can be extracted by the trained image classification model can indicate: the local features of each image region, the global features of the object in the image, and the positional relationship between the image regions. Image classification is performed based on global features, that is, image classification is performed according to the contour of the object in the image; image classification is performed based on the local features of each image region and the positional relationship between the image regions, that is, image classification is performed according to the details of the object. Image classification is performed by combining the contour of the object and the details of the object, that is, combining more abundant information indicating the category to which the object in the image belongs, and obtaining the category detection result together, which can make up for the defect that image classification is performed based only on the contour of the object, resulting in the object being misdetected as belonging to the preset category when the contour of the object is similar to the contour corresponding to a preset category. That is, the trained image classification model obtained by the image classification model training method provided by the present invention has higher robustness.

[0164] Based on the same inventive concept as the above-mentioned image classification method, the present invention also provides an image classification device, see Figure 7 , Figure 7 A structural diagram of an image classification device provided by an embodiment of the present invention. The device comprises:

[0165] The first acquisition module 701 is used to acquire an image to be classified, and divide the image to be classified according to a preset division method to obtain image regions to be classified;

[0166] The second acquisition module 702 is used to acquire a region sequence composed of the divided image regions to be classified;

[0167] An extraction module 703 is used to extract features from the acquired region sequence using a conversion encoder in a trained image classification model to obtain features of the image to be classified; wherein the image classification model is obtained based on any image classification model training method in the aforementioned embodiments; and the image classification model also includes a category detection network;

[0168] The detection module 704 uses the category detection network to perform category detection on the features of the image to be classified to obtain a category detection result; wherein the category detection result indicates: the probability that the object in the image to be classified belongs to each preset category.

[0169] Based on the image classification device provided by the present invention, since the image classification model is obtained by training based on the positional relationship between each image area divided according to a preset division method and the preset category to which the object in the image belongs, the trained image classification model can learn the local information of each image area, the global information of the image, and the positional relationship between each image area. Therefore, when image classification is performed based on the trained image classification model, the features extracted by the trained image classification model can indicate: the local features of each image area, the global features of the object in the image, and the positional relationship between each image area. Image classification is performed based on global features, that is, image classification is performed according to the contour of the object in the image; image classification is performed based on the local features of each image area and the positional relationship between each image area, that is, image classification is performed according to the details of the object. Image classification is performed based on the contour of the object and the details of the object, that is, more abundant information indicating the category to which the object in the image belongs is synthesized to obtain the category detection result, which can make up for the defect that image classification is performed based only on the contour of the object, resulting in the object being misdetected as belonging to the preset category when the contour of the object is similar to the contour corresponding to a preset category, thereby improving the accuracy of the obtained image classification result.

[0170] The embodiment of the present invention further provides an electronic device, such as Figure 8As shown, it includes a processor 801, a communication interface 802, a memory 803 and a communication bus 804, wherein the processor 801, the communication interface 802 and the memory 803 communicate with each other through the communication bus 804. The memory 803 is used to store computer programs; the processor 801 is used to implement the steps of the image classification model training method described in any of the above embodiments when executing the program stored in the memory 803, or implement the steps of the image classification method described in any of the above embodiments.

[0171] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0172] The communication interface is used for communication between the above electronic device and other devices.

[0173] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0174] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0175] In another embodiment provided by the present invention, a computer-readable storage medium is also provided, which stores a computer program. When the computer program is executed by a processor, it implements the steps of any of the above-mentioned image classification model training methods, or implements the steps of any of the above-mentioned image classification methods.

[0176] In another embodiment provided by the present invention, a computer program product containing instructions is also provided. When the computer is run on a computer, the computer executes any image classification model training method in the above embodiments, or executes any image classification method in the above embodiments.

[0177] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk (SSD)), etc.

[0178] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0179] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, electronic device, computer-readable storage medium, and computer program product embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0180] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A method for training an image classification model, characterized in that: The method comprises: Acquire a sample image, and divide the sample image according to a preset division method to obtain each sample image area; Obtaining a region sequence composed of each divided sample image region; Using the conversion encoder in the image classification model of the initial structure to extract features from the acquired region sequence to obtain sample image features; wherein the image classification model also includes a category detection network and a position detection network; Using the category detection network to perform category detection on the sample image features to obtain a category detection result; wherein the category detection result represents: the probability that the sample object in the sample image belongs to each preset category; Performing position detection on the sample image features using the position detection network to obtain a position detection result; wherein the position detection result represents: a predicted probability that each sample image region is located at each preset position in the sample image; each preset position is determined by dividing the sample image according to the preset division method; Based on the difference between the category label and the category detection result, and the difference between the position label and the position detection result, the model parameters of the image classification model of the initial structure are adjusted until the preset convergence conditions are reached to obtain a trained image classification model; wherein the category label indicates the preset category to which the sample object belongs; and the position label indicates the preset position of each sample image area in the sample image.

2. The method according to claim 1, characterized in that The region sequence formed by each sample image region obtained by division includes: Arranging the sample image regions according to a preset order to obtain a region sequence; or, Arrange the sample image regions according to a preset order, and shuffle the arrangement results using a shuffling algorithm; combine the arrangement results and the shuffling results to obtain a region sequence; or, The sample image regions are arranged in a preset order; the arrangement result is shuffled using a shuffling algorithm to obtain a region sequence.

3. The method according to claim 1, characterized in that The method of adjusting the model parameters of the image classification model of the initial structure based on the difference between the category label and the category detection result, and the difference between the position label and the position detection result, until a preset convergence condition is reached to obtain a trained image classification model, includes: Using a preset loss function, calculating a position loss value representing a difference between a position label and the position detection result, and a category loss value representing a difference between a category label and the category detection result; According to the preset position loss weight and category loss weight, a weighted sum of the position loss value and the category loss value is calculated to obtain a total loss value; Based on the total loss value, the model parameters of the image classification model of the initial structure are adjusted until a preset convergence condition is reached to obtain a trained image classification model.

4. The method according to claim 1, characterized in that: The image classification model is a visual transformer VIT model; the conversion encoder is a conversion network Transformer encoder; the category detection network includes a fully connected layer; and the position detection network is a multi-layer perceptron MLP.

5. An image classification method, characterized in that: The method comprises: Acquire an image to be classified, and divide the image to be classified according to a preset division method to obtain regions of the image to be classified; Obtaining a region sequence composed of each image region to be classified obtained by division; Using the conversion encoder in the trained image classification model to extract features from the acquired region sequence to obtain features of the image to be classified; wherein the image classification model is obtained based on the method described in any one of claims 1 to 4 above; the image classification model also includes a category detection network; The category detection network is used to perform category detection on the features of the image to be classified to obtain a category detection result; wherein the category detection result represents: the probability that the object in the image to be classified belongs to each preset category.

6. An image classification model training device, characterized in that: The device comprises: An image acquisition module is used to acquire a sample image and divide the sample image according to a preset division method to obtain each sample image area; A sequence acquisition module is used to acquire a region sequence composed of each sample image region obtained by division; A feature extraction module, used to extract features from the acquired region sequence using a conversion encoder in an image classification model of an initial structure, to obtain sample image features; wherein the image classification model further includes a category detection network and a position detection network; A category detection module, used to perform category detection on the sample image features using the category detection network to obtain a category detection result; wherein the category detection result represents: the probability that the sample object in the sample image belongs to each preset category; A position detection module, used to perform position detection on the sample image features using the position detection network to obtain a position detection result; wherein the position detection result represents: a predicted probability that each sample image region is located at each preset position in the sample image; each preset position is determined by dividing the sample image according to the preset division method; A training module is used to adjust the model parameters of the image classification model of the initial structure based on the difference between the category label and the category detection result, and the difference between the position label and the position detection result, until the preset convergence condition is reached to obtain a trained image classification model; wherein the category label indicates the preset category to which the sample object belongs; and the position label indicates the preset position of each sample image area in the sample image.

7. An image classification device, characterized in that: The device comprises: A first acquisition module is used to acquire an image to be classified, and divide the image to be classified according to a preset division method to obtain image regions to be classified; The second acquisition module is used to acquire a region sequence composed of the divided image regions to be classified; An extraction module, used to extract features from the acquired region sequence using a conversion encoder in a trained image classification model to obtain features of the image to be classified; wherein the image classification model is obtained based on the method described in any one of claims 1 to 4 above; the image classification model also includes a category detection network; The detection module uses the category detection network to perform category detection on the features of the image to be classified to obtain a category detection result; wherein the category detection result represents: the probability that the object in the image to be classified belongs to each preset category.

8. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when used to execute a program stored in a memory, implements the method described in any one of claims 1 to 4, or the method described in claim 5.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 4 or the method described in claim 5 is implemented.

10. A computer program product, characterized in that When the computer program product is run on a computer, the computer is enabled to execute the method according to any one of claims 1 to 4, or the method according to claim 5.