Image Classification Model Training Classification Method, Device, Roadside Equipment and Cloud Control Platform

By adding a second classification layer to the basic classification model, building a neural network model and training, the problem of difficulty in completing fine-grained image classification in the existing technology is solved, and efficient image classification performance is achieved.

CN114282583BActive Publication Date: 2025-06-20SHANDONG LONGDU INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110473572.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-29
Publication Date
2025-06-20
Estimated Expiration
2041-04-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the task of fine-grained image classification, especially because the basic classification model cannot complete complex fine-grained classification, and the problem of high complexity and difficulty in collecting additional information is present with the help of detection or segmentation methods.

Method used

Add a second classification layer to the existing basic classification model, and build a neural network model including feature extraction layer, pooling layer, first fully connected layer, first classification layer and second classification layer. Use multiple images and their labels to train the model to realize the classification of the parent and subcategories of the target in the image.

Benefits of technology

The training steps of the image classification model are simplified, the classification performance of the image classification model is improved, and the fine-grained image classification task can be effectively completed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114282583B_ABST
    Figure CN114282583B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for training an image classification model and image classification, which relates to the fields of artificial intelligence technologies such as intelligent transportation, computer vision, and deep learning. Among them, the method for training the image classification model includes: obtaining training data; constructing a neural network model including a feature extraction layer, a pooling layer, a first fully connected layer, a first classification layer, and a second classification layer, where the first classification layer is used to obtain a parent category according to the output result of the first fully connected layer, and the second classification layer is used to obtain a sub-category according to the output result of the first fully connected layer; training the neural network model using multiple first images and the first labels and second labels of the multiple first images to obtain an image classification model. The method for image classification includes: obtaining an image to be processed; using the image to be processed as the input of the image classification model, and using the parent category and sub-category output by the image classification model as the classification result of the image to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to the fields of intelligent transportation, computer vision, and deep learning technology. A method, device, roadside device, and cloud control platform for training an image classification model and image classification are provided. Background Art

[0002] As one of the basic tasks of computer vision, image classification has been widely studied and has achieved exciting research results. Since the basic classification models in the prior art only study the classification tasks of general categories and cannot complete the fine-grained classification tasks with high requirements, some fine-grained classification methods have emerged. For example, strategies such as detection or segmentation are used to obtain more refined classification features, or additional auxiliary information is used to enhance feature learning. However, the above two methods have the problems of relatively high complexity and difficulty in collecting additional information. Summary of the Invention

[0003] According to a first aspect of the present disclosure, a method for training an image classification model is provided, including: obtaining training data, where the training data includes multiple first images and first labels and second labels of the multiple first images; constructing a neural network model including a feature extraction layer, a pooling layer, a first fully connected layer, a first classification layer, and a second classification layer, where the first classification layer is used to obtain a parent category according to the output result of the first fully connected layer, and the second classification layer is used to obtain a subcategory according to the output result of the first fully connected layer; using the multiple first images and the first labels and second labels of the multiple first images to train the neural network model to obtain an image classification model.

[0004] According to a second aspect of the present disclosure, a method for image classification is provided, including: obtaining an image to be processed; using the image to be processed as an input to an image classification model, and using the parent category and subcategory output by the image classification model as the classification result of the image to be processed.

[0005] According to a third aspect of the present disclosure, a device for training an image classification model is provided, including: a first obtaining unit for obtaining training data, where the training data includes multiple first images and first labels and second labels of the multiple first images; a constructing unit for constructing a neural network model including a feature extraction layer, a pooling layer, a first fully connected layer, a first classification layer, and a second classification layer, where the first classification layer is used to obtain a parent category according to the output result of the first fully connected layer, and the second classification layer is used to obtain a subcategory according to the output result of the first fully connected layer; a training unit for using the multiple first images and the first labels and second labels of the multiple first images to train the neural network model to obtain an image classification model.

[0006] According to a fourth aspect of the present disclosure, there is provided an apparatus for image classification, including: a second acquisition unit configured to acquire an image to be processed; and a classification unit configured to use the image to be processed as an input to an image classification model, and use the parent category and the sub-category output by the image classification model as the classification result of the image to be processed.

[0007] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method as described above.

[0008] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.

[0009] According to a seventh aspect of the present disclosure, there is provided a computer program product including a computer program, which implements the method as described above when executed by a processor.

[0010] It can be seen from the above technical solutions that by constructing a neural network model by adding a second classification layer to an existing basic classification model, the constructed neural network model can output the parent category and the sub-category of the target in the first image according to the input first image, and since the added second classification layer uses the same input as the original first classification layer, this embodiment does not require major modifications to the architecture of the basic classification model, thereby simplifying the training steps of the image classification model and improving the classification performance of the image classification model.

[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0013] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;

[0014] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;

[0015] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;

[0016] Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure;

[0017] Figure 5 is a schematic diagram according to the fifth embodiment of the present disclosure;

[0018] Figure 6 is a schematic diagram according to the sixth embodiment of the present disclosure;

[0019] Figure 7 is a block diagram of an electronic device for implementing the method of training an image classification model and image classification in an embodiment of the present disclosure. Detailed implementation manners

[0020] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and mechanisms are omitted below for clarity and conciseness.

[0021] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure. As Figure 1 shown, the method for training an image classification model in this embodiment may specifically include the following steps:

[0022] S101. Obtain training data, where the training data includes multiple first images and first labels and second labels of the multiple first images;

[0023] S102. Construct a neural network model including a feature extraction layer, a pooling layer, a first fully connected layer, a first classification layer, and a second classification layer. The first classification layer is used to obtain a parent category according to the output result of the first fully connected layer, and the second classification layer is used to obtain a subcategory according to the output result of the first fully connected layer;

[0024] S103. Use the multiple first images and the first labels and second labels of the multiple first images to train the neural network model to obtain an image classification model.

[0025] The method for training an image classification model in this embodiment constructs a neural network model by adding a second classification layer to an existing basic classification model, so that the constructed neural network model can output the parent category and subcategory of the target in the first image according to the input first image. And since the added second classification layer uses the same input as the original first classification layer, this embodiment does not need to make major changes to the architecture of the basic classification model, thereby simplifying the training steps of the image classification model and improving the classification performance of the image classification model.

[0026] In the training data obtained by this embodiment in S101, the first label of the first image represents the parent category to which the target in the first image belongs, and the second label represents the subcategory to which the target in the first image belongs. Among them, there is a corresponding relationship between the parent category and the subcategory of the target, and there is one or more subcategories under each parent category.

[0027] For example, if parent category 1 is "motor vehicle" and parent category 2 is "non-motor vehicle", then the subcategories corresponding to parent category 1 include subcategory 1 "truck", subcategory 2 "car", subcategory 3 "bus", etc., and the subcategories corresponding to parent category 2 include subcategory 1 "bicycle", subcategory 2 "tricycle", subcategory 3 "person", etc.

[0028] After this embodiment obtains the training data including multiple first images and the first labels and second labels of multiple first images in S101, it executes S102 to construct a neural network model including a feature extraction layer, a pooling layer, a first fully connected layer, a first classification layer, and a second classification layer, where the first classification layer is the original classification layer in the basic classification model.

[0029] In the neural network model constructed by this embodiment in S102, the feature extraction layer is used to extract the features of the input image, and the feature extraction layer can be a backbone network such as ResNet, MobileNet, ResNext, etc.; the pooling layer is used to perform pooling processing on the output result of the feature extraction layer, for example, average pooling processing; the first fully connected layer is used to convert the output result of the pooling layer into a vector for the classification layer to perform classification according to the obtained vector.

[0030] In the neural network model constructed by this embodiment in S102, the first classification layer is used to obtain the parent category according to the output result of the first fully connected layer, and the second classification layer is used to obtain the subcategory according to the output result of the first fully connected layer. That is to say, the first classification layer and the second classification layer in this embodiment respectively obtain the parent category and the subcategory of the target in the image according to the same input.

[0031] The first classification layer in this embodiment is the original classification layer in the basic classification model. It obtains the probability that the target in the image belongs to each parent category according to the output result of the first fully connected layer, and then takes the parent category with the highest probability as the parent category of the target; the first classification layer includes a first classifier, and the first classifier can be a classifier such as softmax.

[0032] Specifically, when this embodiment executes S102 and the first classification layer obtains the parent category according to the output result of the first fully-connected layer, an optional implementation method that can be adopted is: input the output result of the first fully-connected layer into the first classifier, and obtain the parent category according to the output result of the first classifier.

[0033] The second classification layer in this embodiment is a classification layer additionally added to the basic classification model. It obtains the probability that the target in the image belongs to each subcategory corresponding to the parent category of the target according to the output result of the first fully-connected layer, and then takes the subcategory with the highest probability as the subcategory of the target; the second classification layer includes a second fully-connected layer, a weight coefficient generation layer, and a second classifier. The weight coefficient generation layer can be a sigmoid function for generating weight coefficients between (0, 1), and the second classifier can be a softmax classifier.

[0034] Specifically, when this embodiment executes S102 and the second classification layer obtains the subcategory according to the output result of the first fully-connected layer, an optional implementation method that can be adopted is: input the output result of the first fully-connected layer into the second fully-connected layer to obtain the output result of the second fully-connected layer; input the output result of the second fully-connected layer into the weight coefficient generation layer to obtain the output result of the weight coefficient generation layer; input the output result of the first fully-connected layer and the output result of the weight coefficient generation layer into the second classifier, and obtain the subcategory according to the output result of the second classifier.

[0035] Among them, when this embodiment executes S102 and inputs the output result of the first fully-connected layer and the output result of the weight coefficient generation layer into the second classifier, after performing a dot product operation on the two output results, the obtained processing result can be used as a prediction feature to be input into the second classifier for subcategory classification.

[0036] The output result of the first fully-connected layer in the neural network model is a feature for parent category classification. The second classification layer in this embodiment does not depend on the classification result of the first classification layer. Through the processing of the second fully-connected layer and the weight coefficient generation layer, appropriate features can be selected from the features for parent category classification as features for subcategory classification, and then the subcategory of the target can be obtained according to the selected features.

[0037] That is to say, since the second classification layer in this embodiment is independent of the first classification layer, this embodiment does not need to make major changes to the architecture of the existing basic classification model, and the second classification layer can be embedded into the existing basic classification model, so that the constructed neural network model has the ability of coarse classification and fine classification, thereby improving the classification performance of the basic classification model.

[0038] After constructing a neural network model including a feature extraction layer, a pooling layer, a first fully connected layer, a first classification layer, and a second classification layer in S102, this embodiment executes S103 to train the neural network model using multiple first images and the first labels and second labels of the multiple first images to obtain an image classification model.

[0039] The training process of the neural network model in S103 of this embodiment is specifically a process of continuously adjusting the parameters in the neural network model according to the loss function obtained from the output result of the neural network model for each first image and the label of each first image.

[0040] Specifically, when this embodiment executes S103 to train the neural network model using multiple first images and the first labels and second labels of the multiple first images to obtain an image classification model, an optional implementation method that can be adopted is: input the multiple first images into the neural network model respectively to obtain the parent category and subcategory output by the neural network model for each first image; calculate the loss function according to the parent category and the first label, and the subcategory and the second label of each first image, and use the calculated loss function to adjust the parameters in the neural network model until the neural network model converges to obtain an image classification model.

[0041] The image classification model obtained by this embodiment executing S103 can output the parent category and subcategory of the target in the input image, and the output subcategory is one of the multiple subcategories corresponding to the output parent category.

[0042] According to the above method, this embodiment constructs a neural network model by adding a second classification layer to an existing basic classification model, so that the neural network model can output the parent category and subcategory of the target in the input first image. And because the added second classification layer uses the same input as the original first classification layer, this embodiment does not need to make major changes to the architecture of the basic classification model, thereby simplifying the training steps of the image classification model and improving the classification performance of the image classification model.

[0043] Figure 2 It is a schematic diagram according to the second embodiment of the present disclosure. As Figure 2 shown, when this embodiment executes S103 "train the neural network model using multiple first images and the first labels and second labels of the multiple first images to obtain an image classification model", the following steps may be included:

[0044] S201. Select and save the neural network model that meets the preset conditions during the training process;

[0045] S202. Select one of the saved neural network models as the image classification model.

[0046] It can be understood that the training process of the neural network model is a process of continuously adjusting the parameters in the neural network model. After training the neural network model once with the training data, a new neural network model can be obtained.

[0047] Therefore, in this embodiment, multiple neural network models are selected and saved along with the training process of the neural network model. Since each of the saved neural network models corresponds to a different training stage, this embodiment can select one from multiple neural network models in different training stages as the final image classification model.

[0048] When this embodiment executes S201 to select and save the neural network models that meet the preset conditions during the training process, it can select the neural network models whose training times reach the preset number of times during the training process. For example, this embodiment can select the neural network models when the training times reach 50,000 times, 100,000 times, 150,000 times, 200,000 times, etc. for saving.

[0049] In addition, when this embodiment executes S201 to select and save the neural network models that meet the preset conditions during the training process, it can also select the neural network models whose classification accuracy and / or classification speed reach the preset values for saving. For example, this embodiment can select the neural network models with a classification accuracy of more than 85% for saving.

[0050] When this embodiment executes S202 to select one of the saved neural network models as the image classification model, an optional implementation method that can be adopted is: obtain test data, where the obtained test data contains multiple second images; input the obtained multiple second images into the saved neural network models respectively to obtain the parent category and sub-category output by each neural network model for each second image; according to the parent category and sub-category of each second image, obtain the classification performance of each neural network model, and the obtained classification performance is classification accuracy and / or classification speed; select the neural network model with the highest classification performance as the image classification model.

[0051] That is to say, this embodiment tests the saved neural network models by obtaining test data, so as to select the neural network model with the highest classification accuracy and / or classification speed in the test results as the image classification model, further improving the classification performance of the image classification model.

[0052] Figure 3 It is a schematic diagram according to the third embodiment of the present disclosure. As Figure 3 shown in Figure 3On the left is the structure diagram of the existing basic classification model, Figure 3 and on the right is the structure diagram of the image classification model of the present disclosure; the basic classification model includes a feature extraction layer, a pooling layer, a fully connected layer, and a classification layer, and the image classification model includes a feature extraction layer, a pooling layer, a first fully connected layer, a first classification layer, and a second classification layer. The first classification layer includes a first classifier, and the second classification layer includes a second fully connected layer, a weight coefficient generation layer, and a second classifier.

[0053] Through Figure 3 the structure diagrams of the two neural networks in, it can be seen that the image classification model obtained in this embodiment does not change the main architecture of the basic classification model. On the basis of being able to obtain the parent category and subcategory of the target in the image, the construction steps of the image classification model are simplified, and the training efficiency of the image classification model is improved.

[0054] Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure. As Figure 4 shown, the method for image classification in this embodiment may specifically include the following steps:

[0055] S401. Obtain an image to be processed;

[0056] S402. Use the image to be processed as the input of the image classification model, and use the parent category and subcategory output by the image classification model as the classification result of the image to be processed.

[0057] For the method for image classification in this embodiment, the above-mentioned embodiment is used to pre-train the obtained image classification model to obtain the classification result. Since the image classification model can obtain the parent category and subcategory of the target in the image, the accuracy and richness of the obtained classification result are improved.

[0058] The image to be processed obtained by this embodiment in S401 may be an existing image or a real-time captured image.

[0059] Figure 5 is a schematic diagram according to the fifth embodiment of the present disclosure. As Figure 5 shown, the training device 500 of the image classification model in this embodiment includes:

[0060] A first acquisition unit 501, configured to acquire training data, where the training data includes multiple first images, as well as first labels and second labels of the multiple first images;

[0061] The building unit 502 is used to build a neural network model including a feature extraction layer, a pooling layer, a first fully connected layer, a first classification layer, and a second classification layer. The first classification layer is used to obtain the parent category according to the output result of the first fully connected layer, and the second classification layer is used to obtain the subcategory according to the output result of the first fully connected layer;

[0062] The training unit 503 is used to train the neural network model with multiple first images and the first labels and second labels of the multiple first images to obtain an image classification model.

[0063] In the training data obtained by the first acquisition unit 501, the first label of the first image represents the parent category to which the target in the first image belongs, and the second label represents the subcategory to which the target in the first image belongs. Among them, there is a corresponding relationship between the parent category of the target and the subcategory of the target, and there is one or more subcategories under each parent category.

[0064] In this embodiment, after the training data including multiple first images and the first labels and second labels of the multiple first images is obtained by the first acquisition unit 501, the building unit 502 builds a neural network model including a feature extraction layer, a pooling layer, a first fully connected layer, a first classification layer, and a second classification layer. The first classification layer among them is the original classification layer in the basic classification model.

[0065] In the neural network model built by the building unit 502, the feature extraction layer is used to extract the features of the input image; the pooling layer is used to perform pooling processing on the output result of the feature extraction layer; the first fully connected layer is used to convert the output result of the pooling layer into a vector for the classification layer to perform classification according to the obtained vector.

[0066] In the neural network model built by the building unit 502, the first classification layer is used to obtain the parent category according to the output result of the first fully connected layer, and the second classification layer is used to obtain the subcategory according to the output result of the first fully connected layer.

[0067] The first classification layer in this embodiment is the original classification layer in the basic classification model. It obtains the probability that the target in the image belongs to each parent category according to the output result of the first fully connected layer, and then takes the parent category with the largest probability as the parent category of the target; the first classification layer includes a first classifier, and this first classifier can be a classifier such as softmax.

[0068] Specifically, when the first classification layer built by the building unit 502 obtains the parent category according to the output result of the first fully connected layer, an optional implementation method that can be adopted is: input the output result of the first fully connected layer into the first classifier, and obtain the parent category according to the output result of the first classifier.

[0069] The second classification layer in this embodiment is a classification layer additionally added to the basic classification model. According to the output result of the first fully connected layer, it obtains the probabilities that the target in the image belongs to each subcategory corresponding to the parent category of the target, and then takes the subcategory with the highest probability as the subcategory of the target; the second classification layer includes a second fully connected layer, a weight coefficient generation layer, and a second classifier. The weight coefficient generation layer can be a sigmoid function for generating weight coefficients between (0, 1), and the second classifier can be a softmax classifier.

[0070] Specifically, for the second classification layer constructed by the construction unit 502, when obtaining the subcategory according to the output result of the first fully connected layer, an optional implementation method can be: input the output result of the first fully connected layer into the second fully connected layer to obtain the output result of the second fully connected layer; input the output result of the second fully connected layer into the weight coefficient generation layer to obtain the output result of the weight coefficient generation layer; input the output result of the first fully connected layer and the output result of the weight coefficient generation layer into the second classifier, and obtain the subcategory according to the output result of the second classifier.

[0071] After the construction unit 502 in this embodiment constructs a neural network model including a feature extraction layer, a pooling layer, a first fully connected layer, a first classification layer, and a second classification layer, the training unit 503 uses multiple first images and the first labels and second labels of the multiple first images to train the neural network model to obtain an image classification model.

[0072] The training process of the training unit 503 for the neural network model is specifically a process of continuously adjusting the parameters in the neural network model according to the loss function obtained from the output result of the neural network model for each first image and the label of each first image.

[0073] Specifically, when the training unit 503 uses multiple first images and the first labels and second labels of the multiple first images to train the neural network model to obtain an image classification model, an optional implementation method can be: input the multiple first images into the neural network model respectively to obtain the parent category and subcategory output by the neural network model for each first image; calculate the loss function according to the parent category and the first label, and the subcategory and the second label of each first image, and use the calculated loss function to adjust the parameters in the neural network model until the neural network model converges to obtain an image classification model.

[0074] In addition, when the training unit 503 uses multiple first images and the first labels and second labels of the multiple first images to train the neural network model to obtain an image classification model, an optional implementation method can be: select the neural network models that meet the preset conditions during the training process for preservation; select one from the saved neural network models as the image classification model.

[0075] When the training unit 503 selects a neural network model that meets a preset condition during training for saving, it can select a neural network model whose number of training times reaches a preset number during training for saving.

[0076] In addition, when the training unit 503 selects a neural network model that meets a preset condition during training for saving, it can also select a neural network model whose classification accuracy and / or classification speed reaches a preset value for saving.

[0077] When the training unit 503 selects one of the saved neural network models as an image classification model, an optional implementation method that can be adopted is: obtaining test data, where the obtained test data includes multiple second images; inputting the obtained multiple second images into the saved neural network models respectively to obtain the parent category and subcategory output by each neural network model for each second image; obtaining the classification accuracy and / or classification speed of each neural network model according to the parent category and subcategory of each second image; selecting the neural network model with the highest classification accuracy and / or classification speed as the image classification model.

[0078] That is to say, the training unit 503 tests the saved neural network models by obtaining test data, so as to select the neural network model with the highest classification accuracy and / or classification speed in the test results as the image classification model, further improving the classification performance of the image classification model.

[0079] The image classification model obtained by the training unit 503 can output the parent category and subcategory of the target in the input image, and the output subcategory is one of the multiple subcategories corresponding to the output parent category.

[0080] Figure 6 It is a schematic diagram according to the sixth embodiment of the present disclosure. As Figure 6 shown, the image classification device 600 of this embodiment includes:

[0081] A second acquisition unit 601, configured to acquire an image to be processed;

[0082] A classification unit 602, configured to use the image to be processed as the input of the image classification model, and use the parent category and subcategory output by the image classification model as the classification result of the image to be processed.

[0083] The image to be processed acquired by the second acquisition unit 601 can be an existing image or a real-time captured image.

[0084] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0085] As Figure 7 shown, it is a block diagram of an electronic device for a method of training an image classification model and image classification according to an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0086] As Figure 7 shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0087] A plurality of components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0088] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the method for training an image classification model and image classification. For example, in some embodiments, the method for training an image classification model and image classification can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708.

[0089] In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the method for training an image classification model and image classification described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the method for training an image classification model and image classification by any other suitable means (e.g., by means of firmware).

[0090] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0091] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0092] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0093] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0094] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0095] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability in traditional physical hosts and VPS services (“Virtual Private Server”, or simply “VPS”). The server can also be a server of a distributed system, or a server combined with blockchain.

[0096] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0097] According to another aspect of the present disclosure, a roadside device is provided, including the electronic device 700 described above. Optionally, in addition to including the electronic device 700, the roadside device can further include communication components, etc. The electronic device 700 can be integrated with the communication components as a whole or can be separately arranged. The electronic device 700 can obtain data from sensing devices (such as roadside cameras), such as pictures and videos, etc., and thus perform image and video processing and data calculation. Optionally, the electronic device 700 itself can also have the functions of obtaining sensing data and communication functions. For example, it is an AI camera, and the electronic device 700 can directly perform image and video processing and data calculation based on the obtained sensing data.

[0098] According to another aspect of the present disclosure, a transportation control platform is provided, which includes the electronic device 700 as described above. Optionally, the cloud control platform performs processing in the cloud. The electronic device 700 included in the cloud control platform can obtain data from sensing devices (such as roadside cameras), such as pictures and videos, etc., so as to perform image and video processing and data calculation; the cloud control platform can also be referred to as a vehicle-road collaborative management platform, an edge computing platform, a cloud computing platform, a central system, a cloud server, etc.

[0099] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for training an image classification model, comprising: Obtain training data, where the training data includes multiple first images and first labels and second labels of the multiple first images; Construct a neural network model including a feature extraction layer, a pooling layer, a first fully connected layer, a first classification layer, and a second classification layer. The first classification layer is used to obtain a parent category according to the output result of the first fully connected layer, and the second classification layer is used to obtain a sub-category according to the output result of the first fully connected layer; Use the multiple first images and the first labels and second labels of the multiple first images to train the neural network model to obtain an image classification model; Wherein, the second classification layer includes a second fully connected layer, a weight coefficient generation layer, and a second classifier; The second classification layer obtaining a sub-category according to the output result of the first fully connected layer includes: Input the output result of the first fully connected layer into the second fully connected layer to obtain the output result of the second fully connected layer; Input the output result of the second fully connected layer into the weight coefficient generation layer to obtain the output result of the weight coefficient generation layer; Perform a dot product operation on the output result of the first fully connected layer and the output result of the weight coefficient generation layer, and use the obtained processing result as a prediction feature to input into the second classifier, and obtain a sub-category according to the output result of the second classifier.

2. The method according to claim 1, wherein, The using the multiple first images and the first labels and second labels of the multiple first images to train the neural network model to obtain an image classification model includes: Select a neural network model that meets a preset condition during the training process for saving; Select one from the saved neural network models as the image classification model.

3. The method according to claim 2, wherein, The selecting one from the saved neural network models as the image classification model includes: Obtain test data, where the test data includes multiple second images; Input the multiple second images into the saved neural network models respectively to obtain the parent category and sub-category output by each neural network model for each second image; Obtain the classification performance of each neural network model according to the parent category and sub-category of each second image; Select the neural network model with the highest classification performance as the image classification model.

4. A method for image classification, comprising: Obtain an image to be processed; Use the image to be processed as the input of the image classification model, and use the parent category and sub-category output by the image classification model as the classification result of the image to be processed; Wherein, the image classification model is pre-trained according to the method of any one of claims 1-3.

5. An apparatus for training an image classification model, comprising: A first obtaining unit, configured to obtain training data, where the training data includes multiple first images and first labels and second labels of the multiple first images; A constructing unit, configured to construct a neural network model including a feature extraction layer, a pooling layer, a first fully connected layer, a first classification layer, and a second classification layer. The first classification layer is used to obtain a parent category according to the output result of the first fully connected layer, and the second classification layer is used to obtain a sub-category according to the output result of the first fully connected layer; A training unit for training the neural network model using multiple first images and first and second labels of the multiple first images to obtain an image classification model; Wherein, the second classification layer constructed by the construction unit includes a second fully-connected layer, a weight coefficient generation layer, and a second classifier; When obtaining sub-categories according to the output result of the first fully-connected layer, the second classification layer specifically performs: Inputting the output result of the first fully-connected layer into the second fully-connected layer to obtain the output result of the second fully-connected layer; Inputting the output result of the second fully-connected layer into the weight coefficient generation layer to obtain the output result of the weight coefficient generation layer; Performing a dot product operation on the output result of the first fully-connected layer and the output result of the weight coefficient generation layer, and using the obtained processing result as a prediction feature to input into the second classifier, and obtaining sub-categories according to the output result of the second classifier.

6. The apparatus according to claim 5, wherein, When the training unit trains the neural network model using multiple first images and first and second labels of the multiple first images to obtain an image classification model, it specifically performs: Selecting a neural network model that meets a preset condition during training for storage; Selecting one from the stored neural network models as the image classification model.

7. The apparatus according to claim 6, wherein, When the training unit selects one from the stored neural network models as the image classification model, it specifically performs: Obtaining test data, where the test data includes multiple second images; Inputting the multiple second images into the stored neural network models respectively to obtain the parent categories and sub-categories output by each neural network model for each second image; Obtaining the classification performance of each neural network model according to the parent categories and sub-categories of each second image; Selecting the neural network model with the highest classification performance as the image classification model.

8. An apparatus for image classification, comprising: A second acquisition unit for acquiring an image to be processed; A classification unit for using the image to be processed as an input to the image classification model, and using the parent category and sub-category output by the image classification model as the classification result of the image to be processed; Wherein, the image classification model is pre-trained according to any one of the devices in claims 5-7.

9. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-4.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.

11. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-4.

12. A roadside device, comprising the electronic device according to claim 9.

13. A cloud control platform, comprising the electronic device according to claim 9.

Citation Information

Patent Citations

  • Image classification method, neural network training method and neural network training device

    CN110309856A

  • Fine-grained Image Classification by Exploring Bipartite-Graph Labels

    US20160307072A1