Image recognition method, device and system and computer readable storage medium
By supervising the intermediate feature layer during the model distillation process and generating target images to improve recognition accuracy, the problem that the existing lightweight image recognition model cannot effectively supervise the intermediate feature layer during the model distillation process is solved, and high accuracy recognition is achieved under a small calculation amount.
Patent Information
- Application Number
- CN202311661815.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-06
AI Technical Summary
The existing lightweight image recognition model cannot effectively supervise the intermediate feature layer during the model distillation process, resulting in a decrease in recognition accuracy.
By supervising the model distillation of the intermediate feature layer, a target image is generated to improve the recognition accuracy of the recognition model.
While ensuring a small amount of calculation, the accuracy of image recognition is effectively improved.
Smart Images

Figure CN120107643A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an image recognition method, device, system and computer-readable storage medium. Background Art
[0002] Image recognition is one of the most important contents in computer vision. Image recognition is to perform semantic segmentation on the original image, assign each pixel in the image to a predefined category label, identify the content in the original image and label it. It is widely used in industrial manufacturing, medical image processing, autonomous driving, smart cities and other fields.
[0003] Current image recognition generally runs on platforms with very limited computing power, memory, and power consumption. This platform has very high requirements for the lightweight nature of the image recognition model. As a result, lightweight image recognition models have emerged. However, existing image recognition segmentation models are obtained by distilling complex image recognition models. During model distillation, the intermediate feature layers of the lightweight image recognition model cannot be effectively supervised, which simplifies the parameters or compresses the channels of the feature map, resulting in weak extraction capabilities of the lightweight image recognition model and reducing the recognition accuracy of the lightweight image recognition model.
[0004] Therefore, how to improve the accuracy of image recognition while ensuring a small amount of calculation is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The embodiments of the present application provide an image recognition method, device, system and computer-readable storage medium, which can improve the accuracy of image recognition while ensuring a small amount of calculation.
[0006] In a first aspect, an embodiment of the present application provides an image recognition method, the method comprising:
[0007] Obtain an image to be recognized;
[0008] Processing the image to be recognized by using a recognition model to obtain output data;
[0009] Based on the output data and the category data, a target image corresponding to the image to be recognized is generated.
[0010] In a second aspect, an embodiment of the present application provides an image recognition device comprising:
[0011] An acquisition unit, used for acquiring an image to be recognized;
[0012] A processing unit, used for processing the image to be recognized by using a recognition model to obtain output data;
[0013] A generating unit is used to generate a target image corresponding to the image to be identified based on the output data and the category data.
[0014] In a third aspect, an embodiment of the present application also provides an image recognition system, including a memory storing multiple instructions; a processor loads instructions from the memory to execute the steps of any image recognition method provided in the embodiment of the present application.
[0015] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores a plurality of instructions suitable for loading by a processor to execute the steps of any one of the image recognition methods provided in the embodiments of the present application.
[0016] In a fifth aspect, an embodiment of the present application further provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the steps of any image recognition method provided in the embodiment of the present application.
[0017] The scheme of the embodiment of the application is adopted to obtain an image to be recognized; the image to be recognized is processed by a recognition model to obtain output data; based on the output data and the category data, a target image corresponding to the image to be recognized is generated. The model distillation of the supervised intermediate feature layer effectively reduces the computational complexity of the recognition model and improves the recognition accuracy of the recognition model, and the image to be recognized is processed by the recognition model to obtain output data, and based on the output data and the category data, a target image corresponding to the image to be recognized is generated, ensuring that the accuracy of image recognition is improved under the premise of a small amount of computation. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 It is a schematic diagram of the first embodiment of the flow chart of the image recognition method provided in the embodiments of the present application;
[0020] Figure 2 is a schematic flow chart of a second embodiment of the image recognition method provided in the embodiments of the present application;
[0021] Figure 3 is a schematic diagram of selecting a distillation node provided in an embodiment of the present application;
[0022] Figure 4 is a schematic diagram of the structure of an image recognition device provided in an embodiment of the present application;
[0023] Figure 5 It is a structural diagram of the image recognition system provided in the embodiment of the present application. DETAILED DESCRIPTION
[0024] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application. At the same time, in the description of the embodiments of the present application, the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.
[0025] Embodiments of the present application provide an image recognition method, device, system, and computer-readable storage medium.
[0026] Specifically, the embodiments of the present application will be described from the perspective of an image recognition system, which can be integrated into an image recognition device, that is, the image recognition method of the embodiments of the present application can be executed by the image recognition system.
[0027] The image recognition method provided in the embodiment of the present application can be applied to an image recognition device, which may include devices such as a mobile terminal, a PC terminal, etc.
[0028] The following is a detailed description in conjunction with the accompanying drawings. In the embodiments of the present application, the execution subject is an image recognition system as an example. It should be noted that the description order of the following embodiments is not intended to limit the preferred order of the embodiments. Although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in an order different from that shown in the accompanying drawings.
[0029] Please refer to Figure 1 The specific process of the first embodiment of the image recognition method includes the following steps:
[0030] Step 101, obtaining an image to be recognized;
[0031] Step 102, processing the image to be recognized by using a recognition model to obtain output data;
[0032] Step 103: Generate a target image corresponding to the image to be recognized based on the output data and the category data.
[0033] In this embodiment, when it is necessary to identify an image and identify various objects contained in the image, the image recognition system obtains the image to be identified input by the relevant personnel, and performs feature extraction, feature fusion, upsampling, feature analysis and other processing on the image to be identified through the recognition model obtained by the model distillation of the supervised intermediate feature layer in advance, and obtains output data. Based on the output data and the category data, the target image corresponding to the image to be identified is generated. The image recognition system generates the target image corresponding to the image to be identified based on the output data and the category data of the recognition model. It should be noted that the recognition model learns the features of certain objects and can identify whether these objects exist in the image to be identified. The category data is the category annotation information of the objects that the recognition model can identify. Different objects correspond to different annotation information. The image recognition system determines the object category corresponding to each pixel in the image to be identified based on the output data, and determines the category annotation information corresponding to each pixel according to the object category and the category data, and then annotates the pixels according to the category annotation information to generate the target image corresponding to the image to be identified.
[0034] The image recognition system of this embodiment obtains an image to be recognized; processes the image to be recognized through a recognition model to obtain output data; and generates a target image corresponding to the image to be recognized based on the output data and category data. The model distillation of the supervised intermediate feature layer effectively reduces the computational complexity of the recognition model and improves the recognition accuracy of the recognition model. The image to be recognized is processed through the recognition model to obtain output data, and the target image corresponding to the image to be recognized is generated based on the output data and category data, ensuring that the accuracy of image recognition is improved under the premise of a small amount of computation.
[0035] Specifically, each step is described in detail below:
[0036] Step 101, obtaining an image to be recognized;
[0037] In this step, when it is necessary to identify the image and various objects contained in the image, the image recognition system obtains the image to be identified input by the relevant personnel.
[0038] Step 102, processing the image to be recognized by using a recognition model to obtain output data;
[0039] In this step, the image recognition system performs feature extraction, feature fusion, upsampling, feature analysis and other processing on the image to be recognized through the recognition model obtained by the model distillation of the supervised intermediate feature layer in advance to obtain output data. It should be noted that before the image to be recognized is input into the recognition model, the image recognition system reads the image to be recognized through pillow or opencv tools, converts the image to be recognized into an input tensor after normalization, and then inputs the input tensor into the recognition model.
[0040] Specifically, step 102 includes:
[0041] Step 1021, inputting the image to be identified into a recognition model, performing feature extraction on the image to be identified through the recognition model, and obtaining depth feature information and semantic feature information of the image to be identified;
[0042] In this step, the image recognition system converts the image to be recognized into an input tensor, and then inputs the input tensor into the recognition model, and performs feature extraction on the image to be recognized through the recognition model to obtain deep feature information and semantic feature information of the image to be recognized; specifically, the semantic feature information includes image detail features, image edge features and image global features. The image recognition system performs detail feature extraction on the input tensor through the recognition model to obtain image detail features of the image to be recognized, performs edge feature extraction on the input tensor through the recognition model to obtain image edge features of the image to be recognized, performs global feature extraction on the input tensor through the recognition model to obtain image global features of the image to be recognized, and performs deep feature extraction on the input tensor through the semantic segmentation network to obtain deep feature information of the image to be recognized.
[0043] It should be noted that the recognition model includes a sparse feature extraction network. The commonly used feature detection algorithms of the sparse feature extraction network include Harris corner detection, Shi-Tomasi corner detection, etc. Through these algorithms, the input tensor of the image to be recognized is detected to determine the pixels with drastic grayscale changes in the image to be recognized. These pixels usually correspond to the image detail features in the image to be recognized. Descriptors are extracted around these detected pixels to describe the local texture features around the pixels. These local texture features usually contain some tiny but very important detail information for the representation and matching of the image to be recognized, thereby obtaining representative image detail features in the image to be recognized.
[0044] The recognition model includes an edge feature extraction network, which performs edge detection on the input tensor of the image to be recognized. Common edge detection algorithms include Canny edge detection and Sobel operator to find possible edge regions. For the detected edge regions, it is necessary to calculate the descriptor of each edge region. Common edge descriptors include shape context descriptors and gradient histogram descriptors. These descriptors can capture image edge features such as edge shape, direction and intensity.
[0045] The recognition model includes a global feature extraction network, which uses a convolutional neural network or other feature extraction methods to extract features from the input tensor of the image to be recognized. Convolutional neural networks are usually used, such as VGGNet, ResNet, etc. These networks can convert the image to be recognized into a high-dimensional feature vector, match the extracted high-dimensional feature vectors, and aggregate the high-dimensional features of multiple images according to the matching results to obtain a global feature representation. This can be achieved by weighted summation, splicing, averaging, etc. of the feature vectors, and then the global feature extraction network obtains the image global features with global semantic information of the image to be recognized.
[0046] The recognition model includes a deep feature information extraction network, which uses a deep learning network or other depth estimation methods to perform depth estimation on the input tensor of the image to be recognized. The goal of depth estimation is to predict the depth value of each pixel of the image to be recognized, reflecting the distance of the object in the camera coordinate system. For the depth value of each pixel obtained by depth estimation, related features can be extracted. These features can include visual features such as color and texture, as well as geometric features such as normals and surface normal vectors related to depth. The extracted depth feature information is matched, and the depth feature information of multiple images is fused according to the matching results to obtain a global and comprehensive representation of the depth feature information. Fusion can be achieved through weighted summation, splicing, etc., and then the deep feature information extraction network obtains the image depth feature information of the image to be recognized.
[0047] Step 1022: Process the depth feature information and the semantic feature information through the recognition model to obtain output data.
[0048] In this step, after obtaining the depth feature information and semantic feature information corresponding to the image to be recognized, the image recognition system performs feature fusion on the depth feature information and the semantic feature information to obtain a feature map, and then performs upsampling and feature analysis on the feature map to obtain output data.
[0049] Further, step 1022 includes:
[0050] Step 10221: perform feature fusion on the depth feature information and the semantic feature information through the recognition model to obtain a down-sampled feature map corresponding to the image to be recognized.
[0051] In this step, the image recognition system performs feature fusion on the deep feature information and semantic feature information through the recognition model to obtain the downsampled feature map corresponding to the image to be recognized; specifically, the recognition model includes a feature fusion network, which fuses the image detail features, image edge features, and image global features in the deep feature information and semantic feature information to obtain a preliminary first-level fusion feature map. In order to allow the features of different dimensions to be further more fully integrated, a multiplication-addition nonlinear structure is introduced here. The image recognition system uses the recognition model to learn deeper fusion weights through the convolution network for the first-level fusion feature map after preliminary fusion, that is, after 2 layers of convolution layers, and after each convolution calculation layer, a BN layer and a ReLU layer are added to obtain a feature map with a certain fusion weight. This weighted fusion map is multiplied with the first-level fusion feature map through a multiplication layer (Mul) to obtain a second-level fusion feature map. The image recognition system superimposes the fused first-level fusion feature map and the second-level fusion feature map through the recognition model to obtain the final downsampled feature map. It should be understood that the superposition layer here is different from the fusion layer. The fusion layer concatenates the feature maps to increase the dimension of the feature maps, while the overlay layer can increase the amount of information in the feature maps. Although the dimension of the feature maps themselves has not increased, the amount of information in each dimension has increased. During the training process, more robust features can be obtained through learning.
[0052] Step 10222, upsampling the downsampled feature map through the recognition model to obtain an upsampled feature map with the same size as the image to be recognized;
[0053] In this step, the image recognition system upsamples the downsampled feature map through the recognition model to obtain an upsampled feature map of the same size as the image to be recognized. The feature fusion process will cause the loss of spatial resolution, so it is necessary to restore the spatial resolution through the upsampling (also called upsampling or deconvolution) process. The goal of this step is to generate an upsampled feature map of the same size as the image to be recognized.
[0054] Step 10223: Perform feature analysis on the upsampled feature map through the recognition model to obtain output data.
[0055] In this step, the image recognition system performs feature analysis on the upsampled feature map through the recognition model to obtain output data. Specifically, the recognition model obtains dimensional information such as the height, width, and number of channels of the upsampled feature map, obtains the feature vector corresponding to each pixel in the upsampled feature map, takes the feature vector corresponding to each pixel as input, and processes it through the softmax function. For each pixel, the softmax function generates a probability distribution, which represents the set of category prediction probabilities for each pixel, and then obtains the output data.
[0056] Step 103: Generate a target image corresponding to the image to be recognized based on the output data and the category data.
[0057] In this step, the output data of the recognition model includes a set of category prediction probabilities for each pixel in the image to be recognized. The image recognition system generates a target image corresponding to the image to be recognized based on the output tensor, the set of category prediction probabilities and the category data. It should be noted that the recognition model has learned the characteristics of certain objects and can identify whether these objects exist in the image to be recognized. The set of category prediction probabilities contains the probability values of the pixel belonging to each object that the recognition model can recognize. For example, if the recognition model can recognize three categories of objects: people, cars and obstacles, then the set of category prediction probabilities corresponding to each pixel contains the probability of people, cars and obstacles.
[0058] Specifically, step 103 includes:
[0059] Step 1031, obtaining a set of category prediction probabilities corresponding to each pixel point of the image to be identified in the output data;
[0060] Step 1032, determining the category to which each pixel belongs according to the category prediction probability set corresponding to each pixel;
[0061] Step 1033: label each pixel based on the category and category data to which each pixel belongs, and generate a target image corresponding to the image to be identified.
[0062] In step 1031 to step 1033, the image recognition system obtains the set of category prediction probabilities corresponding to each pixel of the image to be recognized in the output data, selects the largest category prediction probability from the set of category prediction probabilities corresponding to each pixel, and determines the category to which each pixel belongs according to the object category corresponding to the largest category prediction probability; the image recognition system annotates each pixel based on the category to which each pixel belongs and the category data, and generates a target image corresponding to the image to be recognized. Exemplarily, the recognition model can recognize three categories of objects: people, cars, and obstacles. Correspondingly, the category data includes category annotation information of people, cars, and obstacles, wherein the category annotation information of people is red, the category annotation information of cars is green, and the category annotation information of obstacles is blue. The image recognition system fills each pixel with the corresponding color according to the category to which each pixel belongs, and generates a target image corresponding to the image to be recognized, and the target image is annotated with category annotation information of different colors for the people, cars, and obstacles in the image to be recognized.
[0063] The image recognition system of this embodiment obtains an image to be recognized; processes the image to be recognized through a recognition model to obtain output data; and generates a target image corresponding to the image to be recognized based on the output data and category data. The model distillation of the supervised intermediate feature layer effectively reduces the computational complexity of the recognition model and improves the recognition accuracy of the recognition model. The image to be recognized is processed through the recognition model to obtain output data, and the target image corresponding to the image to be recognized is generated based on the output data and category data, ensuring that the accuracy of image recognition is improved under the premise of a small amount of computation.
[0064] For further information, please refer to Figure 2 , a second embodiment of the image recognition method is proposed. The difference between the second embodiment of the image recognition method and the first embodiment is that before obtaining the image to be recognized, a recognition model is obtained by supervising the model distillation of the intermediate feature layer. The specific process of obtaining the recognition model by supervising the model distillation of the intermediate feature layer includes the following steps:
[0065] Step a, obtaining a pre-trained recognition model and a first recognition model, and selecting a first distillation node in the pre-trained recognition model and selecting a second distillation node in the first recognition model;
[0066] Step b, setting a target loss function according to the first distillation node, the second distillation node, the pre-trained recognition model and the first recognition model;
[0067] Step c: acquiring training image data, and performing model distillation of the supervised intermediate feature layer of the pre-trained recognition model based on the training image data, the target loss function and the first recognition model to obtain a recognition model.
[0068] In this embodiment, when it is necessary to perform model distillation training of the intermediate feature layer of the selected pre-trained recognition model according to the selected first recognition model, the image recognition system selects the first distillation node in the pre-trained recognition model and the second distillation node in the first recognition model according to the model complexity or the overall structure of the model; the image recognition system determines the distillation node loss function, the model output loss function and the pre-trained recognition model loss function according to the first distillation node, the second distillation node, the pre-trained recognition model and the first recognition model, respectively, and then sets the target loss function according to the distillation node loss function, the model output loss function and the pre-trained recognition model loss function; the image recognition system obtains training image data, and performs model distillation of the intermediate feature layer of the pre-trained recognition model based on the training image data, the target loss function and the first recognition model to obtain a recognition model. It should be noted that the first recognition model and the pre-trained recognition model can be applied to related models in the field of image recognition. The first recognition model is generally a network model with a more complex structure or performs better in image recognition tasks. The pre-trained recognition model is generally a network model with a simpler structure than the first recognition model, and is suitable for scenarios and tasks with lower requirements for computing power and power consumption. The training image data is usually selected based on the actual application scenarios of the first recognition model and the pre-trained recognition model.
[0069] The image recognition system of this embodiment obtains a pre-trained recognition model and a first recognition model, and selects a first distillation node in the pre-trained recognition model and a second distillation node in the first recognition model; sets a target loss function according to the first distillation node, the second distillation node, the pre-trained recognition model and the first recognition model; obtains training image data, and performs image recognition on the pre-trained recognition model based on the training image data, the target loss function and the first recognition model to obtain a recognition model. By adding the loss function corresponding to the distillation node to the loss function while performing image recognition, and then supervising the intermediate feature layer of the pre-trained recognition model, the generalization ability of the pre-trained recognition model obtained after image recognition is effectively improved, the computational complexity of the recognition model is effectively reduced, and the recognition accuracy of the recognition model is improved.
[0070] Specifically, each step is described in detail below:
[0071] Step a, obtaining a pre-trained recognition model and a first recognition model, and selecting a first distillation node in the pre-trained recognition model and selecting a second distillation node in the first recognition model;
[0072] In this step, the image recognition system first selects a pre-trained recognition model and a first recognition model. When selecting the pre-trained recognition model and the first recognition model, the structural similarity between the first recognition model and the pre-trained recognition model should be kept as high as possible, which is not only conducive to the pre-trained recognition model to obtain higher generalization and fitting capabilities, but also convenient for selecting distillation nodes for feature map supervision. After determining the pre-trained recognition model and the first recognition model, the image recognition system selects the first distillation node in the pre-trained recognition model and the second distillation node in the first recognition model.
[0073] Furthermore, before selecting the pre-trained recognition model and the first recognition model, since the pre-trained recognition model needs to perform image recognition based on the first recognition model, the first recognition model needs to be trained so that the first recognition model can achieve better results in the image recognition task; for example: the pre-trained recognition model and the first recognition model are both applied in the field of image processing and are used to identify the content in the image. At this time, it is necessary to obtain massive image data as training image data, and annotate the content in the image data as verification data. The first recognition model is trained through the training image data, and the trained first recognition model is verified through the verification data until the first recognition model can accurately identify the content in the image, thereby obtaining the first recognition model that can be used to supervise the model distillation of the intermediate feature layer of the pre-trained recognition model.
[0074] Specifically, selecting a first distillation node in the pre-trained recognition model and selecting a second distillation node in the first recognition model include the following steps:
[0075] Step a1, the image recognition system selects the first target module in the pre-trained recognition model according to the preset task objectives and the structural information of the pre-trained recognition model; specifically, the preset task objectives are set by relevant personnel, which include covering the important features of the first recognition model, abstract information at different levels, the diversity and relevance of nodes, and the number of distillation nodes. Since the distillation node is usually after each module, the image recognition system includes multiple modules in the pre-trained recognition model according to the preset task objectives and the structural information of the pre-trained recognition model to determine which modules are followed by the distillation node, and these modules are the selected first target module in the pre-trained recognition model.
[0076] For example, Figure 3 As shown, the pre-trained recognition model (student model) contains n modules M with the same or similar structures. S , the image recognition system uses the n-1 modules after the input feature processing unit in the pre-trained recognition model as the first target module according to the preset task objectives and the structural information of the pre-trained recognition model.
[0077] Optionally, the image recognition system determines a first target module for every module interval, or a first target module for every two modules interval, etc. in the pre-trained recognition model according to preset task objectives and structural information of the pre-trained recognition model.
[0078] Step a2, after determining the first target module in the pre-trained recognition model, the image recognition system selects the second target module in the first recognition model according to the structural information of the first target module and the first recognition model; specifically, since a high structural similarity is maintained between the first recognition model and the pre-trained recognition model when selecting the model, the image recognition system can select the second target module in the first recognition model according to the position of the first target module in the pre-trained recognition model and the structural information of the first recognition model.
[0079] For example, Figure 3 As shown, the first recognition model (teacher model) and the pre-trained recognition model (student model) are the same, both of which contain n modules M with the same or similar structures. T Since the n-1 modules after the input feature processing unit in the pre-trained recognition model are the first target modules, the image recognition system regards the n-1 modules after the input feature processing unit in the first recognition model as the second target modules.
[0080] Optionally, in the pre-trained recognition model, a first target module is determined for every module interval, or a first target module is determined for every two modules interval, etc., and in the corresponding first recognition model, a first target module is also determined for every module interval, or a first target module is determined for every two modules interval, etc.
[0081] In step a3, after determining the first target module and the second target module in the pre-trained recognition model, the image recognition system determines a first distillation node after each first target module in the pre-trained recognition model, and determines a second distillation node after each second target module in the first recognition model. Figure 2 As shown, since the n-1 modules after the input feature processing unit in the pre-trained recognition model are the first target modules, and the n-1 modules after the input feature processing unit in the first recognition model are the second target modules, correspondingly, the second distillation node and the second distillation node are determined after the first n-1 modules in the pre-trained recognition model and the first recognition model, while it is not necessary to determine the second distillation node and the second distillation node after the nth module.
[0082] Further, after step a, the method further comprises:
[0083] The image recognition system obtains a first feature map channel corresponding to a first distillation node, and a second feature map channel corresponding to a second distillation node; it can be understood that the first feature map channel corresponding to the first distillation node is determined according to the feature map channel of a previous module of the first distillation node, and the second feature map channel corresponding to the second distillation node is determined according to the feature map channel of the previous module of the second distillation node. For example: the pre-trained recognition model includes three modules with the same or similar structures, there is a first distillation node between the first module and the second module, and there is a first distillation node between the second module and the third module. Correspondingly, the first feature map channel corresponding to the first distillation node between the first module and the second module is the same as the feature map channel of the first module, and the first feature map channel corresponding to the first distillation node between the second module and the third module is the same as the feature map channel of the second module.
[0084] After the image recognition system obtains the first feature map channel corresponding to the first distillation node and the second feature map channel corresponding to the second distillation node, the number of the first feature map channels and the second feature map channels of the corresponding first distillation node and the second distillation node are compared. If it is determined that the number of the first feature map channels is less than the number of the second feature map channels, the first feature map channels are amplified to make the number of the first feature map channels the same as the number of the second feature map channels. It should be noted that the first recognition model is generally a network model with a more complex structure, and the pre-trained recognition model is generally a network model with a simpler structure than the first recognition model. Therefore, the number of the first feature map channels corresponding to the first distillation node is usually less than or equal to the number of the second feature map channels corresponding to the second distillation node.
[0085] Exemplarily, the pre-trained recognition model includes three modules with the same or similar structures, there is a first distillation node between the first module and the second module, and there is a first distillation node between the second module and the third module. Similarly, the first recognition model includes three modules with the same or similar structures, there is a second distillation node between the first module and the second module, and there is a second distillation node between the second module and the third module. At this time, the image recognition system compares the number of first feature map channels corresponding to the first distillation node between the first module and the second module and the number of second feature map channels corresponding to the second distillation node between the first module and the second module, and compares the number of first feature map channels corresponding to the first distillation node between the second module and the third module and the number of second feature map channels corresponding to the second distillation node between the second module and the third module; so that the number of first feature map channels corresponding to the first distillation node between the first module and the second module and the number of second feature map channels corresponding to the second distillation node between the first module and the second module are the same, and the number of first feature map channels corresponding to the first distillation node between the second module and the third module and the number of second feature map channels corresponding to the second distillation node between the second module and the third module are the same.
[0086] Furthermore, the specific steps of amplifying the first feature map channels so that the number of the first feature map channels is the same as that of the second feature map channels include:
[0087] After the image recognition system determines that the number of channels of the first feature map is less than the number of channels of the second feature map, a target convolution kernel is generated according to the number of channels of the second feature map; specifically, the target convolution kernel is a two-dimensional convolution kernel, and the size of the two-dimensional convolution kernel is 1×1×C T , where C T is the number of channels in the second feature map. This two-dimensional convolution kernel is not learnable, that is, it does not participate in gradient update. This two-dimensional convolution kernel is also used to construct the distillation node loss function. The expression of the distillation node loss function is: Where ψ is the two-dimensional convolution kernel.
[0088] The image recognition system performs a convolution operation on the first feature map channel based on the target convolution kernel, and expands the number of the first feature map channels to the same number as the second feature map channels, so as to facilitate the calculation of the distillation node loss function.
[0089] Step b, setting a target loss function according to the first distillation node, the second distillation node, the pre-trained recognition model and the first recognition model;
[0090] In this step, after determining the first distillation node and the second distillation node, the image recognition system determines the distillation node loss function according to the first distillation node and the second distillation node, determines the model output loss function and the pre-trained recognition model loss function according to the pre-trained recognition model and the first recognition model, and then constructs the target loss function according to the distillation node loss function, the model output loss function and the pre-trained recognition model loss function. It should be noted that the distillation node loss function is a feature map loss function, which is used to supervise the intermediate feature layer of the pre-trained recognition model in the model distillation process of the supervised intermediate feature layer, the model output loss function is a loss function that reflects the relationship between the output of the first recognition model and the output of the pre-trained recognition model, and the pre-trained recognition model loss function is a loss function that reflects the relationship between the output of the pre-trained recognition model and the labeled data of the training image data.
[0091] Specifically, setting a target loss function according to the first distillation node, the second distillation node, the pre-trained recognition model and the first recognition model includes the following steps:
[0092] Step b1, obtaining preset weight values and first preset loss functions corresponding to the first distillation node and the second distillation node, and constructing a distillation node loss function by weighted processing based on the preset weight values and the preset first loss function;
[0093] In this step, the image recognition system obtains the preset weight value and the first preset loss function corresponding to the first distillation node and the second distillation node, and performs weighted processing based on the preset weight value and the preset first loss function to construct the distillation node loss function; specifically, the preset weights of the first preset loss function of the first distillation node and the second distillation node can be weighted in a progressive or uniformly distributed manner to construct the distillation node loss function, and the expression of the obtained distillation node loss function is:
[0094]
[0095] Where i is the total number of the first distillation nodes or the second distillation nodes. Since the structures of the pre-trained recognition model and the first recognition model are basically the same, the total number of the first distillation nodes is the same as the total number of the second distillation nodes. i is the preset weight value of the first preset loss function corresponding to the i-th first distillation node or the second distillation node, and are the output values of the i-th distillation node in the first recognition model and the pre-trained recognition model, ψ is the preset two-dimensional convolution kernel, L feat is the first preset loss function operator. It can be understood that is a first preset loss function, and the first distillation node or the second distillation node both adopts the first preset loss function.
[0096] Furthermore, the distillation node loss function can be set as follows according to specific circumstances: First, as the model distillation of the supervised intermediate feature layer proceeds, the preset weight value of the first preset loss function corresponding to each first distillation node or second distillation node is gradually increased, and then the distillation node loss function is constantly changing as the model distillation of the supervised intermediate feature layer proceeds. Second, as the training proceeds, the preset weight value of the first preset loss function corresponding to each first distillation node or second distillation node remains unchanged and evenly distributed, and then the distillation node loss function does not change as the model distillation of the supervised intermediate feature layer proceeds.
[0097] Step b2, obtaining a second preset loss function corresponding to the pre-trained recognition model and the first recognition model, and constructing a model output loss function and a pre-trained recognition model loss function based on the second preset loss function;
[0098] In this step, the image recognition system obtains the second preset loss function corresponding to the pre-trained recognition model and the first recognition model, and constructs the model output loss function and the pre-trained recognition model loss function based on the second preset loss function; it can be understood that the pre-trained recognition model and the first recognition model both use the same second preset loss function, and the model output loss function and the pre-trained recognition model loss function are both constructed based on the second preset loss function using different parameters.
[0099] Specifically, step b2 includes the following steps:
[0100] In step b21, the image recognition system obtains preset input data, and inputs the preset input data into the pre-trained recognition model and the first recognition model respectively, to obtain a first output corresponding to the pre-trained recognition model and a second output corresponding to the first recognition model; specifically, the preset input data input into the pre-trained recognition model and the first recognition model are the same data, which facilitates the subsequent construction of the model output loss function.
[0101] Step b22, the image recognition system constructs a model output loss function based on substituting the first output and the second output into a second preset loss function; specifically, the expression of the obtained model output loss function is: loss2 = L out (O T ,O S ), where O T is the second output of the first recognition model, O S is the first output of the pre-trained recognition model, L out It is the second preset loss function operator.
[0102] Step b23, the image recognition system obtains the labeled data corresponding to the preset input data, and substitutes the first output and the labeled data into the second preset loss function, thereby constructing a pre-trained recognition model loss function; specifically, the obtained pre-trained recognition model loss function is: loss3 = L out (GT,O S ), where GT is the labeled data corresponding to the preset input data, which belongs to the true value, S is the first output of the pre-trained recognition model, L out It is the second preset loss function operator.
[0103] Step b3, setting a target loss function according to the distillation node loss function, the model output loss function and the pre-trained recognition model loss function.
[0104] In this step, the image recognition system adds the distillation node loss function, the model output loss function, and the pre-trained recognition model loss function to obtain the target loss function. The expression of the obtained target loss function is:
[0105] Optionally, according to the actual situation, corresponding weight values can be added to the distillation node loss function, the model output loss function and the pre-training recognition model loss function, and then added together to obtain the expression of the target loss function: Among them, a, b, and c are weight values, and usually, b is the same as c.
[0106] Step c: acquiring training image data, and performing model distillation of the supervised intermediate feature layer of the pre-trained recognition model based on the training image data, the target loss function and the first recognition model to obtain a recognition model.
[0107] In this step, the image recognition system obtains training image data, inputs the training image data into the pre-trained recognition model and the first recognition model, obtains the outputs of the pre-trained recognition model and the first recognition model, the outputs of the first distillation node and the second distillation node, combines the above outputs and the target loss function to calculate the target loss value of the pre-trained recognition model, and back-propagates the target loss value to update the gradient of the pre-trained recognition model to supervise the intermediate feature layer of the pre-trained recognition model as well as the input layer and output layer of the pre-trained recognition model; repeats the above model distillation process of supervising the intermediate feature layer to obtain the recognition model.
[0108] Specifically, step c includes:
[0109] Step c1, inputting the training image data into the pre-trained recognition model and the first recognition model respectively, obtaining a first prediction value corresponding to the pre-trained recognition model, a second prediction value corresponding to the first recognition model, a third prediction value corresponding to the first distillation node, and a fourth prediction value corresponding to the second distillation node;
[0110] In this step, the image recognition system inputs the training image data into the pre-trained recognition model and the first recognition model respectively, and obtains the first prediction value corresponding to the pre-trained recognition model, the second prediction value corresponding to the first recognition model, the third prediction value corresponding to the first distillation node, and the fourth prediction value corresponding to the second distillation node, wherein the first prediction value is the output value of the pre-trained recognition model, the second prediction value is the output value of the first recognition model, the third prediction value is the output value of the first distillation node, that is, the feature map output by the intermediate feature layer of the pre-trained recognition model, and the fourth prediction value is the output value of the second distillation node, that is, the feature map output by the intermediate feature layer of the first recognition model.
[0111] Step c2, obtaining a true value corresponding to the training image data, and substituting the true value, the first predicted value, the second predicted value, the third predicted value and the fourth predicted value into the target loss function to obtain a target loss value;
[0112] In this step, the image recognition system obtains the true value corresponding to the training image data, and substitutes the true value, the first predicted value, the second predicted value, the third predicted value and the fourth predicted value into the target loss function to obtain the target loss value. Specifically, the expression of the target loss function is: in, and are the fourth predicted value and the third predicted value respectively, O T is the second predicted value, O A is the first predicted value and GT is the true value.
[0113] Step c3, until the target loss value meets the preset conditions, the recognition model is obtained.
[0114] In this step, the image recognition system loops through the above image recognition steps, and after each loss value is obtained, determines whether the target loss value meets the preset conditions; if the target loss value meets the preset conditions, it determines that the pre-trained recognition model has completed image recognition and obtains the recognition model; if the target loss value meets the preset conditions, the gradient of the pre-trained recognition model is updated based on the target loss value back-propagation to supervise the intermediate feature layer of the pre-trained recognition model and the input layer and output layer of the pre-trained recognition model until the target loss value meets the preset conditions and obtains the recognition model. The preset condition can be that the target loss value is less than a preset threshold, or that the target loss value obtained for several consecutive times is less than the preset threshold.
[0115] The image recognition system of this embodiment obtains a pre-trained recognition model and a first recognition model, and selects a first distillation node in the pre-trained recognition model and a second distillation node in the first recognition model; sets a target loss function according to the first distillation node, the second distillation node, the pre-trained recognition model and the first recognition model; obtains training image data, and performs image recognition on the pre-trained recognition model based on the training image data, the target loss function and the first recognition model to obtain a recognition model. By adding the loss function corresponding to the distillation node to the loss function while performing image recognition, and then supervising the intermediate feature layer of the pre-trained recognition model, the generalization ability of the pre-trained recognition model obtained after image recognition is effectively improved.
[0116] This embodiment also provides an image recognition device, which may specifically include a background server corresponding to various application programs with comment functions, such as Figure 4 As shown, the image recognition device may include:
[0117] An acquisition unit 1001 is used to acquire an image to be recognized;
[0118] The processing unit 1002 is used to process the image to be recognized by using a recognition model to obtain output data;
[0119] The generating unit 1003 is used to generate a target image corresponding to the image to be identified based on the output data and the category data.
[0120] In an optional example, the processing unit is further configured to:
[0121] Inputting the image to be identified into a recognition model, performing feature extraction on the image to be identified through the recognition model, and obtaining depth feature information and semantic feature information of the image to be identified;
[0122] The depth feature information and the semantic feature information are processed by the recognition model to obtain output data.
[0123] In an optional example, the processing unit is further configured to:
[0124] Performing feature fusion on the depth feature information and the semantic feature information through the recognition model to obtain a down-sampled feature map corresponding to the image to be recognized;
[0125] Upsampling the downsampled feature map by using the recognition model to obtain an upsampled feature map with the same size as the image to be recognized;
[0126] The up-sampled feature map is subjected to feature analysis by the recognition model to obtain output data.
[0127] In an optional example, the generation unit is further configured to:
[0128] Obtaining a set of category prediction probabilities corresponding to each pixel point of the image to be identified in the output data;
[0129] Determine the category to which each pixel belongs according to the category prediction probability set corresponding to each pixel;
[0130] Based on the category and category data of each pixel point, each pixel point is labeled to generate a target image corresponding to the image to be identified.
[0131] In an optional example, the image recognition device further includes a training unit, and the training unit is used to:
[0132] Acquire a pre-trained recognition model and a first recognition model, and select a first distillation node in the pre-trained recognition model and a second distillation node in the first recognition model;
[0133] Setting a target loss function according to the first distillation node, the second distillation node, the pre-trained recognition model, and the first recognition model;
[0134] Training image data is acquired, and based on the training image data, the target loss function and the first recognition model, model distillation of the supervised intermediate feature layer of the pre-trained recognition model is performed to obtain a recognition model.
[0135] In an optional example, the training unit is further configured to:
[0136] Obtaining preset weight values and a first preset loss function corresponding to the first distillation node and the second distillation node, and constructing a distillation node loss function by performing weighted processing based on the preset weight values and the preset first loss function;
[0137] Obtaining a second preset loss function corresponding to the pre-trained recognition model and the first recognition model, and constructing a model output loss function and a pre-trained recognition model loss function based on the second preset loss function;
[0138] A target loss function is set according to the distillation node loss function, the model output loss function and the pre-trained recognition model loss function.
[0139] In an optional example, the training unit is further configured to:
[0140] Inputting the training image data into the pre-trained recognition model and the first recognition model respectively, obtaining a first prediction value corresponding to the pre-trained recognition model, a second prediction value corresponding to the first recognition model, a third prediction value corresponding to the first distillation node, and a fourth prediction value corresponding to the second distillation node;
[0141] Obtaining a true value corresponding to the training image data, and substituting the true value, the first predicted value, the second predicted value, the third predicted value, and the fourth predicted value into the target loss function to obtain a target loss value;
[0142] Until the target loss value meets the preset conditions, the recognition model is obtained.
[0143] By adopting the scheme of this embodiment, an image to be identified is obtained; the image to be identified is processed by a recognition model to obtain output data; based on the output data and category data, a target image corresponding to the image to be identified is generated. The model distillation of the supervised intermediate feature layer effectively reduces the computational complexity of the recognition model and improves the recognition accuracy of the recognition model, and the image to be identified is processed by the recognition model to obtain output data, and based on the output data and category data, a target image corresponding to the image to be identified is generated, ensuring that the accuracy of image recognition is improved under the premise of a small amount of computation.
[0144] Accordingly, the present application also provides an image recognition system, such as Figure 5 As shown, Figure 5 A schematic diagram of the structure of an image recognition system provided in an embodiment of the present application. The image recognition system 1100 includes a processor 1101 having one or more processing cores, a memory 1102 having one or more computer-readable storage media, and a computer program stored in the memory 1102 and executable on the processor. The processor 1101 is electrically connected to the memory 1102. It will be appreciated by those skilled in the art that the image recognition system structure shown in the figure does not constitute a limitation on the image recognition system, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0145] The processor 1101 is the control center of the image recognition system 1100, and uses various interfaces and lines to connect various parts of the entire electronic device 1100. By running or loading software programs and / or units stored in the memory 1102, and calling data stored in the memory 1102, the processor 1101 executes various functions of the electronic device 1100 and processes data, thereby monitoring the image recognition system 1100 as a whole. The processor 1101 can be a processor CPU, a graphics processor GPU, a network processor (Network Processor, NP), etc., and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application.
[0146] In an embodiment of the present application, the processor 1101 in the image recognition system 1100 will load the instructions corresponding to the processes of one or more applications into the memory 1102 in accordance with the following steps, and the processor 1101 will run the application stored in the memory 1102 to execute any image recognition method provided in the embodiment of the present application.
[0147] Optional, such as Figure 5 As shown, the image recognition system 1100 further includes: a touch screen 1103, a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. The processor 1101 is electrically connected to the touch screen 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107, respectively. Those skilled in the art can understand that Figure 4 The image recognition system structure shown in the figure does not constitute a limitation on the electronic device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0148] The touch display screen 1103 can be used to display a graphical user interface and receive operation instructions generated by the user acting on the graphical user interface. The touch display screen 1103 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user and various graphical user interfaces of the electronic device, and these graphical user interfaces can be composed of graphics, text, icons, videos and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD, Liquid Crystal Display), an organic light-emitting diode (OLED, Organic Light-EmittingDiode) and the like. The touch panel can be used to collect the user's touch operation on or near it (such as the user uses any suitable object or attachment such as a finger, a stylus, etc. on the touch panel or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute corresponding programs. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 1101, and can receive the command sent by the processor 1101 and execute it. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 1101 to determine the type of touch event, and then the processor 1101 provides a corresponding visual output on the display panel according to the type of touch event. In an embodiment of the present application, the touch panel and the display panel can be integrated into the touch display screen 1103 to realize the input and output functions. However, in some embodiments, the touch panel and the touch panel can be used as two independent components to realize the input and output functions. That is, the touch display screen 1103 can also be used as a part of the input unit 1106 to realize the input function.
[0149] The radio frequency circuit 1104 may be used to send and receive radio frequency signals, so as to establish wireless communication with a network device or other electronic devices through wireless communication, and to send and receive signals between the network device or other electronic devices.
[0150] The audio circuit 1105 can be used to provide an audio interface between the user and the electronic device through a speaker and a microphone. The audio circuit 1105 can transmit the electrical signal converted from the received audio data to the speaker, which is converted into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 1105 and converted into audio data, and then the audio data is output to the processor 1101 for processing, and then sent to another electronic device through the radio frequency circuit 1104, or the audio data is output to the memory 1102 for further processing. The audio circuit 1105 may also include an earplug jack to provide communication between an external headset and an electronic device.
[0151] The input unit 1106 may be used to receive input numbers, character information or user feature information (such as fingerprint, iris, facial information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0152] The power supply 1107 is used to supply power to various components of the electronic device 1100. Optionally, the power supply 1107 can be logically connected to the processor 1101 through a power management system, so that the power management system can manage charging, discharging, and power consumption. The power supply 1107 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0153] although Figure 5 Not shown, the electronic device 1100 may also include a camera, a sensor, a wireless fidelity module, a Bluetooth module, etc., which will not be described in detail here.
[0154] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0155] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0156] To this end, an embodiment of the present application provides a computer-readable storage medium, in which multiple computer programs are stored. The computer programs can be loaded by a processor to execute any image recognition method provided by the embodiment of the present application.
[0157] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0158] Since the computer program stored in the computer-readable storage medium can execute any image recognition method provided in the embodiments of the present application, the beneficial effects that can be achieved by any image recognition method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0159] According to one aspect of the present application, a computer program product or a computer program is also provided, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods provided in various optional implementations of the above embodiments.
[0160] In the above-mentioned image recognition device, computer-readable storage medium, image recognition system, and computer program product embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process and beneficial effects of the above-described image recognition device, computer-readable storage medium, computer program product, image recognition system, and corresponding units can refer to the description of the image recognition method in the above embodiment, and will not be repeated here.
[0161] The above is a detailed introduction to an image recognition method, device, system, computer-readable storage medium and computer program product provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. An image recognition method, It is characterized in that The image recognition method comprises: Obtain an image to be recognized; Processing the image to be recognized by using a recognition model to obtain output data; Based on the output data and the category data, a target image corresponding to the image to be recognized is generated.
2. The image recognition method according to claim 1, It is characterized in that The step of processing the image to be recognized by using a recognition model to obtain output data includes: Inputting the image to be identified into a recognition model, performing feature extraction on the image to be identified through the recognition model, and obtaining depth feature information and semantic feature information of the image to be identified; The depth feature information and the semantic feature information are processed by the recognition model to obtain output data.
3. The image recognition method according to claim 2, It is characterized in that The processing of the depth feature information and the semantic feature information by the recognition model to obtain output data includes: Performing feature fusion on the depth feature information and the semantic feature information through the recognition model to obtain a down-sampled feature map corresponding to the image to be recognized; Upsampling the downsampled feature map by using the recognition model to obtain an upsampled feature map with the same size as the image to be recognized; The up-sampled feature map is subjected to feature analysis by the recognition model to obtain output data.
4. The image recognition method according to claim 1, It is characterized in that The step of generating a target image corresponding to the image to be identified based on the output data and the category data includes: Obtaining a set of category prediction probabilities corresponding to each pixel point of the image to be identified in the output data; Determining the category to which each pixel belongs according to the category prediction probability set corresponding to each pixel; Based on the category and category data of each pixel point, each pixel point is labeled to generate a target image corresponding to the image to be identified.
5. The image recognition method according to claim 1, It is characterized in that Before obtaining the image to be identified, the method includes: Acquire a pre-trained recognition model and a first recognition model, and select a first distillation node in the pre-trained recognition model and a second distillation node in the first recognition model; Setting a target loss function according to the first distillation node, the second distillation node, the pre-trained recognition model, and the first recognition model; Training image data is acquired, and based on the training image data, the target loss function and the first recognition model, model distillation of the supervised intermediate feature layer of the pre-trained recognition model is performed to obtain a recognition model.
6. The image recognition method according to claim 5, It is characterized in that The setting of the target loss function according to the first distillation node, the second distillation node, the pre-trained recognition model and the first recognition model includes: Obtaining preset weight values and a first preset loss function corresponding to the first distillation node and the second distillation node, and constructing a distillation node loss function by performing weighted processing based on the preset weight values and the preset first loss function; Obtaining a second preset loss function corresponding to the pre-trained recognition model and the first recognition model, and constructing a model output loss function and a pre-trained recognition model loss function based on the second preset loss function; A target loss function is set according to the distillation node loss function, the model output loss function and the pre-trained recognition model loss function.
7. The image recognition method according to claim 5, It is characterized in that The step of performing a model distillation of a supervised intermediate feature layer on the pre-trained recognition model based on the training image data, the target loss function and the first recognition model to obtain a recognition model comprises: Inputting the training image data into the pre-trained recognition model and the first recognition model respectively, obtaining a first prediction value corresponding to the pre-trained recognition model, a second prediction value corresponding to the first recognition model, a third prediction value corresponding to the first distillation node, and a fourth prediction value corresponding to the second distillation node; Obtaining a true value corresponding to the training image data, and substituting the true value, the first predicted value, the second predicted value, the third predicted value, and the fourth predicted value into the target loss function to obtain a target loss value; Until the target loss value meets the preset conditions, the recognition model is obtained.
8. An image recognition device, It is characterized in that The device comprises: An acquisition unit, used for acquiring an image to be recognized; A processing unit, used for processing the image to be recognized by using a recognition model to obtain output data; A generating unit, configured to generate a target image corresponding to the image to be identified based on the output data and the category data; Preferably, the processing unit processes the image to be recognized through a recognition model to obtain output data, including: Inputting the image to be identified into a recognition model, performing feature extraction on the image to be identified through the recognition model, and obtaining depth feature information and semantic feature information of the image to be identified; Processing the depth feature information and the semantic feature information through the recognition model to obtain output data; Preferably, the processing unit processes the depth feature information and the semantic feature information through the recognition model to obtain output data, including: Performing feature fusion on the depth feature information and the semantic feature information through the recognition model to obtain a down-sampled feature map corresponding to the image to be recognized; Upsampling the downsampled feature map by using the recognition model to obtain an upsampled feature map with the same size as the image to be recognized; Performing feature analysis on the up-sampled feature map through the recognition model to obtain output data; Preferably, the generating unit generates a target image corresponding to the image to be identified based on the output data and the category data, including: Obtaining a set of category prediction probabilities corresponding to each pixel point of the image to be identified in the output data; Determining the category to which each pixel belongs according to the category prediction probability set corresponding to each pixel; Based on the category and category data of each pixel point, each pixel point is labeled to generate a target image corresponding to the image to be identified; Preferably, the recognition model is obtained by the processing unit by executing the following steps: Acquire a pre-trained recognition model and a first recognition model, and select a first distillation node in the pre-trained recognition model and a second distillation node in the first recognition model; Setting a target loss function according to the first distillation node, the second distillation node, the pre-trained recognition model, and the first recognition model; Acquire training image data, and perform model distillation of a supervised intermediate feature layer on the pre-trained recognition model based on the training image data, the target loss function, and the first recognition model to obtain a recognition model; Preferably, the processing unit sets a target loss function according to the first distillation node, the second distillation node, the pre-trained recognition model and the first recognition model, including: Obtaining preset weight values and a first preset loss function corresponding to the first distillation node and the second distillation node, and constructing a distillation node loss function by performing weighted processing based on the preset weight values and the preset first loss function; Obtaining a second preset loss function corresponding to the pre-trained recognition model and the first recognition model, and constructing a model output loss function and a pre-trained recognition model loss function based on the second preset loss function; Setting a target loss function according to the distillation node loss function, the model output loss function and the pre-trained recognition model loss function; Preferably, the processing unit performs a model distillation of a supervised intermediate feature layer on the pre-trained recognition model based on the training image data, the target loss function and the first recognition model to obtain a recognition model, comprising: Inputting the training image data into the pre-trained recognition model and the first recognition model respectively, obtaining a first prediction value corresponding to the pre-trained recognition model, a second prediction value corresponding to the first recognition model, a third prediction value corresponding to the first distillation node, and a fourth prediction value corresponding to the second distillation node; Obtaining a true value corresponding to the training image data, and substituting the true value, the first predicted value, the second predicted value, the third predicted value, and the fourth predicted value into the target loss function to obtain a target loss value; Until the target loss value meets the preset conditions, the recognition model is obtained.
9. An image recognition system, It is characterized in that It comprises a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps of the image recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps of the image recognition method according to any one of claims 1 to 7.