Image recognition method, image recognition device and storage medium

By jointly training the feature extraction model with depth estimation and main target segmentation model, the problem of low image recognition accuracy is solved, and higher feature extraction accuracy and image recognition accuracy are achieved.

CN120047686APending Publication Date: 2025-05-27ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510115840.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, the image recognition accuracy is low and is easily affected by shooting angle, lighting, occlusion and background information, resulting in recognition errors.

Method used

The pre-trained feature extraction model is used to train together with the constructed knowledge distillation structure, including depth estimation and main target segmentation model. Through this joint training, the accuracy of the feature extraction model is improved and background interference is reduced.

Benefits of technology

The accuracy of feature extraction model when extracting image features is improved, the accuracy of image recognition is enhanced, and the dependence on background information is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047686A_ABST
    Figure CN120047686A_ABST
Patent Text Reader

Abstract

The invention relates to an image recognition method, an image recognition device and a storage medium, and the method comprises the steps: obtaining a to-be-recognized image; performing image feature extraction on the to-be-recognized image according to a pre-trained feature extraction model to obtain a target image feature; wherein the pre-trained feature extraction model is obtained through combined training with a pre-constructed first knowledge distillation structure and a pre-constructed second knowledge distillation structure; the first knowledge distillation structure comprises a depth estimation teacher model and a depth estimation student model; the second knowledge distillation structure comprises a main target segmentation teacher model and a main target segmentation student model; and obtaining an image recognition result based on the target image features. According to the method and the device, the interference of surrounding environment factors when the feature extraction model extracts the image features is reduced, the accuracy of extracting the image features by the feature extraction model is improved, and the image recognition accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image recognition, and in particular, to an image recognition method, an image recognition device, and a storage medium. Background Art

[0002] Object recognition and visual search technologies can greatly shorten the distance between the physical world and the data world, and help users quickly and conveniently obtain information. In the currently highly concerned Internet field, for the to-be-recognized images obtained by taking pictures through a camera or downloaded from the Internet, image recognition technology can find the matching images from the pre-registered images, and further obtain the relevant information of the to-be-recognized images (this process is usually also called the image recognition process). For example, by retrieving the book cover image, information such as the name and author of the book can be obtained.

[0003] With the development of deep learning technology, great progress has been made in image recognition technology in recent years. Currently, Neural Networks is a commonly used image recognition model. Generally, the original neural network is trained by obtaining image data as training samples to obtain a trained neural network. The image containing the to-be-recognized content is used as the input of the trained neural network to recognize the relevant information of the to-be-recognized content. However, in the actual image recognition process, due to the influence of factors such as shooting angle, illumination, and occlusion, the extraction of image features is easily affected by the background information around the target, and there are cases of false detection in model recognition.

[0004] In view of the problem of low image recognition accuracy in the related art, no effective solution has been proposed yet. Summary of the Invention

[0005] In this embodiment, an image recognition method, an image recognition device, and a storage medium are provided to solve the problem of low image recognition accuracy in the related art.

[0006] In a first aspect, in this embodiment, an image recognition method is provided, including:

[0007] Obtain a to-be-recognized image;

[0008] Extract image features from the to-be-recognized image according to a pre-trained feature extraction model to obtain target image features; wherein, the pre-trained feature extraction model is jointly trained with a pre-constructed first knowledge distillation structure and a second knowledge distillation structure; the first knowledge distillation structure includes a depth estimation teacher model and a depth estimation student model; the second knowledge distillation structure includes a main target segmentation teacher model and a main target segmentation student model;

[0009] Based on the target image features, obtain an image recognition result.

[0010] In some of these embodiments, during the joint training process:

[0011] Based on the output data of the depth estimation teacher model and the output data of the depth estimation student model, the feature extraction model and the depth estimation student model being trained are updated synchronously;

[0012] Based on the output data of the main object segmentation teacher model and the output data of the main object segmentation student model, the feature extraction model and the main object segmentation student model being trained are updated synchronously.

[0013] In some of these embodiments, during the joint training process:

[0014] The input data for model training is simultaneously input into the feature extraction model, the depth estimation teacher model, and the main object segmentation teacher model being trained, and the output data of the feature extraction model being trained is input into the depth estimation student model and the main object segmentation student model;

[0015] Taking the output data of the depth estimation teacher model as the first reference information, based on the first reference information and the output data of the depth estimation student model, the feature extraction model and the depth estimation student model being trained are updated synchronously;

[0016] Taking the output data of the main object segmentation teacher model as the second reference information, based on the second reference information and the output data of the main object segmentation student model, the feature extraction model and the main object segmentation student model being trained are updated synchronously.

[0017] In some of these embodiments, during the process of training the image feature extraction model, it further includes:

[0018] Taking the output data of the main object segmentation student model as the first output data and the output data of the feature extraction model being trained as the second output data, based on the first output data, the second output data, and a preset image feature pool, the second output data is subjected to feature enhancement processing to obtain enhanced features;

[0019] The enhanced features are subjected to dimensionality compression processing to obtain compressed enhanced features;

[0020] Based on the compressed enhanced features, the first extraction accuracy information currently corresponding to the feature extraction model being trained is determined, and the feature extraction model is updated according to the first extraction accuracy information.

[0021] In some of these embodiments, performing feature enhancement processing on the second output data according to the first output data, the second output data, and a preset image feature pool to obtain enhanced features includes:

[0022] Weighting the first output data and the second output data to obtain weighted features;

[0023] Calculating the similarity between the weighted features and reference features in the preset feature pool;

[0024] Obtaining the reference feature in the image feature pool with the highest similarity to the weighted features as the feature to be stitched;

[0025] Stitching the feature to be stitched and the weighted features to obtain enhanced features.

[0026] In some of these embodiments, during the process of training the feature extraction model, it further includes:

[0027] Calculating the usage frequency of each reference feature in the image feature pool;

[0028] In the image feature pool, replacing the reference features with usage frequencies less than a preset frequency threshold based on the weighted features.

[0029] In some of these embodiments, during the process of training the feature extraction model, it further includes:

[0030] Updating the feature extraction model according to the degree of approximation between the second output data and the weighted features.

[0031] In some of these embodiments, during the process of training the feature extraction model, it further includes:

[0032] Performing mean and max pooling processing on the second output data to obtain global features;

[0033] Performing weighted summation processing on the second output data according to the similarity between the global features and the second output data to obtain compressed features;

[0034] Determining the second extraction accuracy information of the feature extraction model during training according to the compressed features;

[0035] Updating the feature extraction model according to the second extraction accuracy information.

[0036] In some of these embodiments, obtaining the image recognition result based on the target image features includes:

[0037] Calculate the similarity between the target image features and the standard features to obtain the target similarity; the standard features are image features extracted in advance from a standard image containing the target recognition object;

[0038] Obtain the image recognition result for the target recognition object according to the target similarity.

[0039] In a second aspect, an image recognition device is provided in this embodiment, including: an acquisition module, a feature extraction module, and an identification module, where,

[0040] The acquisition module is used to acquire the image to be recognized;

[0041] The feature extraction module is used to extract image features from the image to be recognized according to a pre-trained feature extraction model to obtain target image features; wherein, the pre-trained feature extraction model is jointly trained with a pre-constructed first knowledge distillation structure and a second knowledge distillation structure; the first knowledge distillation structure includes a depth estimation teacher model and a depth estimation student model; the second knowledge distillation structure includes a main target segmentation teacher model and a main target segmentation student model;

[0042] The identification module is used to obtain an image recognition result based on the target image features.

[0043] In a third aspect, a storage medium is provided in this embodiment, on which a computer program is stored, and when the program is executed by a processor, it implements the image recognition method described in the first aspect above.

[0044] Compared with the related art, the image recognition method provided in this embodiment obtains the image to be recognized; extracts image features from the image to be recognized according to a pre-trained feature extraction model to obtain target image features; wherein, the pre-trained feature extraction model is jointly trained with a pre-constructed first knowledge distillation structure and a second knowledge distillation structure; the first knowledge distillation structure includes a depth estimation teacher model and a depth estimation student model; the second knowledge distillation structure includes a main target segmentation teacher model and a main target segmentation student model; and obtains an image recognition result based on the target image features. It reduces the interference of surrounding environmental factors when the feature extraction model extracts image features, improves the accuracy of the feature extraction model in extracting image features, and improves the image recognition accuracy.

[0045] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The accompanying drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:

[0047] Figure 1 is a hardware structure block diagram of a terminal of the image recognition method of this embodiment.

[0048] Figure 2 is a flowchart of the image recognition method of this embodiment.

[0049] Figure 3 is a flowchart of feature compression in the process of training a feature extraction model of this embodiment.

[0050] Figure 4 is a training schematic diagram of a feature extraction model of some of these embodiments.

[0051] Figure 5 is a schematic diagram of the data flow of a feature extraction model of some of these embodiments.

[0052] Figure 6 is a flowchart of a training method of a feature extraction model of some of these embodiments.

[0053] Figure 7 is a structure block diagram of the image recognition device of this embodiment. Detailed implementation manners

[0054] To more clearly understand the purpose, technical solution and advantages of the present application, the present application will be described and illustrated below with reference to the accompanying drawings and embodiments.

[0055] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. In this application, words such as "a", "an", "one kind", "the", "these" and the like do not indicate a limitation in quantity, and they can be singular or plural. The terms "including", "containing", "having" and any variants thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device containing a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in this application do not limit to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third" and the like involved in this application only distinguish similar objects and do not represent a specific order for the objects.

[0056] The method embodiments provided in this embodiment can be executed on a terminal, a computer or a similar computing device. For example, running on a terminal, Figure 1 is the hardware structure block diagram of the terminal of the image recognition method in this embodiment. As Figure 1 shown, the terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 and a memory 104 for storing data. Among them, the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA. The above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown.

[0057] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the image recognition method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.

[0058] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by a communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0059] In this embodiment, an image recognition method is provided. Figure 2 is a flowchart of the image recognition method of this embodiment, as Figure 2 shown, and the process includes the following steps:

[0060] Step S201, obtain the image to be recognized;

[0061] Specifically, the image to be recognized can come from different devices, including but not limited to digital cameras, scanners, video surveillance systems, mobile devices (such as smart phones and tablet computers). Among them, the image to be recognized can be an image captured for recognition objects such as natural landscapes, animals, plants, transportation tools, and daily items, and this embodiment does not make specific limitations on this.

[0062] Step S202, perform image feature extraction on the image to be recognized according to a pre-trained feature extraction model to obtain target image features; wherein, the pre-trained feature extraction model is jointly trained with a pre-constructed first knowledge distillation structure and a second knowledge distillation structure; the first knowledge distillation structure includes a depth estimation teacher model and a depth estimation student model; the second knowledge distillation structure includes a main target segmentation teacher model and a main target segmentation student model.

[0063] Specifically, in order to improve the accuracy of the feature extraction model in extracting image features, the feature extraction model is jointly trained with a pre-constructed first knowledge distillation structure and a second knowledge distillation structure. Among them, the first knowledge distillation structure includes a depth estimation teacher model and a depth estimation student model, and the second knowledge distillation structure includes a main object segmentation teacher model and a main object segmentation student model. The joint training of the first knowledge distillation structure improves the attention of the feature extraction model to the main object, and the joint training of the second knowledge distillation structure enables the feature extraction model to reduce the interference of background information during feature extraction, thereby improving the accuracy of the feature extraction model in extracting image features and the accuracy of image recognition. It can be understood that during the joint training process, the feature extraction model is updated together with the depth estimation student model and the main object segmentation student model, while the depth estimation teacher model and the main object segmentation teacher model are not updated.

[0064] Step S203: Obtain an image recognition result based on the target image features.

[0065] Specifically, the image is recognized according to the extracted image features. The target similarity between the target image features and the standard features corresponding to a certain recognition object obtained in advance can be calculated, and the recognition result is obtained based on the target similarity. For example, when it is determined that the target similarity between the target image features and the standard features of "cat" collected in advance is higher than the preset similarity threshold, the main object in the image to be recognized is recognized as "cat".

[0066] It should also be noted that during the training process, the first knowledge distillation structure and the second knowledge distillation structure are constructed in this embodiment as auxiliary training structures, so that the feature extraction model can exclude the interference of background information and improve the main object recognition ability. After the training of the feature extraction model is completed and the inference stage is deployed, only the trained feature extraction model can be used for feature extraction, and the distance metric is used to calculate the distance between features to determine the similarity between different images, so as to minimize the inference in the inference stage, reduce the performance requirements for hardware, and achieve lightweight deployment.

[0067] In the related art, image recognition is often directly performed through a neural network model. The feature extraction process of the image is easily affected by the background information around the target, resulting in inaccurate extracted image features, and further resulting in low accuracy of the image recognition result based on the image feature extraction.

[0068] In this embodiment, through the above steps S201 to S203, first, an image to be recognized is obtained; then, image features of the image to be recognized are extracted according to a pre-trained feature extraction model to obtain target image features; wherein, the pre-trained feature extraction model is jointly trained with a pre-constructed first knowledge distillation structure and a second knowledge distillation structure; the first knowledge distillation structure includes a depth estimation teacher model and a depth estimation student model; the second knowledge distillation structure includes a main target segmentation teacher model and a main target segmentation student model; finally, based on the target image features, an image recognition result is obtained. Compared with directly performing image recognition through a neural network model in the prior art, this embodiment performs image recognition based on the feature extraction model jointly trained with the first knowledge distillation structure and the second knowledge distillation structure, reducing the interference of surrounding environmental factors when the feature extraction model extracts image features, improving the accuracy of the feature extraction model in extracting image features, and thus improving the accuracy of image recognition.

[0069] In some of these embodiments, during the joint training process: based on the output data of the depth estimation teacher model and the output data of the depth estimation student model, the feature extraction model and the depth estimation student model being trained are updated synchronously; based on the output data of the main target segmentation teacher model and the output data of the main target segmentation student model, the feature extraction model and the main target segmentation student model being trained are updated synchronously.

[0070] Specifically, during the training process, each model can be trained with training image samples. Among them, the image features of the training image samples can be any kind of image features in the field of image processing that are convenient for subsequent image recognition, such as natural landscape image features, animal image features, or building image features. Specifically, the feature extraction model A being trained samples the training image samples by category. Each time, p categories of data are selected from the training image sample data, and k samples are selected from each category to form a batch B of p×k samples, and B is input into the feature extraction model A to extract image features. Among them, the feature extraction model A can be a convolutional neural network or a vit network, and image features F 1 ∈R (p×k)×d×M×N are obtained, where d is the image feature dimension, and M and N are the resolutions of the image feature map.

[0071] Currently, in image retrieval technology, it is vulnerable to factors such as shooting angle, illumination, occlusion, and surrounding background information, resulting in errors in the recognition of image targets and inaccurate image recognition. Therefore, to solve this problem, in this embodiment, the feature extraction model A is trained using a joint training strategy to improve the attention of the feature extraction model A to the main target. Specifically, in this embodiment, the joint training strategy adopts a first knowledge distillation structure and a second knowledge distillation structure. The first knowledge distillation structure includes a depth estimation teacher model and a depth estimation student model; the first knowledge distillation structure includes a depth estimation teacher model and a depth estimation student model. Among them, the first knowledge distillation structure mainly calculates the distance of the pixel points of the target subject in the image from the camera, converts the RGB image into a grayscale image, distinguishes the target subject from the surrounding background environment through the grayscale of the pixel points, and a three-dimensional target subject can be extracted therefrom. The training image samples are trained through the first knowledge distillation structure, and combined with the constructed depth estimation loss function, the depth estimation loss is calculated. Then, the training image samples are trained through the second knowledge distillation structure, and combined with the constructed main target segmentation loss function, the main target segmentation loss is calculated. Based on the loss value of the depth estimation loss function, the feature extraction model A and the depth estimation student model are updated; based on the loss value of the main target segmentation loss function, the feature extraction model A and the main target segmentation student model are updated to obtain the trained feature extraction model A. Through this joint training strategy, the feature extraction model A pays more attention to the main target, reduces the interference of the surrounding background information, improves the accuracy of the feature extraction model A in extracting image features, and improves the accuracy of image recognition.

[0072] Additionally, in one embodiment, during the joint training process: the input data for model training is simultaneously input into the feature extraction model being trained, the depth estimation teacher model, and the main target segmentation teacher model, and the output data of the feature extraction model being trained is input into the depth estimation student model and the main target segmentation student model; using the output data of the depth estimation teacher model as the first reference information, based on the first reference information and the output data of the depth estimation student model, the feature extraction model being trained and the depth estimation student model are updated synchronously; using the output data of the main target segmentation teacher model as the second reference information, based on the second reference information and the output data of the main target segmentation student model, the feature extraction model being trained and the main target segmentation student model are updated synchronously.

[0073] Specifically, during the training process, the training image samples can be used as input data and input into the depth estimation teacher model to generate depth estimation supervision information D pe , and this depth estimation supervision information D peAs the first reference information. Input the training image sample into the feature extraction model to obtain the image feature F output by the feature extraction model 1 . In the depth estimation student model, use the image feature F 1 as the input to predict the depth information D of the training image sample ps . The depth information D of the training image sample ps is the distance from the pixel point in the training image sample to the camera. Convert the distance into a grayscale value to distinguish the target object and the surrounding background information through the pixel point grayscale value. Calculate the depth estimation loss L ps between the depth information D pe and the depth estimation supervision information D p . The specific calculation formula is as follows:

[0074]

[0075] Update the parameters of the feature extraction model and the depth estimation student model during training through the loss value L of the depth estimation loss function. Among them, the parameters of the depth estimation teacher model are fixed and do not participate in parameter update. p

[0076] Input the training image sample into the main object segmentation teacher model to obtain the main object supervision information S e . Use the main object supervision information S e as the second reference information; input the image feature F 1 into the main object segmentation student model to obtain the main object prediction information S s . Calculate the loss L s of the main object segmentation loss function between the main object prediction information S e and the main object supervision information S s . The specific calculation formula is as follows:

[0077]

[0078] Update the parameters of the feature extraction model and the main object segmentation student model during training through the loss value L of the main object segmentation loss function. Among them, the main object segmentation teacher model also does not participate in the update of training parameters. Based on this, this embodiment can realize the synchronous update of the main object segmentation student model and the feature extraction model, making the trained feature extraction model pay more attention to the main object. s

[0079] ​​In some of these embodiments, during the training of the image feature extraction model, it further includes: using the output data of the main target segmentation student model as the first output data, using the output data of the feature extraction model during training as the second output data, and performing feature enhancement processing on the second output data according to the first output data, the second output data, and a preset image feature pool to obtain enhanced features; performing dimensionality compression processing on the enhanced features to obtain compressed enhanced features; based on the compressed enhanced features, determining the current corresponding first extraction accuracy information of the feature extraction model during training, and updating the feature extraction model according to the first extraction accuracy information.

[0080] Specifically, in each training iteration, after obtaining the image feature F 1 (the second output data) output by the feature extraction model, the image feature F 1 is input into the main target segmentation student model to obtain the main target prediction information S s , that is, the first output data. The main target prediction information S s ∈R (p×k)×1×M×N . The value range of each element in the high-dimensional matrix is [0 - 1]. The closer the value is to 1, the area is the main target; the smaller the value and the closer it is to 0, it indicates that the area is background information.

[0081] After that, in this embodiment, an image feature pool F p ∈R L×d is set. There are a total of L features in this image feature pool, each feature dimension is d, and the image features in the image feature pool are randomly initialized learnable parameters. The main target prediction information S s and the image feature F 1 are jointly processed with the image feature pool to obtain the enhanced feature F 3 . Dimensionality compression processing is performed on the enhanced image feature F 3 . The global feature F g and the enhanced image feature F 3 are subjected to a feature position attention operation to generate the position weight W ij . Among them, the global feature F g is obtained by first performing average pooling and max pooling on the image feature F 1 of the training image samples. The position weight calculation formula is as follows:

[0082]

[0083] where cos represents the normalized cosine distance, and the value range is [0 - 1], M, N are the resolutions of the target image features; i, j are the position coordinates of the target image features;

[0084] After obtaining the position weight, the position weight Wij With the enhanced image feature F 3 Perform weighted summation to obtain the compressed image feature F d , and the specific formula for position weighting is as follows:

[0085]

[0086] Calculate the triplet loss L 3 of the enhanced image feature F tri3 , to obtain the first extraction accurate information, which can be specifically represented based on the loss of classification or clustering. For example, it is represented according to the above triplet loss and classification loss, so that the image recognition completed by the image features extracted based on the feature extraction model can cluster the samples of the same category together in the feature space and separate the samples of different categories. Calculate the loss L c3 of the classification loss function of the compressed image feature, measure the difference between the category predicted by the model and the true category, separate the positive and negative samples, and punish classification errors to a certain extent. Based on the loss value L tri3 of the triplet loss function and the loss value L c3 of the classification loss function, jointly update the feature extraction model A in the training with the above depth estimation loss and main target segmentation loss.

[0087] In one embodiment, according to the first output data, the second output data, and a preset image feature pool, perform feature enhancement processing on the second output data to obtain enhanced features, including: weighting the first output data and the second output data to obtain a weighted feature; calculating the similarity between the weighted feature and the reference feature in the preset feature pool; obtaining the reference feature with the highest similarity to the weighted feature in the image feature pool as the feature to be spliced; splicing the feature to be spliced and the weighted feature to obtain enhanced features.

[0088] To further reduce the interference of surrounding background factors, perform position weighting on the main target prediction information S s (the first output data) and the image feature F 1 (the second output data) to obtain the weighted feature F 2 . The calculation formula of the weighted feature F 2 is as follows:

[0089] F 2 =S s ·F 1 .

[0090] Calculate the image feature at each position in the weighted feature F 2 and the preset feature pool F pBased on the similarity of the reference features in it, the feature with the highest similarity to this position in the image feature pool is taken out as the feature to be stitched, and is stitched with the image feature at this position in the target image feature to obtain the enhanced image feature F 3 , and its specific calculation formula is as follows:

[0091]

[0092]

[0093] Among them, "cat" represents the stitching operation, and maxcos represents the calculation of the maximum cosine similarity. By calculating the similarity between the weighted feature and the reference feature in the feature pool, the reference feature in the feature pool with the highest similarity and the weighted feature are stitched, and feature matching regularization is used to improve the generalization of the features.

[0094] In some of the embodiments, during the process of training the feature extraction model, it further includes: calculating the usage frequency of each reference feature in the image feature pool; in the image feature pool, replacing the reference features with usage frequencies less than the preset frequency threshold based on the weighted feature.

[0095] Specifically, after strengthening the weighted feature F 2 using the feature pool, in order to make the feature training more stable and ensure that the feature pool does not collapse during the training process, this embodiment records the usage of each image feature in the feature pool and constructs an image feature usage record loop table Rec p ∈R L×T , which records the usage of each image feature in the most recent T iterations. In the initial state, all the element values in Rec p are 0. Among them, when the M-th iteration is performed, the cosine similarity between the image feature at each position in the weighted feature F 2 and the image features in the feature pool is calculated, and the image feature F pk with the maximum similarity is taken. Then Rec p [M % L, k] = 1, where "%" represents the modulo operation, L is the number of features in the image feature pool, and k is the number of samples for each category in the image to be recognized.

[0096] After each round of parameter update, calculate the usage frequency P rec of each image feature:

[0097]

[0098] Among them, sum 2 represents summation in the second dimension; that is, after calculation, sum 2 (Rec p) ∈ R L , p rec Among the L positions of, for each position i, the element represents the usage frequency of the i-th image feature in the feature pool. When the usage frequency of this feature is less than the threshold η, use F 2 The element at a random position in will replace it. The threshold η of the usage frequency can be set according to the actual situation, and this embodiment does not make specific limitations on this.

[0099] In another embodiment, during the training of the feature extraction model, it further includes: updating the feature extraction model according to the degree of approximation between the second output data and the weighted feature.

[0100] Specifically, in order to make the image feature F 1 and the weighted feature F 2 feature similar, an inference optimization loss function is also set in this embodiment. This inference optimization loss function is used to calculate the loss L 2 between the image feature F1 and the weighted feature F inf , and its specific calculation formula is as follows:

[0101] L inf = ||F 1 - F 2 || 2 ;

[0102] The parameter of the feature extraction model A is updated through the loss value L inf of this inference optimization loss function to obtain the trained feature extraction model.

[0103] In some of these embodiments, during the training of the feature extraction model, it further includes: performing mean and max pooling processing on the second output data to obtain global features; performing weighted summation processing on the second output data according to the similarity between the global features and the second output data to obtain compressed features; determining the second extraction accuracy information of the feature extraction model during training according to the compressed features; and updating the feature extraction model according to the second extraction accuracy information.

[0104] This second extraction accuracy information is used to characterize the accuracy of the features extracted by the feature extraction model, and can also be represented based on the loss of classification or clustering. Specifically, Figure 3 is the feature compression flow chart during the training of the feature extraction model in this embodiment. As Figure 3 shown, first perform mean pooling and max pooling processing on the image feature F 1 (second output data) of the training image sample to obtain global features F g , and combine the global features F g with the image feature F1 Perform feature location attention operation to generate location weight W ij , and the calculation formula of the location weight is as follows:

[0105]

[0106] where cos represents the normalized cosine distance, and the value range is [0-1], M and N are the resolutions of the image feature F 1 ; i and j are the position coordinates of the image feature F 1 ;

[0107] After obtaining the location weight, perform weighted summation on the location weight W ij and the image feature F of the training image sample 1 to obtain the compressed feature F c , and the specific formula for location weighting is as follows:

[0108]

[0109] Calculate the triplet loss L c of the compressed feature F trj , so that the image recognition model clusters samples of the same category together in the feature space and separates samples of different categories. Calculate the classification loss L c of the compressed image feature, measure the difference between the category predicted by the model and the true category, separate positive and negative samples, and punish classification errors to a certain extent. The triplet loss L tri and the classification loss L c constitute the second extraction of accurate information. Based on the triplet loss L tri and the classification loss L c , update the feature extraction model A.

[0110] In some embodiments, based on the target image feature, obtain the image recognition result, including: calculate the similarity between the target image feature and the standard feature to obtain the target similarity; the standard feature is the image feature extracted in advance from the standard image containing the target recognition object; according to the target similarity, obtain the image recognition result for the target recognition object.

[0111] Specifically, take the image feature extracted from the standard image containing the target recognition object as the standard feature, calculate the similarity between the target image feature and the standard feature to obtain the target similarity, and identify the target recognition object corresponding to the standard feature with the highest similarity as the image recognition result of the object to be recognized. Exemplarily, in an autonomous driving system, by extracting the features of the scene image and calculating the similarity with the standard scene features, the recognition and analysis of roads, traffic signs, etc. are realized.

[0112] In some of these embodiments, the total loss L can be calculated based on the depth estimation loss L p in the above embodiments, the main target segmentation loss L s the classification loss L c the classification loss L c3 the triplet loss L tri the triplet loss L tri3 and the inference optimization loss L inf to obtain the total loss L:

[0113] L = L tri + L c + L tri3 + L c3 + L inf + L s + L p ;

[0114] Update the feature extraction model A according to the total loss L.

[0115] Figure 4 is a training schematic diagram of the feature extraction model in some of these embodiments. As Figure 4 shown, the training image samples are used as Figure 4 input data in, and are respectively input into the feature extraction model, the main target segmentation teacher model, and the depth estimation teacher model. The output data F 1 of the feature extraction model are respectively input into the depth estimation student model and the main target segmentation student model. According to the output data of the depth estimation teacher model, the output data of the depth estimation student model, and in combination with the depth estimation loss function, calculate the depth estimation loss. According to the output data of the main target segmentation teacher model, the output data of the main target segmentation student model, and in combination with the main target segmentation loss function, calculate the main target segmentation loss. Additionally, the output data F 1 of the feature extraction model is also weighted based on the processing of the main target segmentation student model and the output result of the main target segmentation student model to obtain the weighted feature F 2 . Based on the stitching feature with the highest similarity to this weighted feature F 2 in the image feature pool, it is stitched with this weighted feature to form a strengthened feature. After compressing the strengthened feature, calculate the triplet loss (classification loss may also be included) for the compressed strengthened feature. Synchronously update the main target segmentation student model, the feature extraction model, and the depth estimation learning model based on various losses.

[0116] Figure 5 is a data flow schematic diagram of the feature extraction model in some of these embodiments. As Figure 5 shown, the input data that needs to be feature-extracted is input into the feature extraction model to achieve feature extraction and obtain the image feature F1 .

[0117] Figure 6 A flowchart of a training method for a feature extraction model of some embodiments is shown as Figure 6 shown, and the training method includes the following steps:

[0118] Step S601, constructing a first knowledge distillation structure and a second knowledge distillation structure; wherein, the first knowledge distillation structure includes a depth estimation teacher model and a depth estimation student model; the second knowledge distillation structure includes a main target segmentation teacher model and a main target segmentation student model.

[0119] Step S602, inputting the input data for model training into the feature extraction model being trained, the depth estimation teacher model, and the main target segmentation teacher model simultaneously, and inputting the output data of the feature extraction model being trained into the depth estimation student model and the main target segmentation student model.

[0120] Step S603, using the output data of the main target segmentation student model as the first output data and the output data of the feature extraction model being trained as the second output data, and performing feature enhancement processing on the second output data according to the first output data, the second output data, and a preset image feature pool to obtain enhanced features.

[0121] Step S604, performing dimensionality compression processing on the enhanced features to obtain compressed enhanced features.

[0122] Step S605, determining the current corresponding first extraction accuracy information of the feature extraction model being trained based on the compressed enhanced features, and updating the feature extraction model according to the first extraction accuracy information.

[0123] Step S606, calculating the usage frequency of each reference feature in the image feature pool.

[0124] Step S607, in the image feature pool, replacing the reference features with a usage frequency less than a preset frequency threshold based on the weighted features.

[0125] Step S608, updating the feature extraction model according to the proximity between the second output data and the weighted features.

[0126] Step S609, performing mean and max pooling processing on the second output data to obtain global features; performing weighted summation processing on the second output data according to the similarity between the global features and the second output data to obtain compressed features; determining the second extraction accuracy information of the feature extraction model being trained according to the compressed features; and updating the feature extraction model according to the second extraction accuracy information. Among them, there is no fixed execution order among S605, S607, S608, and S609.

[0127] Based on the above steps S601 to S609, through the training of depth estimation and main target segmentation auxiliary features, focus on the main target, suppress background activation, and at the same time introduce the feature pool technology, use feature matching regularization to improve feature generalization. Finally, a feature compression structure is proposed, using spatial attention to better extract the main target information in the image, improve feature generalization, and ultimately improve the model recognition effect.

[0128] In this embodiment, an image recognition device is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated here. The following terms "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0129] Figure 7 is the structural block diagram of the image recognition device of this embodiment, as Figure 7 shown, the device 70 includes: an acquisition module 71, a feature extraction module 72, and an identification module 73, where,

[0130] The acquisition module 71 is used to acquire the image to be recognized;

[0131] The feature extraction module 72 is used to perform image feature extraction on the image to be recognized according to a pre-trained feature extraction model to obtain target image features; among them, the pre-trained feature extraction model is jointly trained with a pre-constructed first knowledge distillation structure and a second knowledge distillation structure; the first knowledge distillation structure includes a depth estimation teacher model and a depth estimation student model; the second knowledge distillation structure includes a main target segmentation teacher model and a main target segmentation student model; the identification module 73 is used to obtain an image recognition result based on the target image features.

[0132] It should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combination form.

[0133] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and alternative implementation manners, and will not be repeated in this embodiment.

[0134] In addition, in combination with the image recognition method provided in the above embodiments, a storage medium can also be provided in this embodiment to implement it. A computer program is stored on the storage medium; when the computer program is executed by a processor, any one of the image recognition methods in the above embodiments is implemented.

[0135] It should be understood that the specific embodiments described here are only used to explain this application, rather than to limit it. According to the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of this application.

[0136] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0137] Obviously, the accompanying drawings are only some examples or embodiments of this application. For those of ordinary skill in the art, this application can also be applied to other similar situations based on these drawings without creative work. Additionally, it can be understood that although the work done during this development process may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient disclosure of this application.

[0138] The term "embodiment" in this application means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various positions in the specification and does not necessarily mean the same embodiment, nor does it mean being mutually exclusive with other embodiments and having independence or being optional. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in this application can be combined with other embodiments without conflict.

[0139] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0140] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. An image recognition method, characterized in that: include: Obtain an image to be recognized; Image features of the image to be identified are extracted according to a pre-trained feature extraction model to obtain target image features; wherein the pre-trained feature extraction model is jointly trained with a pre-constructed first knowledge distillation structure and a second knowledge distillation structure; the first knowledge distillation structure includes a depth estimation teacher model and a depth estimation student model; the second knowledge distillation structure includes a main target segmentation teacher model and a main target segmentation student model; Based on the target image features, an image recognition result is obtained.

2. The image recognition method according to claim 1, characterized in that: During the joint training process: Based on the output data of the depth estimation teacher model and the output data of the depth estimation student model, synchronously updating the feature extraction model and the depth estimation student model in training; Based on the output data of the main target segmentation teacher model and the output data of the main target segmentation student model, the feature extraction model and the main target segmentation student model in training are synchronously updated.

3. The image recognition method according to claim 2, characterized in that: During the joint training process: Inputting input data for model training into the feature extraction model, the depth estimation teacher model and the main target segmentation teacher model in training at the same time, and inputting output data of the feature extraction model in training into the depth estimation student model and the main target segmentation student model; Taking the output data of the depth estimation teacher model as first reference information, and based on the first reference information and the output data of the depth estimation student model, synchronously updating the feature extraction model and the depth estimation student model in training; The output data of the main target segmentation teacher model is used as the second reference information, and based on the second reference information and the output data of the main target segmentation student model, the feature extraction model and the main target segmentation student model in training are synchronously updated.

4. The image recognition method according to any one of claims 1 to 3, characterized in that: In the process of training the image feature extraction model, the method further includes: The output data of the main target segmentation student model is used as the first output data, and the output data of the feature extraction model in training is used as the second output data. According to the first output data and the second output data, and a preset image feature pool, the second output data is subjected to feature enhancement processing to obtain enhanced features; Performing dimension compression processing on the enhanced features to obtain compressed enhanced features; Based on the compressed enhanced features, the first extraction accurate information currently corresponding to the feature extraction model in training is determined, and the feature extraction model is updated according to the first extraction accurate information.

5. The image recognition method according to claim 4, characterized in that: The step of performing feature enhancement processing on the second output data according to the first output data and the second output data and a preset image feature pool to obtain enhanced features includes: Weighting the first output data and the second output data to obtain a weighted feature; Calculating the similarity between the weighted feature and a reference feature in a preset feature pool; Obtaining a reference feature with the highest similarity to the weighted feature in the image feature pool as a feature to be spliced; The feature to be spliced ​​is spliced ​​with the weighted feature to obtain a strengthened feature.

6. The image recognition method according to claim 5, characterized in that: In the process of training the feature extraction model, the method further includes: Calculating the usage frequency of each reference feature in the image feature pool; In the image feature pool, the reference features whose usage frequency is less than a preset frequency threshold are replaced based on the weighted features.

7. The image recognition method according to claim 5, characterized in that: In the process of training the feature extraction model, the method further includes: The feature extraction model is updated according to the degree of proximity between the second output data and the weighted feature.

8. The image recognition method according to claim 4, characterized in that: In the process of training the feature extraction model, the method further includes: Performing mean and maximum pooling processing on the second output data to obtain global features; According to the similarity between the global feature and the second output data, performing weighted sum processing on the second output data to obtain a compressed feature; Determining second extraction accuracy information of the feature extraction model in training according to the compression feature; The feature extraction model is updated according to the second extraction accuracy information.

9. The image recognition method according to claim 1, characterized in that: The obtaining of an image recognition result based on the target image feature includes: Calculating the similarity between the target image feature and the standard feature to obtain the target similarity; the standard feature is an image feature extracted in advance from a standard image containing the target recognition object; The image recognition result for the target recognition object is obtained according to the target similarity.

10. An image recognition device, characterized in that: include: Acquisition module, feature extraction module and recognition module, among which, The acquisition module is used to acquire the image to be identified; The feature extraction module is used to extract image features of the image to be identified according to a pre-trained feature extraction model to obtain target image features; wherein the pre-trained feature extraction model is jointly trained with a pre-constructed first knowledge distillation structure and a second knowledge distillation structure; the first knowledge distillation structure includes a depth estimation teacher model and a depth estimation student model; the second knowledge distillation structure includes a main target segmentation teacher model and a main target segmentation student model; The recognition module is used to obtain an image recognition result based on the target image features.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image recognition method according to any one of claims 1 to 8 are implemented.