A target recognition method and device

By identifying discriminative regions in an image, generating occluded and unoccluded feature images, and utilizing an attention model and dilated separable convolution, the accuracy of target recognition is improved, solving the problem of image target recognition accuracy in complex environments.

CN114429561BActive Publication Date: 2026-01-30HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202011479454.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-14
Filing Date
2020-12-15
Publication Date
2026-01-30
Estimated Expiration
2040-12-15

AI Technical Summary

Technical Problem

Existing technologies suffer from low accuracy in target recognition in complex environments due to factors such as occlusion and lighting, making it difficult to accurately obtain target information.

Method used

By identifying the discriminative regions of an image, two feature images—one occluded and one unoccluded—are generated. An attention model is used to configure scores, and dilated separable convolution and feature image fusion are employed to improve the accuracy of target recognition.

Benefits of technology

To improve the accuracy of target recognition under occlusion and environmental influences, reduce the impact of adverse factors, and enhance the recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114429561B_ABST
    Figure CN114429561B_ABST
Patent Text Reader

Abstract

A target recognition method and apparatus are disclosed in this application. The target recognition apparatus acquires an input image, which includes a target to be recognized. The apparatus first determines a discriminative region of the input image, which is a subset of regions in the input image that indicate the category to which the target belongs. After determining the discriminative region, a first feature image is obtained by occluding the discriminative region in the input image. Subsequently, the target recognition apparatus performs target recognition based on the first feature image to determine the category to which the target in the input image belongs. The target recognition apparatus considers the situation where the discriminative region in the input image is occluded. By acquiring the first feature image in this case, the apparatus analyzes the regions in the first feature image other than the discriminative region to perform target recognition. By strengthening the analysis of the regions outside the discriminative region during the target recognition process, the accuracy of target recognition can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing

[0002] This application claims priority to Chinese Patent Application No. 202011097641.5, filed on October 14, 2020, entitled "An Image Recognition Method and Apparatus", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of communication technology, and in particular to a target recognition method and apparatus. Background Technology

[0004] Currently, image-based target recognition has a wide range of applications. For example, it is involved in scenarios such as traffic violation management, commodity recognition, endangered species protection, traffic monitoring and investigation.

[0005] Taking the management of vehicles violating traffic rules as an example, image-based target recognition can identify vehicles violating traffic rules by using images of roads or other locations, and obtain vehicle information such as license plates and vehicle logos.

[0006] However, due to the complexity of real-world situations, such as the lighting at the time the image was taken, road conditions, and the influence of the environment around the offending vehicle, the offending vehicle may be obscured in the image, making it impossible to perform accurate target recognition based on the image, that is, to accurately identify the offending vehicle and obtain accurate vehicle information. Summary of the Invention

[0007] This application provides a target recognition method and apparatus to improve the accuracy of target recognition.

[0008] Firstly, embodiments of this application provide a target recognition method, which can be executed by a target recognition device. In this method, the target recognition device acquires an input image containing a target to be recognized. To recognize the target, the target recognition device first determines a discriminative region of the input image, whereby the discriminative region is a subset of regions in the input image that indicate the category to which the target belongs. After determining the discriminative region, the target recognition device can obtain a first feature image by occluding the discriminative region in the input image; that is, the first feature image is an image in which the discriminative region of the input image is occluded. Subsequently, the target recognition device can recognize the target based on the first feature image.

[0009] By using the above method, the target recognition device takes into account the situation where the distinguishable region in the input image is occluded when performing target recognition. It obtains the first feature image under such circumstances and analyzes the region other than the distinguishable region in the first feature image to determine the category to which the target belongs. In this target recognition process, by strengthening the analysis of the region other than the distinguishable region, the target recognition effect is achieved, thus ensuring the accuracy of target recognition.

[0010] In one possible implementation, the target recognition device can generate a second feature image by either occluding the discriminative regions of the input image or by not occluding them, such as by displaying the discriminative regions normally or by highlighting them. In other words, the second feature image is a feature image of the input image where the discriminative regions are not occluded; when performing target recognition, the target recognition device can identify the target based on the first and second feature images.

[0011] By using the above method, the target recognition device considers both the case where the distinguishable region is occluded and the case where the distinguishable region is not occluded, and obtains a first feature image and a second feature image, which correspond to the case where the input image may be occluded (or the image cannot present the true situation due to the influence of the shooting environment) and the case where the input image is not occluded (or the input image can present the true situation). Target recognition based on the first feature image and the second feature image can reduce the influence of occlusion or environment, thereby improving the accuracy of target recognition.

[0012] In one possible implementation, the target recognition device can determine the discriminative region based on the spatial features of the input image. For example, regions in the input image whose spatial features are greater than a threshold or fall within a certain range can be selected as discriminative regions. This application does not limit the method of determining the discriminative region based on the spatial features of the input image.

[0013] Using the above method, spatial features of the input image can be used to determine more discriminative regions that can better represent the target's category from other targets.

[0014] In one possible implementation, when the target recognition device determines the discriminative region based on the spatial features of the input image, it can assign scores to the spatial features of the input image; for example, an attention model can be used to assign scores to the spatial features of the input image, and regions with spatial feature scores greater than a threshold can be used as discriminative regions.

[0015] Using the above methods, the higher the score of the spatial feature, the more information it contains, and the more discriminative the identified region is.

[0016] In one possible implementation, when the target recognition device generates a first feature image based on the discriminative regions of the input image, it can configure first coefficient values ​​for the pixels in the input image. For example, the first coefficient values ​​of the pixels belonging to the discriminative regions in the input image can be configured to a smaller first value, and the first coefficient values ​​of the remaining pixels can be configured to a larger second value. The graph formed by the first coefficient values ​​of each pixel is the first coefficient graph. Then, the first coefficient graph is applied to the input image (in specific applications, it can be applied to the feature image of the input image or the processed feature image of the input image, such as Fout or B in the embodiment) to generate the first feature image.

[0017] By applying the coefficient map to the input image using the above method, the pixel values ​​of each pixel in the discriminative region of the input image can be reduced, thereby occluding the discriminative region and making it easier to obtain the first feature image.

[0018] In one possible implementation, the target recognition device generates a second feature image based on the discriminative regions of the input image. It can configure second coefficient values ​​for pixels in the input image. For example, the second coefficient values ​​of pixels belonging to the discriminative regions can be configured as larger first values, while the coefficients of other pixels can be configured as smaller second values. Alternatively, the second coefficient values ​​of pixels belonging to the discriminative regions can be configured as scores representing the spatial features of the pixels. The graph formed by the second coefficient values ​​of all pixels is called the second coefficient graph. Then, the second coefficient graph is applied to the input image (in specific applications, it can be applied to the feature image of the input image or the feature image of a processed input image, such as Fout or B in the embodiment) to generate the second feature image.

[0019] Using the above method, the target recognition device can change the pixel values ​​of each region in the input image in a variety of different ways to highlight the distinguishable region, and thus obtain the second feature image.

[0020] In one possible implementation, when the target recognition device identifies a target based on a first feature image and a second feature image, it can aggregate and reduce the dimensionality of the first and second feature images in the channel dimension to generate a third feature image. Then, based on the third feature image, it determines multiple candidate feature images with different receptive fields, wherein each candidate feature image is the same size. Then, it fuses the multiple candidate feature images into a fourth feature image. The target is then identified based on the fourth feature image.

[0021] Using the above method, the fourth feature image is formed by fusing multiple candidate feature images with different receptive fields. In this way, the receptive field of the obtained fourth feature image can cover more effective information that is conducive to target recognition and reduce invalid information that is not conducive to target recognition, so that the target recognition device can recognize the target more accurately through the fourth feature image.

[0022] In one possible implementation, when the target recognition device aggregates the first feature image and the second feature image in the channel dimension to generate the third feature image, it can first aggregate and reduce the dimensionality of the first feature image and the second feature image in the channel dimension to generate an aggregated image. This aggregated image can be the same size as the first feature image or the second feature image. Then, weights are assigned to the aggregated image in the channel dimension to generate the third feature image. The weights assigned to the candidate feature images can achieve the following effects: when the discriminative region is occluded in the input image, the weight of the part of the aggregated image belonging to the first feature image in the channel is greater than the weight of the part of the aggregated image belonging to the second feature image in the channel; or when the discriminative region is not occluded in the input image, the weight of the part of the aggregated image belonging to the first feature image in the channel is less than the weight of the part of the aggregated image belonging to the second feature image in the channel.

[0023] By using the above method and the weight configuration in the channel dimension, the part of the aggregated image belonging to the second feature image can be highlighted when the discriminative region is occluded, and the part of the aggregated image belonging to the first feature image can be highlighted when the discriminative region is not occluded. This makes the weight of the third feature image in the channel dimension more consistent with the state of the discriminative region in the input image being occluded or not.

[0024] In one possible implementation, the target recognition device determines multiple candidate feature images based on a third feature image. Multiple different convolution kernels can be applied to the third feature image respectively, and multiple candidate feature images can be obtained by dilating the convolution (that is, by padding the third feature image with 0).

[0025] Using the above method, the target recognition device obtains multiple candidate feature images of the same size by dilated separable convolution, which facilitates the subsequent fusion of these multiple candidate feature images.

[0026] In one possible implementation, when the target recognition device fuses multiple candidate feature images into a fourth feature image, it can configure a weight for each candidate feature image. This weight can be obtained through pre-learning and training. Then, based on each candidate feature image and the weight corresponding to each candidate feature image, the fourth feature image is obtained.

[0027] By configuring corresponding weights for each candidate feature image using the above method, when the fourth feature image is subsequently fused and generated, information from each candidate feature image can be retained with different emphases, so that the receptive field of the fourth feature image can cover more effective information that is beneficial to target recognition.

[0028] Secondly, embodiments of this application also provide a target recognition device, which has the function of implementing the behavior in the method example of the first aspect described above. The beneficial effects can be found in the description of the first aspect and will not be repeated here. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. In one possible design, the device structure includes an acquisition unit, an image generation unit, a recognition unit, and a determination unit. These units can perform the corresponding functions in the method example of the first aspect described above, as detailed in the method example description, and will not be repeated here.

[0029] Thirdly, embodiments of this application also provide an apparatus that performs the functions described in the method example of the first aspect above. The beneficial effects are described in the first aspect and will not be repeated here. The apparatus includes a processor and a memory. The processor is configured to support the target identification device in performing the corresponding functions of the method in the first aspect above. The memory is coupled to the processor and stores necessary program instructions and data for the communication device. The communication device also includes a communication interface for communicating with other devices.

[0030] Fourthly, this application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the methods described in the first aspect and various possible implementations of the first aspect.

[0031] Fifthly, this application also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods described in the first aspect and various possible implementations of the first aspect.

[0032] In a sixth aspect, this application also provides a computer chip connected to a memory, the chip being used to read and execute a software program stored in the memory, and to execute the methods described in the first aspect and various possible implementations of the first aspect. Attached Figure Description

[0033] Figure 1 A schematic diagram of a feature image provided in this application;

[0034] Figure 2A schematic diagram of the system architecture provided in this application;

[0035] Figure 3A A schematic diagram of a target recognition method provided in this application;

[0036] Figure 3B A schematic diagram of another target recognition method provided in this application;

[0037] Figure 4 A schematic diagram illustrating the method for determining the distinguishable region provided in this application;

[0038] Figure 5A This application provides a schematic diagram of a method for assigning scores to spatial features using an attention model.

[0039] Figure 5B A schematic diagram illustrating a method for assigning scores to spatial and temporal features using an attention model, as provided in this application;

[0040] Figure 6A A schematic diagram illustrating a method for generating a first feature image provided in this application;

[0041] Figure 6B A schematic diagram illustrating the effect of a first feature image provided in this application;

[0042] Figure 7A A schematic diagram illustrating a method for generating a second feature image provided in this application;

[0043] Figure 7B A schematic diagram illustrating the effect of a second feature image provided in this application;

[0044] Figure 8 A schematic diagram illustrating the conversion of a third feature image into a fourth feature image, as provided in this application;

[0045] Figure 9 A schematic diagram illustrating a method for generating a sixth feature image provided in this application;

[0046] Figure 10A A schematic diagram of the ResNet50 structure provided in this application;

[0047] Figure 10B A schematic diagram of a CNN structure is provided for this application;

[0048] Figure 11 A schematic diagram of the structure of a target recognition device provided in this application;

[0049] Figure 12 This is a schematic diagram of a device provided in this application. Detailed Implementation

[0050] Before describing the target recognition method and device provided in the embodiments of this application, some concepts designed in the embodiments of this application will be explained first:

[0051] 1. Image features, feature images

[0052] Image features are used to characterize the attributes of an image. There are many types of image features, which can be divided into spatial features and visual features. Different image features can represent an image from different perspectives. Image features can be quantified into numerical values, which are called feature values.

[0053] Different regions of an image have different image features, meaning that different regions of an image correspond to different feature values. An image composed of the feature values ​​corresponding to each region of the image is called a feature image.

[0054] 2. Channel dimension, spatial dimension, visual features, and spatial features

[0055] like Figure 1 As shown, a feature image can be abstracted into a cube in space with length, width, and height C, H, and W respectively, where the direction of C is the channel dimension, and the planes containing W and H are the spatial dimensions.

[0056] from Figure 1 As can be seen from the data, the feature image has a length of C in the channel dimension, which can be understood as the feature image having C channels. Visual features describe features in the channel dimension; one channel corresponds to one visual feature. The number of channels may vary for different feature images. There are many types of visual features, such as color features and texture features.

[0057] This feature image displays the distances or relationships between people or objects in the image in a spatial dimension. Spatial feature description refers to features in the spatial dimension. A feature value in the spatial feature corresponds to a region (composed of multiple pixels) in the image and is used to describe the characteristics of that region in the spatial dimension.

[0058] 3. Distinguishing regions

[0059] Discriminative regions are used to distinguish different categories of targets. They are subsets of regions that can represent differences from the categories to which other targets belong. In other words, there are many regions in an image that can represent differences from the categories to which other targets belong; a subset of these regions can be selected as discriminative regions. For example, in an image of a vehicle, the areas containing the front of the car, the logo, the rear lights, and the tires are all regions that can represent differences from the categories to which other targets belong. Therefore, the areas containing the front of the car, the logo, and the rear lights can be considered discriminative regions.

[0060] There are many methods for determining the discriminative region. For example, the attention model provided in the embodiments of this application can be used to determine the discriminative region. Alternatively, each pixel in the feature image can be clustered, and the region where the value of each pixel is greater than a set value can be used as the discriminative region. Or, the region where the value of each pixel is within a preset range can be used as the discriminative region.

[0061] The discriminative regions in this application are suitable for coarse-grained target recognition (i.e., identifying the broad category to which the target belongs, such as the identification of plants, animals, and people), and also for fine-grained target recognition (i.e., identifying the specific subcategory to which the target belongs, such as identifying the categories to which different birds belong). Taking a fine-grained target recognition scenario as an example, a discriminative region is a region that can distinguish the category to which a target in an image belongs from the same broad category. Simply put, for example, distinguishing different birds (parrots, sparrows, orioles, etc.) within the broad category of birds, many birds within the same broad category are extremely similar in appearance and size, with only extremely minor differences. These differences mostly exist in areas such as the bird's beak, claws, feather color, eyes, and tail, and these areas are called discriminative regions, which are the areas that can distinguish the bird.

[0062] 4. Aggregation and Dimensionality Reduction

[0063] In the embodiments of this application, multiple feature images can be aggregated in the channel dimension. Aggregation in the channel dimension refers to superimposing two feature images in the channel dimension.

[0064] In this embodiment, dimensionality reduction can also be performed on the feature image along the channel dimension. Dimensionality reduction along the channel dimension refers to reducing the length of the feature image along the channel dimension so that the length of the reduced feature image along the channel dimension meets specific requirements. In this embodiment, when multiple feature images are aggregated along the channel dimension, the length of the aggregated image along the channel dimension is equal to the sum of the lengths of the multiple feature images along the channel dimension. To ensure that the length of the output image along the channel dimension is consistent with that of the feature images before aggregation, dimensionality reduction can be performed again after aggregating multiple feature images along the channel dimension to obtain a feature image of the same size as the multiple feature images.

[0065] 5. Receptive field

[0066] The receptive field refers to the size of the region on the original image that a pixel in the feature image maps to.

[0067] 6. Dilated Separating Convolution, Dilation Rate

[0068] Dilated separable convolution is a combination of dilated convolution and depthwise separable convolution. Dilated convolution, also known as atted convolution or dilated convolution, injects holes (filled with zeros) into the standard convolution kernel to increase the receptive field. Compared to the original normal convolution operation, dilated convolution has an additional parameter: the dilation rate (or simply rate). The dilation rate refers to the number of zeros between points in the convolution kernel. Dilated convolution not only increases the size of the convolution kernel by a certain dilation rate but also fills the feature map by padding with zero values, so that the image after convolution has the same size as the image before convolution but has a larger receptive field.

[0069] Depthwise separable convolution is a lightweight convolutional operation that can extract information from both channels and spatially. Compared to standard convolution, depthwise separable convolution has a much lower parameter count and computational cost. It consists of two parts: depthwise convolution and pointwise convolution. Unlike standard convolution, in depthwise convolution, one kernel operates on one channel, with each channel having its own kernel. In pointwise convolution, each kernel effectively fuses information from multiple channels to generate a feature image. This allows multiple kernels to extract different features, resulting in multi-dimensional feature outputs. In summary, while depthwise convolution operates on each channel independently, it lacks interaction between information at the same spatial location across different channels. Therefore, pointwise convolution is needed to facilitate information exchange between different channels.

[0070] Dilated separable convolution applies dilated convolution to depthwise separable convolution. It first performs depthwise convolution on each channel and then uses pointwise convolution to fuse the information from each channel. This dilated separable convolution operation can reduce the amount of computation and parameters without reducing classification accuracy.

[0071] like Figure 2 The diagram shown is a system architecture diagram applicable to an embodiment of this application. The system includes an image collection device 200 and a target recognition device 100.

[0072] Image collection device 200 is used to collect images. After collecting the images, image collection device 200 feeds them back to target recognition device 100. The deployment location and type of image collection device 200 will vary in different application scenarios. For example, in a scenario of identifying illegal vehicles, image collection device 200 could be a camera device deployed on both sides of the road or a monitoring device deployed at a traffic intersection. Image collection device 200 can capture images of the road and send the captured images to target recognition device 100. As another example, in a scenario of species classification recognition, image collection device 200 could be a camera device deployed in a forest or ocean. Image collection device 200 can capture images of various plants and animals in the forest or ocean and send the captured images to target recognition device 100.

[0073] The target recognition device 100 can receive images from the image collection device 200 and execute the target recognition method provided in this application embodiment. This application embodiment does not limit the deployment location of the target recognition device 100. For example, the target recognition device 100 can be deployed in an edge data center, such as an edge computing node (multi-access edge computing, MEC) in the edge data center, or in a cloud data center, or on a terminal computing device. The target recognition device 100 can also be distributed and deployed in some or all of the environments of edge data centers, cloud data centers, and terminal computing devices.

[0074] The target identification device 100 can be a hardware device, such as a server, service cluster, or terminal computing device, or it can be a software device, specifically a software module running on a hardware computing device.

[0075] In this embodiment, the target recognition device 100 can perform both coarse-grained and fine-grained target recognition. For example, coarse-grained target recognition means the device 100 can perform simple classification of targets, identifying the broad categories to which they belong. For instance, the device 100 can identify humans, vehicles, animals, and plants in an image. Fine-grained target recognition means the device 100 can perform detailed classification of targets, identifying the specific subcategories to which they belong. For example, the device 100 can identify the model and brand of a vehicle in an image, or the species of different birds in an image.

[0076] In addition to receiving images from the image collection device 200, the target recognition device 100 can also receive data from other devices. Taking the scenario of identifying illegal vehicles as an example, the target recognition device 100 can also receive data from roadside units (RSUs) and radar measurements. The roadside unit (RSU) can identify vehicles passing through it and obtain their information. The RSU can then send this information to the target recognition device 100. After performing target recognition on the images sent by the image collection device, the target recognition device 100 can also label the targets in the images based on the vehicle information sent by the roadside unit. The radar can perform ranging, measuring the distance between vehicles and the distance from a vehicle to an object. The radar can send the measured information to an edge sensing unit, which then sends the information to the target recognition device. After performing target recognition on the images sent by the image collection device 200, the target recognition device 100 can also annotate the targets in the images with the information measured by the radar.

[0077] In addition to receiving data (such as data from image collection devices, roadside units, or radar), the target recognition device 100 sends the recognized information (such as information about the target, or an image labeled with the target and information from roadside units, radar, etc.) to other devices. For example, in the scenario of traffic violation measurement and recognition, the target recognition device can identify the vehicle violating the traffic rules in the input image, obtain the information of the vehicle violating the traffic rules, and send the information of the vehicle violating the traffic rules to the traffic control center system.

[0078] The target recognition method provided in this application will be described below with reference to the accompanying drawings. Figure 3A The method includes:

[0079] Step 101: The target recognition device 100 acquires an input image from the image collection device 200, the input image including the target to be recognized.

[0080] Step 102: The target recognition device 100 obtains a first feature image by occluding a distinguishable region in the input image. In other words, the first feature image is a feature image of the input image with the distinguishable region occluded. In this first feature image, the distinguishable region is hidden, and the feature values ​​of the pixels in the distinguishable region can be significantly smaller than the feature values ​​of pixels in other regions, such as zero feature values.

[0081] For methods of determining distinctive regions, please refer to... Figure 3B The relevant descriptions in step 202 of the illustrated embodiment, and the method by which the target recognition device 100 acquires the first feature image, can be found in the following example. Figure 3B The following is a description of steps 203 to 204 in the illustrated embodiment.

[0082] Step 103: The target recognition device 100 uses the first feature image to recognize the target in the input image.

[0083] The target recognition process performed by the target recognition device 100 can be seen in the following example: Figure 3B The following is a description of steps 207 to 211 in the illustrated embodiment.

[0084] To ensure accurate target recognition and obtain relevant information about the target in an image even when it is occluded or the image does not accurately reflect its true state due to environmental factors, this application embodiment addresses this issue. After acquiring the input image, a first feature image is generated using a discriminative region within the input image. This first feature image is a feature image of the occluded discriminative region. After acquiring this first feature image, the target is then identified based on it. As can be seen from the above process, this application embodiment considers the situation where the discriminative region is occluded, corresponding to situations where the image being recognized is occluded (or the image does not accurately reflect its true state due to environmental factors). Based on this first feature image, the target in the input image can be identified more accurately, reducing the impact of occlusion or environmental factors on target recognition and improving the accuracy of recognition.

[0085] It should be noted that the target recognition process provided in this application mainly involves the field of deep learning. The modules or neural networks used can be trained before use, that is, the modules or neural networks are first trained using a training set, and the parameters of the modules or neural networks are continuously adjusted so that the modules or neural networks can output more accurate results. After training is completed, the modules or neural networks can be put into use. The modules or neural networks can process the input data and output results, such as input feature images. However, the process of processing the input data during training and use is the same. The difference is that during training, the parameters of the modules or neural networks need to be adjusted according to the output results each time, while during use, the output is tended to be obtained through the modules or neural networks. The method described below uses use as an example to introduce the target recognition method provided in this application.

[0086] See Figure 3B To ensure the efficiency of target recognition, this application embodiment takes the recognition of targets using a first feature image and a second feature image as an example. The method specifically includes:

[0087] Step 201: The target recognition device 100 acquires the input image. The type of input image is not limited here. The input image can be an image directly sent after being acquired by the image collection device 200, or it can be an image processed based on the image acquired by the image collection device 200.

[0088] Step 202: The target recognition device 100 determines the distinguishable region in the input image.

[0089] Before calculating the discriminative regions in an image, feature extraction must first be performed on the image to obtain the feature image of the input image. This application does not limit the method by which the target recognition device 100 obtains the feature image of the input image. For example, the target recognition device 100 can use a ResNet50 neural network or a VGG16 network to obtain the feature image of the input image. The feature image of the input image can be the bottleneck layer of the ResNet50 neural network, or the feature images output by the mid-to-high layers of the VGG16 network, such as conv3_x, conv4_x, conv5_x, and conv6 of the VGG16 network.

[0090] This application does not limit the method by which the target recognition device 100 determines the distinguishable region; any method capable of determining the distinguishable region is applicable to the embodiments of this application. The following describes a method for determining the distinguishable region provided by an embodiment of this application, such as... Figure 4 As shown, the method includes:

[0091] Step 301: The target recognition device 100 determines the spatial features of the input image.

[0092] After the target recognition device 100 acquires the feature image of the input image, the value of each pixel in the feature image of the input image in the spatial dimension represents the spatial features of the input image, and the value of each pixel is the feature value.

[0093] Step 302: The target recognition device 100 can assign scores to spatial features based on an attention model.

[0094] Attention models can measure multiple pieces of information from a specific perspective and determine the value of each piece of information. In this embodiment, when the target recognition device 100 performs step 302, it can use the attention model to measure the spatial features of the feature image corresponding to the input image, determine the value of each spatial feature, such as determining the amount of information contained in the spatial feature, and assigning a score to each spatial feature. For example, a higher score is assigned to spatial features containing a lot of information, and a lower score is assigned to spatial features containing less information.

[0095] Optionally, in addition to assigning scores to spatial features, the target recognition device 100 can also assign scores to visual features. That is, it scores and assigns scores to each visual feature in the channel dimension.

[0096] The method by which the target recognition device 100 assigns scores to visual features is similar to the method by which it assigns scores to spatial features; it can also use an attention model to assign scores to visual features. The attention model used to assign scores to spatial features and the attention model used to assign scores to visual features are two independent attention models.

[0097] like Figure 5A The diagram shows a flowchart of the target recognition device 100 using an attention model to score spatial features and configure the scores.

[0098] Figure 5A In the process, for the feature image Fin of the input image, an attention model can be used to assign scores to the spatial features of the input image. These scores can be directly applied to the feature image of the input image, that is, the feature image B is obtained by multiplying the score with the corresponding feature value in the feature image. The feature image B and Fin have the same size, and the feature image B is the feature image after applying the spatial feature scores to Fin. The feature image B can be the feature image applied by the first coefficient map and the second coefficient map in steps 204 and 206.

[0099] like Figure 5B The diagram shows a flowchart of the target recognition device 100 using two independent attention models to score spatial features and visual features, and configuring the scores.

[0100] Figure 5B In this process, for the feature image Fin of the input image, two attention models can be used simultaneously to assign scores to the spatial and temporal features of the input image. The score assigned to the spatial features can be directly applied to the feature image of the input image, that is, the score is multiplied by the corresponding feature value in the feature image to obtain feature image A; the score assigned to the visual features of the input image can be directly applied to the input image to obtain feature image B. Then, feature images A and B are aggregated and dimensionality reduced to obtain feature image Fout. Fin and Fout have the same size, and Fout is the feature image after applying the scores of spatial features and visual features to Fin. This feature image Fout can be the feature image applied by the first coefficient map and the second coefficient map in steps 204 and 206.

[0101] It should be noted that the value of each visual or spatial feature lies in its contribution to target recognition (such as fine-grained or coarse-grained target recognition). Some visual or spatial features can directly indicate the attributes (such as category) of the target and contribute significantly to target recognition, while others do not highlight the attributes of the target and contribute less. In this embodiment, taking the determination of discriminative regions based on spatial features as an example, an attention model is used in the spatial dimension to enhance the expression of spatial features beneficial to target recognition and weaken the expression of spatial features with little impact on target recognition by configuring scores. This results in the acquisition of effective feature representations, thereby improving the accuracy of target recognition.

[0102] Step 303: The target recognition device 100 determines the discriminative region based on the spatial feature scores in the input image. This embodiment does not limit the manner in which the target recognition device 100 performs step 303. For example, the target recognition device 100 can use regions in the input image where the spatial feature scores are greater than a threshold as discriminative regions. This threshold can be an empirical value or a value determined through simulation or other methods. Alternatively, regions in the input image where the spatial feature scores are greater than a specific range can be used as discriminative regions. This specific range can be a fixed value or a manually set value.

[0103] After determining the discriminative regions in the input image, the first feature image (see steps 203-204) and the second feature image (see steps 205-206) can be determined respectively.

[0104] Step 203: The target recognition device 100 occludes the distinguishable region in the input image and assigns a first coefficient value to each pixel in the input image. The first coefficient values ​​of each pixel constitute a first coefficient map.

[0105] In order to mask the distinguishable regions in the input image, the pixel values ​​in the distinguishable regions of the input image can be configured with a lower first coefficient value, and the remaining pixels excluding the pixels in the distinguishable regions can be configured with a higher first coefficient value.

[0106] For example, for a pixel Fi within a distinguishable region, if pixel Fi needs to be occluded, the first coefficient value corresponding to that pixel can be configured to be zero. If the spatial feature score of that pixel is less than a threshold, the first coefficient value of that pixel is set to 1, that is:

[0107]

[0108] Among them, Att(F i ) represents the score of the spatial features of the pixel determined based on the attention model, and t is the threshold.

[0109] Step 204: The target recognition device 100 applies the first coefficient map to the feature image of the input image to obtain the first feature image. The target recognition device 100 may also apply the first coefficient map to... Figure 5A The feature image B or feature image Fout shown here is used as an example to illustrate the application of the first coefficient image to the feature image of the input image.

[0110] The target recognition device 100 multiplies the value of each pixel in the feature image of the input image with the first coefficient value of that pixel in the first coefficient image to obtain a first feature image. The size of the first feature image is C*H*W, where C is the length of the channel, H is the spatial height, and W is the spatial width.

[0111] like Figure 6A The flowchart shown is a process for the target recognition device 100 to generate a first feature image (wherein, the part of scoring the spatial features of the input image using an attention model and determining the discriminative region corresponds to steps 301-303, and the part of generating the first coefficient map corresponds to step 203). The target recognition device 100 uses an attention model to score the spatial features of the input image and configures the scores (for determining the discriminative region). Then, it generates a first coefficient map based on the scores of the spatial features of each pixel. After that, the first coefficient map is applied to the feature image of the input image to generate the first feature image.

[0112] like Figure 6B The image shown is a rendering of the first feature image. The distinguishable regions in the input image can be the vehicle's rearview mirror, headlights, and license plate. After these distinguishable regions are occluded, they turn black in the input image, while other regions are displayed normally.

[0113] Optionally, the target recognition device 100 may also acquire a second feature image without obscuring the distinguishable region in the input image. The second feature image is a feature image in which the distinguishable region in the input image is not obscured, and the distinguishable region can be displayed normally within this second feature image. The second feature image can be a feature image of the input image itself, meaning it does not obscure the distinguishable region within the feature image of the input image, allowing it to be displayed normally. In one possible implementation, to further highlight the difference between the distinguishable region and other regions, the second feature image can be a feature image in which the feature values ​​of the pixels in the distinguishable region are significantly higher than the feature values ​​of the pixels in other regions. The method by which the target recognition device 100 acquires the second feature image can be found in the relevant descriptions of steps 205 to 206.

[0114] Step 205: The target recognition device 100 does not obscure the distinguishable region in the input image, and configures a second coefficient value for each pixel in the input image. The second coefficient values ​​of each pixel constitute a second coefficient map.

[0115] The target recognition device 100's "not obscuring the distinguishable region in the input image" means keeping the distinguishable region clear. To ensure the distinguishable region can be displayed normally, or even highlighted, the target recognition device 100, when configuring second coefficient values ​​for each pixel in the input image, can configure higher second coefficient values ​​for pixels in the distinguishable region and lower second coefficient values ​​for pixels outside the distinguishable region.

[0116] For example, if a pixel is located within a discriminative region of the feature image of the input image and needs to be highlighted, the second coefficient value of that pixel can be configured to be 1. For pixels outside the discriminative region of the input image, the second coefficient value of that pixel can be configured to be 0.

[0117] For example, for each pixel of the feature image of the input image, the spatial feature score of each pixel can be normalized so that the spatial feature score of each pixel can be distributed in [0,1]. After normalization, the spatial feature score of each pixel can be used as the second coefficient value of each pixel to form the second coefficient map.

[0118] The spatial feature scores of each pixel are normalized using the following method:

[0119] MAP 突显(i) =σ(Att(F) i )),in,

[0120] Step 206: The target recognition device 100 applies the second coefficient map to the feature image of the input image to obtain the second feature image. Similar to step 204, the target recognition device 100 can also apply the second coefficient map to... Figure 5A The feature image B or feature image Fout shown here is used as an example to illustrate the application of the second coefficient image to the feature image of the input image.

[0121] The target recognition device 100 multiplies the value of each pixel in the input image with the second coefficient value of that pixel in the second coefficient image to obtain a second feature image. The size of the second feature image is C*H*W, where C is the length of the channel, H is the spatial height, and W is the spatial width.

[0122] like Figure 7AThe flowchart shown is a process for the target recognition device 100 to generate a second feature image (wherein, the part of scoring the spatial features of the input image using an attention model and determining the discriminative region corresponds to steps 301-303, and the part of generating the first coefficient map corresponds to step 205). The target recognition device 100 uses an attention model to score the visual features of the input image and configures the scores. Then, it generates a second coefficient map based on the scores of the spatial features of each pixel. After that, the second coefficient map is applied to the feature image of the input image to generate the second feature image.

[0123] like Figure 7B The image shown is a rendering of the second feature image. The distinguishable regions in the input image can be the vehicle's rearview mirror, headlights, and license plate. These distinguishable regions are not obscured. Furthermore, these distinguishable regions can be enhanced, with increased brightness in the input image and darker brightness in other areas.

[0124] After the above steps, the target recognition device 100 acquires the first feature image and the second feature image.

[0125] Step 207: The target recognition device 100 aggregates and reduces the dimensionality of the first feature image and the second feature image in the channel dimension to generate the third feature image.

[0126] Typically, the size of the first and second feature images aggregated along the channel dimension is 2C*H*W. To ensure that the size of the third feature image is consistent with the size of the first or second feature image, the aggregated image can be dimensionality reduced, i.e., compressed along the channel dimension, to generate the third feature image, making the length of the third feature image along the channel dimension C.

[0127] In other words, the third feature image is equivalent to the first feature image and the second feature image being stitched together in the channel dimension and then compressed to generate the feature image. The third feature image includes the first feature image and the second feature image. That is, in the channel dimension, the part belonging to the first feature image and the part belonging to the second feature image can be distinguished from the third feature image.

[0128] Step 208: The target recognition device 100 assigns weights to the third feature image in the channel dimension and generates a fourth feature image.

[0129] When the target recognition device 100 performs step 208, it can first configure weights for the third feature image in the channel dimension, that is, configure the weights to the channels of the third feature image to generate the fourth feature image.

[0130] The embodiments of this application are not limited to the way of configuring weights for the third feature image in the channel dimension. For example, the efficient channel attention (ECA) model based on channel relationship modeling can be used to configure weights for the third feature image in the channel dimension.

[0131] The ECA model is pre-trained and can be built based on the channel dimension. It can model the channel relationship and learn the connection between visual features, thereby obtaining a more efficient visual feature representation by configuring weights for the third feature image in the channel dimension.

[0132] At the start of training, the parameters of the ECA model are randomly initialized. As training progresses, the classifier feeds back the classification results to the ECA model, allowing it to adaptively adjust the weights based on these results. Visual features that contribute significantly to classification receive larger weights, while those contributing less receive smaller weights. This continuous learning and adjustment of weight allocation during training continues until a stable state is reached, resulting in a weight distribution most conducive to object recognition. In short, during ECA model training, based on the gradient descent algorithm, the model learns from the feature images in the training set. This allows the ECA model to redistribute the feature weights across all channels of the feature images. When a discriminative region is occluded, the weight of the occluded region is increased, enabling the classifier to learn the discriminative features of other regions. Conversely, the weight of the unoccluded region is increased to identify the features of the discriminative region, allowing the classifier to make the correct judgment.

[0133] The weights configured for the third feature image satisfy the following conditions: when the discriminative region is occluded in the input image, the weight of the part of the third feature image belonging to the first feature image in the channel is greater than the weight of the part of the third feature image belonging to the second feature image in the channel; or when the discriminative region is not occluded in the input image, the weight of the part of the third feature image belonging to the first feature image in the channel is less than the weight of the part of the third feature image belonging to the second feature image in the channel.

[0134] like Figure 8 The diagram shown illustrates how the ECA model is used to convert the third feature image into the fourth feature image. Figure 8 In the diagram, the portion located between the third and fourth feature images is the ECA model. Only some operations included in this ECA model are illustrated as examples, such as the average pooling operation (GAP) and the sigmoid activation function.

[0135] In steps 201 to 208, the portion of the fourth feature image belonging to the first feature image and the portion belonging to the second feature image in the final fourth feature image obtained by the target recognition device 100 are assigned different weights. This corresponds to whether the discriminative region in the input image is occluded or not. Based on this, when the discriminative region in the input image is occluded, a higher weight can be assigned to the first feature image and a lower weight to the second feature image. This allows for the acquisition of more information from other regions besides the discriminative region during subsequent target recognition, aiding in the identification of targets within the discriminative region and determining the target's category. When the discriminative region in the input image is not occluded, a lower weight can be assigned to the portion of the fourth feature image belonging to the first feature image and a higher weight to the portion belonging to the second feature image. This allows for the acquisition of more information from the discriminative region during subsequent target recognition, enabling a more comprehensive analysis of the discriminative region and accurate identification of targets within it.

[0136] In this embodiment, steps 202 to 208 can be executed by a discriminative fine-grained feature representation method based on double attention (DMF) device. The DMF device can be embedded in a neural network, for example, after the network layer used to extract image features. This embodiment does not limit the position or number of DMF devices. For example, the DMF device can be embedded after each network layer in the neural network that can extract image features. Taking ResNet50 as an example, the DMF device can be embedded in the CNN, after each stage of the CNN. Typically, in a neural network, between the network layer that extracts image features and the classifier, there are other network layers that can process the fourth feature image output by the DMF device. After a series of processing (the specific type of processing is not limited here; for example, it can be a convolution operation, a pooling operation, or a combination of convolution and pooling operations), a fifth feature image is obtained, which can then be transmitted to the classifier for classification.

[0137] In order to further improve the accuracy of the classifier's classification, that is, to improve the accuracy of target recognition, the target recognition device 100 can further process the fifth feature image. The further processing method of the fifth feature image is described below.

[0138] Step 209: Based on the fifth feature image, the target recognition device 100 determines multiple candidate feature images. Each candidate feature image corresponds to a different receptive field; that is, each candidate feature image corresponds to one receptive field. The receptive field refers to the size of the region mapped by a pixel on the feature image onto the input image. The number of candidate feature images is not limited here and can be determined according to the actual application scenario.

[0139] To obtain multiple candidate feature images, the target recognition device 100 can apply multiple convolution kernels of different sizes to the fifth feature image, and obtain multiple candidate feature images through dilation and separation convolution. Since the convolution kernels are of different sizes, the receptive fields of the multiple candidate feature images obtained by dilation and separation convolution are also different.

[0140] Step 210: The target recognition device 100 fuses multiple candidate feature images into a sixth feature image. The receptive field of the sixth feature image is a region with fewer redundant areas. These redundant areas are unfavorable for target recognition, meaning they contain little or no information characterizing the target category. From another perspective, the receptive field of the sixth feature image contains more effective information. This effective information indicates information that can be used for target recognition, such as information that can be extracted by a classifier and whose target type can be determined based on it.

[0141] Taking bird category recognition as an example, the receptive field of this sixth feature image can include fewer non-bird areas. For bird category recognition, the current image shows the bird's head, but the color of the bird's head feathers will vary depending on the shooting scene. That is, the color of the head feathers is not conducive to target recognition and belongs to redundant areas. However, areas such as the bird's beak and eyes do not easily change due to different shooting scenes. Areas such as the bird's beak and eyes contain more effective information, and the receptive field of this sixth feature image can include areas such as the bird's beak and eyes.

[0142] When the target recognition device 100 fuses the multiple candidate feature images, it can aggregate and reduce the dimensions of the multiple candidate feature images, and assign weights to the parts of the aggregated and reduced feature images that belong to each candidate feature image to obtain a sixth feature image. The weights assigned to each candidate feature image are obtained in advance through training.

[0143] Through steps 209-210, the target recognition device 100 acquires multiple candidate feature images, each of which corresponds to a receptive field. The size of the receptive field varies for each candidate feature image. Some receptive fields may only cover a portion of the target, while others may cover the target but also include a significant portion of non-target areas. By fusing (aggregating, reducing dimensionality, and configuring weights) these multiple candidate feature images, an adaptive process between the receptive field and the target can be achieved, resulting in a receptive field more conducive to target recognition. This receptive field is the receptive field of the sixth feature image. Thus, when target recognition is performed based on this sixth feature image, the target's features can be extracted more accurately, and the target's type can be determined.

[0144] like Figure 9 The diagram shown is a flowchart of the target recognition device 100 generating the sixth feature image. Figure 8 In this process, the target recognition device 100 utilizes six different convolution kernels to perform convolution operations on the sixth feature image.

[0145] The six convolution kernels are a 1*1 convolution kernel, a 3*3 convolution kernel with a dilation rate of 1, a 3*3 convolution kernel with a dilation rate of 2, a 3*3 convolution kernel with a dilation rate of 3, a 3*3 convolution kernel with a dilation rate of 4, and a 3*3 convolution kernel with a dilation rate of 5.

[0146] The kernel size refers to the dimensions of the kernel, which are its length and width. Commonly used sizes include 3x3 and 5x5.

[0147] The fifth feature image, after passing through one convolutional kernel, will output a candidate feature image of size C*H*W. The fifth feature image, after passing through six convolutional kernels, will produce six candidate feature images of size C*H*W.

[0148] The target recognition device 100 can aggregate and reduce the dimensionality of the six candidate feature images of size C*H*W in the channel dimension to obtain a 6C / N*H*W feature image (where N is the parameter for dimensionality reduction in the channel dimension). Figure 9 (Taking N=16 as an example for illustration), after dimensionality reduction, weights can be redistributed for each candidate feature image along the channel dimension. Figure 9 The diagram only illustrates a few operations involved in weight redistribution, such as Global Average Pooling (GAP), a 1x1 convolutional kernel (con1x1), BN+ReLU, and the sigmoid function. The configured weights can be obtained through pre-training or learning.

[0149] In this system, a 1x1 convolutional kernel performs convolution operations, and the number of 1x1 kernels can be adjusted to increase or decrease dimensions. GAP (Gross Average) refers to the summation and averaging of all feature values ​​in a feature image, resulting in a numerical value that represents the overall feature information of the image. BN+ReLU is the normalization and activation function in a convolutional neural network, primarily used for normalization and enhancing non-linear operations. FC (Fully Connected) is a common layer in neural networks, acting as a "classifier" within the overall neural network.

[0150] It should be noted that, in Figure 9 In this method, a convolutional kernel of size 1*1 is used. By setting a 1*1 convolutional kernel and performing inverse convolution, the original feature information can be effectively preserved. Furthermore, global information can be obtained through GAP, which effectively compensates for the discontinuity of information obtained that may be caused by dilated convolution, thereby obtaining a more complete and efficient feature representation.

[0151] The target recognition device 100 can directly assign weights to the dimensionality-reduced feature image along the channel dimension to obtain the sixth feature image, or it can assign weights to the dimensionality-reduced feature image along the channel dimension and then aggregate and reduce the dimensionality with another feature image to generate the sixth feature image. This other feature image can be a feature image generated by performing average pooling on the fifth feature image, and its size is C*H*W. The purpose of aggregating and reducing the dimensionality with another feature image is to ensure consistency between the input and output feature dimensions, and appropriate dimensionality reduction can effectively improve computational efficiency and recognition accuracy.

[0152] Step 211: The target recognition device 100 performs target recognition based on the sixth feature image.

[0153] When performing step 211, the target recognition device 100 can use a classifier. The classifier can be pre-trained and can determine the category of the target in the feature image based on the feature image, thereby achieving target recognition.

[0154] It should be understood that in steps 209 and 210, the target recognition device 100 processes the fifth feature image. Of course, in actual application scenarios, the target recognition device 100 can also directly process the fourth feature image after acquiring it.

[0155] In this embodiment of the application, steps 209 to 210 can be performed by a multi-scale feature fusion method based on receptive field adaptive adjustment (RFAM) device. The RFAM device can be placed before the classifier to process the feature image that needs to be input to the classifier so that the classifier can finally output an accurate result.

[0156] The following describes the implementation of the target recognition method provided in this application embodiment in ResNet50 from the perspective of overall application. See [link to relevant documentation]. Figure 10A This is a flowchart illustrating the image recognition method applied to ResNet50. In this method, the image recognition device can be divided into three parts, which for ease of distinction are referred to as the DFM device, the RFAM device, and the classifier. The DFM device executes steps 201-208 as shown in the embodiment above. The RFAM device executes steps 209-210 as shown in the embodiment above. The classifier executes step 211 as shown in the embodiment above.

[0157] ResNet50 includes a main thread comprising a main CNN and a main RFAM unit. The main CNN extracts features from the input image and outputs a feature image. A DFM unit can be added to the main CNN, such as... Figure 10BThe diagram illustrates the structure of the main CNN, which includes four stages (each stage is essentially a convolutional layer used for feature extraction). A Design for Rendering (DFM) device can be added after each stage. Each DFM device can process the feature images output by the stages preceding it, as in steps 201-208 of this embodiment. The main RFAM device can process the feature images output by the main CNN. The main CNN can output multiple feature images, one of which is a feature image for the entire input image. This feature image can be transmitted to the main RFAM device for processing. These multiple feature images also include feature images for different regions of the input image. These feature images, containing more information, can be transmitted to multiple branches following the main CNN for processing. Each branch processes one feature image; for example, four branches are connected after the main CNN. Each branch includes a branch CNN and a branch RFAM device. For any given branch, it can process one feature image output by the main CNN. Specifically, the branch CNN can continue to extract features from the feature image and output a new feature image. A DFM device can be added to the branch CNN. The method of adding a DFM device to the branch CNN is the same as that of adding a DFM device to the main CNN, as explained above, and will not be repeated here. The branch RFAM device can process the feature image output by the branch CNN.

[0158] The feature images output by the RFAM device in the main line and the branch RFAM devices in each branch can be input into the classifier. The classifier can perform target recognition based on the feature images. Then, the structures of each classifier are summarized and the final result is output, which can indicate the target in the input image.

[0159] Based on the same inventive concept as the method embodiments, this application also provides a target recognition device for performing the above-described... Figures 3A-3B The method executed by the target recognition device in the method embodiment shown in Figure 4 has related features that can be found in the above method embodiments, and will not be repeated here. Figure 11 As shown, a target recognition device 1100 provided in an embodiment of this application includes an acquisition unit 1101, an image generation unit 1102 and a recognition unit 1103, and optionally, a determination unit 1104 may also be included.

[0160] The acquisition unit 1101 is used to acquire an input image, which includes the target to be identified. The acquisition unit 1101 can perform... Figure 3A Step 101 in the method embodiment shown. The acquisition unit 1101 can execute... Figure 3B Step 201 in the method embodiment shown.

[0161] Image generation unit 1102 is configured to generate a first feature image based on a discriminative region of an input image. The first feature image is a feature image in which the discriminative region of the input image is occluded. The discriminative region of the input image is a subset of regions in the input image that can indicate the category to which the target belongs. Image generation unit 1102 can perform... Figure 3A Step 102 in the method embodiment shown. The image generation unit 1102 can perform... Figure 3B Steps 203-204 in the method embodiment shown.

[0162] The recognition unit 1103 is used to recognize a target based on a first feature image. The recognition unit 1103 can perform... Figure 3A Step 103 in the method embodiment shown.

[0163] As one possible implementation, the image generation unit 1102 can also obtain a second feature image without occluding the discriminative regions in the input image; that is, the second feature image is an image of the input image where the discriminative regions are not occluded. The image generation unit 1102 can perform... Figure 3B Steps 205-206 in the method embodiment shown.

[0164] When the recognition unit 1103 recognizes a target based on the first feature image, the recognition unit 1103 can simultaneously consider the first feature image and the second feature image, and recognize the target based on the first feature image and the second feature image. The recognition unit 1103 can perform... Figure 3B Steps 207 to 211 in the method embodiment shown.

[0165] As one possible implementation, after the image generation unit 1102 generates the first feature image and the second feature image, the determination unit 1104 can also determine the discriminative region based on the spatial features of the input image.

[0166] As one possible implementation, when determining a discriminative region based on the spatial features of the input image, the determining unit 1104 can assign scores to the spatial features of the input image; the determining unit 1104 can use regions where the scores of the spatial features are greater than a threshold as discriminative regions. Alternatively, regions where the scores are within a preset range can be used as discriminative regions. The determining unit 1104 can perform... Figure 3B Step 202 in the method embodiment shown. The determining unit 1104 can execute... Figure 4 The method embodiment shown.

[0167] In one possible implementation, when generating a first feature image based on the discriminative regions of the input image, the image generation unit 1102 can configure first coefficient values ​​for the pixels in the input image. The graph formed by the first coefficient values ​​of each pixel is called the first coefficient graph. There are many ways to configure the first coefficient values ​​for the pixels. For example, the first coefficient values ​​of the pixels belonging to the discriminative regions in the input image can be configured as smaller first values; the first coefficient values ​​of the remaining pixels can be configured as larger second values, wherein the first value is less than the second value. The graph formed by the first coefficient values ​​of each pixel is called the first coefficient graph. After obtaining the first coefficient graph, the image generation unit 1102 applies the first coefficient graph to the input image to generate the first feature image.

[0168] In one possible implementation, when generating a second feature image based on the discriminative regions of the input image, the image generation unit 1102 can configure second coefficient values ​​for the pixels in the input image. The image composed of the second coefficient values ​​of each pixel is called a second coefficient image. There are many ways to configure second coefficient values ​​for pixels. For example, the image generation unit 1102 can configure the second coefficient values ​​of pixels belonging to the discriminative regions in the input image as the spatial feature scores of the pixels; or, for another example, the image generation unit 1102 can configure the second coefficient values ​​of pixels belonging to the discriminative regions in the input image as a larger third value, and the second coefficient values ​​of the remaining pixels as a smaller fourth value, wherein the third value is greater than the fourth value. After obtaining the second coefficient image, the second coefficient image is applied to the input image to generate the second feature image.

[0169] As one possible implementation, when recognizing a target based on a first feature image and a second feature image, the recognition unit 1103 may first aggregate the first feature image and the second feature image along the channel dimension to generate a third feature image. Then, based on the third feature image, multiple candidate feature images of the same size are determined, wherein each candidate feature image has a different receptive field; the multiple candidate feature images are fused into a fourth feature image; and the fourth feature image is used for target recognition.

[0170] As one possible implementation, when the recognition unit 1103 aggregates the first feature image and the second feature image in the channel dimension to generate the third feature image, it can aggregate the first feature image and the second feature image in the channel dimension to reduce the dimensionality and generate an aggregated image; and configure weights for the aggregated image in the channel dimension to generate the third feature image. The weights configured for the candidate feature images can satisfy the following conditions: when the discriminative region is occluded in the input image, the weight of the part of the aggregated image belonging to the first feature image in the channel is greater than the weight of the part of the aggregated image belonging to the second feature image in the channel; or when the discriminative region is not occluded in the input image, the weight of the part of the aggregated image belonging to the first feature image in the channel is less than the weight of the part of the aggregated image belonging to the second feature image in the channel.

[0171] As one possible implementation, when the recognition unit 1103 determines multiple candidate feature images of the same size based on the third feature image, it can apply multiple different convolution kernels to the third feature image respectively, and obtain multiple candidate feature images by expanding and separating convolution.

[0172] As one possible implementation, when the recognition unit 1103 fuses multiple candidate feature images into a fourth feature image, it can configure a corresponding weight for each candidate feature image, and then obtain the fourth feature image based on each candidate feature image and the weight corresponding to each candidate feature image.

[0173] It should be noted that the division of units in the embodiments of this application is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods. The functional units in the embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units.

[0174] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive (SSD).

[0175] In a simplified embodiment, those skilled in the art will conceive of, as follows: Figures 3A-3B In the embodiments shown, the target recognition device may employ Figure 12 As shown in the figure.

[0176] like Figure 12 The device 1200 shown includes at least one processor 1201 and a memory 1202, and optionally, may also include a communication interface 1203.

[0177] Memory 1202 may be volatile memory, such as random access memory; memory may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD) or solid-state drive; or memory 1202 may be any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 1202 may be a combination of the above-described memories.

[0178] The specific connection medium between the processor 1201 and the memory 1202 described in this application embodiment is not limited.

[0179] Processor 1201 can be a central processing unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, artificial intelligence chips, on-chip chips, etc. General-purpose processors can be microprocessors or any conventional processor. It has data transmission and reception capabilities and can communicate with other devices, such as... Figure 12 The device can also be equipped with a separate data transceiver module, such as the communication interface 1203, for sending and receiving data. When the processor 1201 communicates with other devices, it can transmit data through the communication interface 1203, such as acquiring input images.

[0180] When the target recognition device adopts Figure 12 When in the form shown, Figure 12 The processor 1201 can call computer execution instructions stored in the memory 1202, so that the target identification device can execute the method executed by the target identification device in any of the above method embodiments.

[0181] Specifically, Figure 11 The functions / implementation processes of the acquisition unit, image generation unit, recognition unit, and determination unit can all be achieved through... Figure 12 The processor 1201 in the memory calls computer execution instructions stored in memory 1202 to implement this. Or, Figure 11 The functions / implementation processes of the image generation unit, recognition unit, and determination unit in the image can be understood through... Figure 12 The processor 1201 in the memory calls computer execution instructions stored in the memory 1202 to implement this. Figure 11 The functions / implementation process of the acquisition unit and the sending unit can be obtained through Figure 12 It is implemented using the communication interface 1203.

[0182] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0183] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0184] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0185] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0186] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method of target recognition, characterized in that, The method comprises: obtaining an input image, the input image comprising a target to be recognized; generating a first feature image and a second feature image according to a distinctive region of the input image, the first feature image being a feature image of the distinctive region of the input image being occluded, the second feature image being a feature image of the distinctive region of the input image not being occluded, the distinctive region of the input image being a subset of regions of the input image capable of indicating a category to which the target belongs; aggregating the first feature image and the second feature image in a channel dimension to generate a third feature image; determining a plurality of candidate feature images based on the third feature image, wherein each of the candidate feature images has the same size, and each of the candidate feature images has a different receptive field; fusing the plurality of candidate feature images into a fourth feature image; recognizing the target according to the fourth feature image.

2. The method of claim 1, wherein, The method further comprises: determining the distinctive region according to a spatial feature of the input image.

3. The method of claim 2, wherein, The determination of the distinctive region according to the spatial feature of the input image comprises: configuring a score for the spatial feature of the input image; regarding a region with a score greater than a threshold value of the spatial feature as the distinctive region.

4. The method according to any one of claims 1 to 3, characterized in that, The generation of the first feature image according to the distinctive region of the input image comprises: configuring a first coefficient value of a pixel point belonging to the distinctive region in the input image as a first value, and configuring a first coefficient value of the rest of the pixel points as a second value, wherein the first value is less than the second value, and a graph constituted by the first coefficient values of each of the pixel points is a first coefficient graph; applying the first coefficient graph to the input image to generate the first feature image.

5. The method of claim 1, wherein, The generation of the second feature image according to the distinctive region of the input image comprises: configuring a second coefficient value of a pixel point belonging to the distinctive region in the input image as a score of a spatial feature of the pixel point to generate a second coefficient graph; applying the second coefficient graph to the input image to generate the second feature image.

6. The method according to any one of claims 1 to 3, 5, wherein The generation of the second feature image according to the distinctive region of the input image comprises: configuring a second coefficient value of a pixel point belonging to the distinctive region in the input image as a third value, and configuring a second coefficient value of the rest of the pixel points as a fourth value, wherein the third value is greater than the fourth value, and a graph constituted by the second coefficient values of each of the pixel points is a second coefficient graph; applying the second coefficient graph to the input image to generate the second feature image.

7. The method of claim 1, wherein, The aggregation of the first feature image and the second feature image in the channel dimension to generate the third feature image comprises: aggregating the first feature image and the second feature image in the channel dimension to reduce dimensionality and generate an aggregated image; The third feature image is generated by configuring a weight for the aggregated image in a channel dimension, and the weight configured for the candidate feature image satisfies a condition that a weight of a part belonging to the first feature image in the aggregated image is greater than a weight of a part belonging to the second feature image in the aggregated image in a channel when the distinctive region is occluded in the input image, or the weight of the part belonging to the first feature image in the aggregated image is less than the weight of the part belonging to the second feature image in the aggregated image in the channel when the distinctive region is not occluded in the input image.

8. The method of claim 1, wherein, The third feature image is used to determine a plurality of candidate feature images, including: The plurality of different convolution kernels are applied to the third feature image to obtain the plurality of candidate feature images by using dilated separable convolution.

9. The method of claim 1, 7 or 8, wherein, The fourth feature image is obtained by fusing the plurality of candidate feature images, including: The fourth feature image is obtained based on each candidate feature image and a weight corresponding to each candidate feature image.

10. A target recognition device, characterized by The device includes: An acquisition unit configured to acquire an input image, the input image including a target to be recognized; An image generation unit configured to generate a first feature image and a second feature image according to a distinctive region of the input image, the first feature image being a feature image in which the distinctive region of the input image is occluded, the second feature image being a feature image in which the distinctive region of the input image is not occluded, the distinctive region of the input image being a subset of regions of the input image that can indicate a category to which the target belongs; An identification unit configured to aggregate the first feature image and the second feature image in a channel dimension to generate a third feature image, determine a plurality of candidate feature images based on the third feature image, wherein each candidate feature image has a same size and a different receptive field, fuse the plurality of candidate feature images into a fourth feature image, and identify the target according to the fourth feature image.

11. The apparatus of claim 10, wherein, The device further includes a determination unit configured to: determine the distinctive region according to a spatial feature of the input image.

12. The apparatus of claim 11, wherein, When determining the distinctive region according to the spatial feature of the input image, the determination unit is specifically configured to: configure a score for the spatial feature of the input image; regard a region with a score greater than a threshold value as the distinctive region.

13. The apparatus of any one of claims 10-12, wherein, When generating the first feature image according to the distinctive region of the input image, the image generation unit is specifically configured to: configure a first coefficient value of a pixel point belonging to the distinctive region in the input image as a first value, and configure a first coefficient value of a remaining pixel point as a second value, wherein the first value is less than the second value, and a graph formed by the first coefficient values of the pixel points is a first coefficient graph; apply the first coefficient graph to the input image to generate the first feature image.

14. The apparatus of claim 13, wherein, When generating the second feature image according to the distinctive region of the input image, the image generation unit is specifically configured to: The second coefficient value of a pixel point belonging to the distinctive region in the input image is configured as a score of a spatial feature of the pixel point, to generate a second coefficient map; The second coefficient map is applied to the input image to generate the second feature image.

15. The apparatus of any of claims 10-12, 14, wherein, In the process of generating the second feature image according to the distinctive region of the input image, the image generation unit is specifically configured to: The second coefficient value of a pixel point belonging to the distinctive region in the input image is configured as a third value, and the second coefficient value of the remaining pixel points is configured as a fourth value, wherein the third value is greater than the fourth value, and the second coefficient value of each pixel point constitutes a second coefficient map; The second coefficient map is applied to the input image to generate the second feature image.

16. The apparatus of claim 10, wherein, In the process of generating the third feature image by aggregating the first feature image and the second feature image in the channel dimension, the recognition unit is specifically configured to: The first feature image and the second feature image are aggregated in the channel dimension to generate an aggregated image; The aggregated image is configured with a weight in the channel dimension to generate the third feature image, and the weight configured for the candidate feature image satisfies the following condition: when the distinctive region is occluded in the input image, the weight of the part belonging to the first feature image in the channel of the aggregated image is greater than the weight of the part belonging to the second feature image in the channel of the aggregated image, or when the distinctive region is not occluded in the input image, the weight of the part belonging to the first feature image in the channel of the aggregated image is less than the weight of the part belonging to the second feature image in the channel of the aggregated image.

17. The apparatus of claim 10, wherein, In the process of determining a plurality of candidate feature images based on the third feature image, the recognition unit is specifically configured to: A plurality of different convolution kernels are applied to the third feature image to obtain the plurality of candidate feature images through dilated separable convolution.

18. The apparatus of claims 10, 16, or 17, wherein, In the process of fusing the plurality of candidate feature images into the fourth feature image, the recognition unit is specifically configured to: The fourth feature image is obtained based on each candidate feature image and the weight corresponding to each candidate feature image.

19. An apparatus, comprising: A computer readable storage medium stores instructions, when the instructions are run on a computer, the computer executes the method of any one of claims 1-9.

20. A computer-readable storage medium, characterized in that, A computer readable storage medium stores instructions, when the instructions are run on a computer, the computer executes the method of any one of claims 1-9.