Object recognition method and device, electronic equipment, medium and program product
By performing multi-level feature extraction and aggregation of finger vein image data, integrating feature information of different scales and resolutions, the problem of insufficient recognition accuracy in the prior art is solved, and higher recognition accuracy and feature robustness are achieved.
Patent Information
- Application Number
- CN202510208053.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-13
AI Technical Summary
There is still room for improvement in the recognition accuracy of existing venous recognition technology, especially when processing feature information of different scales and resolutions.
By performing feature extraction at different levels of feature images of the target object, the initial feature map sequence is obtained and multi-level aggregation is performed to integrate feature information of different scales and resolutions. Aggregation processing at each level includes feature preprocessing and aggregation processing, which improves the robustness and accuracy of features through technical means such as channel attention weighting and dense connection.
The recognition accuracy of finger vein images is improved, the robustness and accuracy of target features are enhanced, and the reliability of recognition results is improved.
Smart Images

Figure CN119992612A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, specifically to the field of artificial intelligence, deep learning and image detection technology, and more specifically to an object recognition method, device, electronic device, medium and program product. Background Art
[0002] Finger vein recognition technology is a biometric technology that uses the vein pattern image obtained after near-infrared rays penetrate the finger to identify individuals. Among various biometric technologies, finger vein recognition technology has a higher anti-counterfeiting ability because it uses the internal characteristics of the organism that cannot be seen from the outside.
[0003] With the rapid development of artificial intelligence and biometrics, finger vein recognition technology has been widely used in payment, security, subway, door locks, consumer electronics and other fields. However, the recognition accuracy of finger vein recognition still needs to be improved. Summary of the invention
[0004] In view of the above problems, the present disclosure provides an object recognition method, apparatus, device, medium and program product.
[0005] According to the first aspect of the present disclosure, an object recognition method is provided, including: performing feature extraction at different levels on the finger vein image data of the target object to obtain an initial feature map sequence representing different feature scales. The finger vein image data is collected under infrared irradiation, and the initial feature map sequence is obtained by sorting multiple initial feature maps obtained by feature extraction according to the feature scale. The initial feature map sequence is subjected to multi-level aggregation processing to obtain a target feature map. The aggregation processing at each level includes the following operations: performing feature preprocessing on the aggregated feature map of the previous level to obtain a feature map to be aggregated with the same feature map size as the adjacent aggregated feature map of the previous level, wherein the aggregated feature map of the previous level and the adjacent aggregated feature map of the previous level are respectively obtained by the aggregation processing of the previous level; performing aggregation processing on the feature map to be aggregated of the previous level and the adjacent aggregated feature map of the previous level to obtain an aggregated feature map of the current level. Based on the target feature map and the reference feature map of the target object, the target object is identified to obtain an identification result.
[0006] According to an embodiment of the present disclosure, feature preprocessing of the aggregate feature map of the previous level includes: when the aggregate feature map of the previous level is at the first position in the sequence of aggregate feature maps of the previous level, sequentially performing feature extraction processing and downsampling processing on the aggregate feature map of the previous level, wherein the sequence of aggregate feature maps of the previous level is obtained by sorting multiple aggregate feature maps of the previous level according to the feature scale.
[0007] According to an embodiment of the present disclosure, feature preprocessing of the aggregate feature map of the previous level also includes: when the aggregate feature map of the previous level is between the first and the last part of the aggregate feature map sequence of the previous level, downsampling the aggregate feature map of the previous level.
[0008] According to an embodiment of the present disclosure, aggregating the feature map to be aggregated at the previous level with the adjacent aggregated feature map at the previous level to obtain an aggregated feature map at the current level includes: when the adjacent aggregated feature map at the previous level is an initial feature map, performing channel attention weighted processing on the adjacent aggregated feature map at the previous level to obtain the adjacent feature map to be aggregated at the previous level; aggregating the feature map to be aggregated at the previous level with the adjacent feature map to be aggregated at the previous level to obtain the aggregated feature map at the current level.
[0009] According to an embodiment of the present disclosure, aggregating the feature map to be aggregated at the previous level and the adjacent aggregate feature map at the previous level to obtain the aggregate feature map at the current level also includes: cascading the feature map to be aggregated at the previous level and the adjacent aggregate feature map at the previous level according to channels to obtain a densely connected aggregate feature map; and obtaining the aggregate feature map of the current level based on the densely connected aggregate feature map.
[0010] According to an embodiment of the present disclosure, obtaining an aggregate feature map of the current level based on a densely connected aggregate feature map includes: when the aggregate feature map of the previous level is between the first and the last part of the aggregate feature map sequence of the previous level, performing dimensionality reduction processing and downsampling processing on the densely connected aggregate feature map in sequence to obtain an aggregate feature map of the current level.
[0011] According to an embodiment of the present disclosure, based on the target feature and the reference feature map of the target object, the target object is identified, and obtaining the identification result includes: matching the target feature map with the reference feature map to obtain a matching degree. The reference feature map is obtained by processing the reference finger vein image data of the target object in the same manner as the target feature map; and obtaining the identification result based on the matching degree.
[0012] According to an embodiment of the present disclosure, an extraction module is used to perform feature extraction at different levels on the finger vein image data of the target object to obtain an initial feature map sequence representing different feature scales. The finger vein image data is collected under infrared irradiation, and the initial feature map sequence is obtained by sorting multiple initial feature maps obtained by feature extraction according to the feature scale. A processing module performs multi-level aggregation processing on the initial feature map sequence to obtain a target feature map. The aggregation processing at each level includes the following operations: performing feature preprocessing on the aggregated feature map of the previous level to obtain a feature map to be aggregated with the same feature map size as the adjacent aggregated feature map of the previous level, wherein the aggregated feature map of the previous level and the adjacent aggregated feature map of the previous level are respectively obtained through the aggregation processing of the previous level; performing aggregation processing on the feature map to be aggregated of the previous level and the adjacent aggregated feature map of the previous level to obtain an aggregated feature map of the current level. An identification module is used to identify the target object based on the target feature map and the reference feature map of the target object to obtain an identification result.
[0013] A third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0014] The fourth aspect of the present disclosure further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the above computer program or instructions are executed by a processor.
[0015] The fifth aspect of the present disclosure further provides a computer program product, including a computer program or instructions, which implement the steps of the above method when the above computer program or instructions are executed by a processor.
[0016] According to the embodiments of the present disclosure, by performing feature extraction at different levels on the finger vein image data of the target object, and performing multi-level aggregation processing on the extracted initial feature map sequence, at each level, the aggregated feature map of the previous level is feature preprocessed and then aggregated with the adjacent aggregated feature map, thereby integrating feature information of different scales and resolutions, improving the robustness and accuracy of the target features, and thereby improving the recognition accuracy of the finger vein image. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0018] Figure 1 The application scenario diagram of the object recognition method, apparatus, device, medium and program product according to the embodiments of the present disclosure is schematically shown;
[0019] Figure 2 A flowchart of an object recognition method according to an embodiment of the present disclosure is schematically shown;
[0020] Figure 3 A flowchart of generating an initial feature map according to an embodiment of the present disclosure is schematically shown;
[0021] Figure 4 A schematic diagram of a network structure for performing multi-layer aggregation according to an embodiment of the present disclosure is shown;
[0022] Figure 5 The structure diagram of the aggregation module according to the embodiment of the present disclosure is schematically shown;
[0023] Figure 6 A schematic diagram of a network structure for performing multi-layer aggregation according to an embodiment of the present disclosure is shown;
[0024] Figure 7 A schematic diagram of a structure of an object recognition device according to an embodiment of the present disclosure is shown; and
[0025] Figure 8 A block diagram of an electronic device suitable for implementing an object recognition method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0026] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0027] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0028] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0029] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0030] In the technical solution of the present disclosure, the user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0031] The embodiment of the present disclosure provides an object recognition method, including: performing feature extraction at different levels on the finger vein image data of the target object to obtain an initial feature map sequence representing different feature scales. The finger vein image data is collected under infrared irradiation, and the initial feature map sequence is obtained by sorting multiple initial feature maps obtained by feature extraction according to the feature scale. The initial feature map sequence is subjected to multi-level aggregation processing to obtain a target feature map. The aggregation processing at each level includes the following operations: performing feature preprocessing on the aggregated feature map of the previous level to obtain a feature map to be aggregated with the same feature map size as the adjacent aggregated feature map of the previous level, and the aggregated feature map of the previous level and the adjacent aggregated feature map of the previous level are respectively obtained by the aggregation processing of the previous level; performing aggregation processing on the feature map to be aggregated of the previous level and the adjacent aggregated feature map of the previous level to obtain an aggregated feature map of the current level. Based on the target feature map and the reference feature map of the target object, the target object is identified to obtain an identification result.
[0032] Figure 1 The application scenario diagram of the object recognition method and device according to the embodiments of the present disclosure is schematically shown.
[0033] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.
[0034] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).
[0035] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0036] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0037] It should be noted that the object recognition method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the object recognition device provided in the embodiment of the present disclosure can generally be set in the server 105. The object recognition method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the object recognition device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0038] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0039] The following will be based on Figure 1 The scene described by Figure 2~Figure 6 The object recognition method of the disclosed embodiment is described in detail.
[0040] Figure 2 The flowchart of the object recognition method according to the embodiment of the present disclosure is schematically shown.
[0041] like Figure 2 As shown, the object recognition of this embodiment includes operations S210 to S230.
[0042] In operation S210, feature extraction at different levels is performed on the finger vein image data of the target object to obtain an initial feature map sequence representing different feature scales.
[0043] According to an embodiment of the present disclosure, the finger vein image data may be collected under infrared irradiation. For example, an infrared light source may be used to illuminate the finger to make the venous blood vessel image inside the finger appear, and then a CCD camera (Charge-coupled Device Camera) may be used to capture the finger vein image data.
[0044] According to an embodiment of the present disclosure, the initial feature map sequence is obtained by sorting a plurality of initial feature maps obtained by feature extraction according to feature scales.
[0045] Initial feature maps of different feature scales differ in semantic information and target location. For example, the initial feature map at a low level has a high positioning accuracy in the target area, but relatively less semantic information. The initial feature map at a high level has a relatively vague positioning in the target area, but is richer in semantic information.
[0046] According to an embodiment of the present disclosure, the initial image data can be input into the skeleton network, and the feature maps of the finger vein image data at different processing stages of the skeleton network are obtained as the initial feature map sequence. The skeleton network can be, for example, a skeleton convolutional neural network, but the present disclosure is not limited thereto, as long as it can extract features of different scales from the initial image data.
[0047] In operation S220, a multi-level aggregation process is performed on the initial feature map sequence to obtain a target feature map.
[0048] In some embodiments, the initial feature map sequence can be input into the aggregation network, and the aggregation network can be composed of multiple levels of aggregation layers, and each level of aggregation layer is used to aggregate the aggregated feature map in the aggregation layer of the previous level to obtain the aggregated feature map of the current level, and then pass the obtained aggregated feature map downward until the aggregation layer of the last level. Based on the aggregated feature map of the aggregation layer of the last level, the target feature map is obtained.
[0049] According to an embodiment of the present disclosure, the aggregation process of each level may include operations S221~S222.
[0050] In operation S221, feature preprocessing is performed on the aggregated feature map of the previous level to obtain a feature map to be aggregated having the same feature map size as that of an adjacent aggregated feature map of the previous level.
[0051] In operation S222, the feature map to be aggregated at the previous level and the adjacent aggregated feature map at the previous level are aggregated to obtain an aggregated feature map at the current level.
[0052] According to an embodiment of the present disclosure, the aggregation feature map of the previous level and the adjacent aggregation feature map of the previous level are respectively obtained through the aggregation processing of the previous level.
[0053] The aggregated feature graphs of the previous level may include multiple ones, and the feature graph to be aggregated of the previous level and the adjacent aggregated feature graph of the previous level represent two aggregated feature graphs in adjacent positions in the same level.
[0054] Different preprocessing methods can be selected according to the scale information and size information of the aggregated feature map, so that the preprocessed feature map to be aggregated can match the size of the feature map of the adjacent aggregated feature map, thereby providing a basis for the next step of aggregation.
[0055] The initial feature map sequence can be used as the aggregate feature map of the first level. At the first level, the multiple initial feature maps in the initial feature map sequence are subjected to feature preprocessing respectively, and then the multiple initial feature maps subjected to feature preprocessing are aggregated with their respective adjacent initial feature maps to obtain multiple aggregate feature maps of the second level. The multiple aggregate feature maps of the second level are passed to the third level, and the feature preprocessing and aggregation operations are repeated until the aggregate feature map of the last level is obtained.
[0056] In operation S230, the target object is recognized based on the target feature map and the reference feature map of the target object to obtain a recognition result.
[0057] The reference feature map may be obtained by processing a finger vein image pre-stored on the target object. Based on the comparison between the target feature map and the reference feature map, it is determined whether the target feature map and the reference feature map are from the same finger.
[0058] According to the embodiments of the present disclosure, by performing feature extraction at different levels on the finger vein image data of the target object, and performing multi-level aggregation processing on the extracted initial feature map sequence, at each level, the aggregated feature map of the previous level is feature preprocessed and then aggregated with the adjacent aggregated feature map, thereby integrating feature information of different scales and resolutions, improving the robustness and accuracy of the target features, and further improving the recognition accuracy of the finger veins.
[0059] Figure 3The flowchart of generating an initial feature map according to an embodiment of the present disclosure is schematically shown.
[0060] like Figure 3 As shown, the finger vein image data 310 is processed by the skeleton network. After the finger vein image data 310 is input into the skeleton network, it can be sequentially processed through the first processing stage 320, the second processing stage 330, the third processing stage 340, the fourth processing stage 350 and the fifth processing stage 360. In the first processing stage 320, the finger vein image data can be processed by the filling layer, the convolution layer, the normalization layer and the activation layer in sequence to obtain the initial feature map of the first processing stage 310. In the second processing stage 330, the initial feature map of the first processing stage is processed by the filling layer, the pooling layer and the dense connection layer in sequence to obtain the initial feature map of the second processing stage. In the third processing stage, the fourth processing stage and the fifth processing stage, the transmitted initial feature map is processed by the transition layer and the dense connection layer in sequence to obtain the initial feature map of the third processing stage, the initial feature map of the fourth processing stage and the initial feature map of the fifth processing stage, respectively.
[0061] The initial feature map of the first processing stage is large in size and has a relatively small number of channels. The initial feature map of the second processing stage is halved in size, and the number of channels increases significantly due to dense connections. The initial feature map of the third processing stage is halved in size again, and the number of channels increases further. Similarly, the initial feature map of the fourth processing stage and the initial feature map of the fifth processing stage are halved in size again compared to the previous processing stage, and the number of channels increases further. At the same time, the image scale information is higher, and the features are more abstract and advanced. Therefore, by inputting the image data of the target opponent into the skeleton network, the initial feature maps of different feature scales of the five processing stages can be obtained, and then the initial feature map sequence can be obtained.
[0062] Figure 4 A schematic diagram of a network structure for performing multi-layer aggregation according to an embodiment of the present disclosure is shown schematically.
[0063] like Figure 4As shown, the network structure for multi-layer aggregation includes a processing layer 410, a first aggregation layer 420, a second aggregation layer 430, a third aggregation layer 440 and a fourth aggregation layer 450. The processing layer 410 includes multiple processing modules, corresponding to multiple processing stages, and respectively generates initial feature maps of different feature scales. The initial feature maps of the processing layer 410 in multiple processing stages can be used as the aggregated feature maps of the first level. Each aggregation layer includes at least one aggregation module. Each aggregation module in the first aggregation layer 420 is used to aggregate the initial feature maps generated by two adjacent processing modules. The aggregation modules in the second aggregation layer 430, the third aggregation layer 440 and the fourth aggregation layer 450 are respectively used to aggregate the aggregated feature maps of two adjacent aggregation modules in the previous layer. A target feature map can be generated in the fourth aggregation layer 450.
[0064] Since there is a difference in size between two adjacent aggregated feature maps at each level, when aggregating adjacent aggregated feature maps, the aggregated feature maps may be preprocessed to make the sizes of the two aggregated feature maps consistent.
[0065] According to an embodiment of the present disclosure, feature preprocessing of the aggregate feature graph of the previous level includes, when the aggregate feature graph of the previous level is at the first position of the aggregate feature graph sequence of the previous level, sequentially performing feature extraction processing and downsampling processing on the aggregate feature graph of the previous level. The aggregate feature graph sequence of the previous level is obtained by sorting multiple aggregate feature graphs of the previous level according to the feature scale.
[0066] According to an embodiment of the present disclosure, the feature extraction process can perform feature extraction through a convolution operation to further extract and refine features in the feature map. For example, the aggregated feature map of the previous level can be convolved through a 1×1 convolution kernel, a 3×3 convolution kernel, etc.
[0067] Downsampling processing can be performed by downsampling operations such as maximum pooling and average pooling to reduce the size of the feature map, which can reduce the amount of calculation and control overfitting.
[0068] For example, Figure 4 As shown, when the fourth aggregation module 431 aggregates the aggregated feature maps output by the first aggregation module 421 and the second aggregation module 422, it can perform 3×3 convolution and average downsampling processing on the aggregated feature maps output by the first aggregation module 421 in sequence, and then perform subsequent aggregation processing after obtaining a feature map to be aggregated with stronger expressive ability and achieving size alignment with the aggregated feature map output by the second aggregation module 422.
[0069] According to the embodiments of the present disclosure, since the aggregate feature map at the first position of the aggregate feature map sequence is obtained by processing based on the initial feature map of the first processing module, and the initial feature map of the first processing module has a large size, a relatively small number of channels, and relatively rough feature information, the receptive field can be increased and the feature representation capability can be improved by extracting features from the aggregate feature map at the first position of the aggregate feature map sequence, and the size of the aggregate feature map can be reduced by downsampling processing, thereby facilitating subsequent aggregation.
[0070] According to an embodiment of the present disclosure, feature preprocessing of the aggregate feature map of the previous level may also include downsampling the aggregate feature map of the previous level when the aggregate feature map of the previous level is between the first and last part of the aggregate feature map sequence of the previous level.
[0071] According to an embodiment of the present disclosure, the aggregated feature map between the first and last position of the aggregated feature map sequence of the previous level represents a feature map obtained after the vein image data has been subjected to multiple feature extractions. The feature map sequence can be aggregated in the middle position by maximum pooling, average pooling, etc. to perform downsampling operations to reduce the size of the feature map.
[0072] For example, Figure 4 As shown, when the fifth aggregation module 432 aggregates the aggregated feature maps output by the second aggregation module 422 and the third aggregation module 423, it can perform average downsampling processing on the aggregated features output by the second aggregation module 422 to obtain a feature map to be aggregated of the same size as the aggregated feature map output by the second aggregation module 422.
[0073] According to an embodiment of the present disclosure, by downsampling the aggregated feature map in the middle position of the aggregated feature map sequence, the spatial size of the feature map can be reduced, and the feature representation can be further refined, redundant information can be reduced, and a more suitable feature map size can be provided for subsequent aggregation operations.
[0074] According to an embodiment of the present disclosure, a feature map to be aggregated at a previous level and an adjacent aggregate feature map at the previous level are aggregated to obtain an aggregate feature map at the current level, including: when the adjacent aggregate feature map at the previous level is an initial feature map, channel attention weighted processing is performed on the adjacent aggregate feature map at the previous level to obtain the adjacent feature map to be aggregated at the previous level; and the feature map to be aggregated at the previous level and the adjacent feature map to be aggregated at the previous level are aggregated to obtain the aggregate feature map at the current level.
[0075] According to an embodiment of the present disclosure, channel attention weighted processing of adjacent aggregated feature maps of the previous level can be performed by using a 1×1 convolution kernel to perform an independent convolution operation on each channel of the adjacent features to be aggregated, so that each channel undergoes an independent linear transformation to obtain an adjacent feature map to be aggregated with the same number of channels as the initial feature map.
[0076] For example, when the adjacent aggregated feature map is the initial feature map generated by the third processing module 413, the second aggregation module 422 is used to aggregate the initial feature map generated by the second processing module 412 and the initial feature map generated by the third processing module 413, and the initial feature map generated by the second processing module 412 can be averaged pooled to reduce the size to obtain the feature map to be aggregated of the second processing module. For the initial feature map generated by the third processing module 413, a 1×1 convolution kernel can be used to perform independent convolution on each channel of the initial feature map generated by the third processing module 413, and the channels are weighted to obtain the adjacent feature map to be aggregated of the third processing module 413, and further extract information.
[0077] According to an embodiment of the present disclosure, when the adjacent aggregated feature map of the previous level is the initial feature map, channel attention weighted processing is performed on it, and different weights can be assigned to each channel of the initial feature map to emphasize or suppress the feature representation of different channels, so that the channel information is more refined, thereby improving the representation ability of the feature map to be aggregated.
[0078] According to an embodiment of the present disclosure, the feature map to be aggregated at the previous level and the adjacent aggregate feature map at the previous level are aggregated to obtain the aggregate feature map at the current level, and also includes: cascading the feature map to be aggregated at the previous level and the adjacent aggregate feature map at the previous level according to channels to obtain a densely connected aggregate feature map; and obtaining the aggregate feature map of the current level based on the densely connected aggregate feature map.
[0079] According to the embodiments of the present disclosure, during the aggregation process, the two feature maps can be concatenated in the channel dimension by cascading according to the channels, so that the different channel information of each feature map is integrated together to form a densely connected aggregate feature map with more channels.
[0080] For example, Figure 4 As shown, when the fourth aggregation module 431 is used to aggregate the aggregated feature graphs transmitted from the first aggregation module 421 and the second aggregation module 422, since the first aggregation module 421 and the second aggregation module 422 have undergone one aggregation, it is assumed that the first aggregation module 421 and the second aggregation module 422 include a total of multiple feature graphs such as F1, F2, ..., Fn, etc., and after special preprocessing, their dimensions are all R H×W×C, H is the height of the aggregate feature map, W is the size of the aggregate feature map, and C is the channel of each aggregate feature map. The channels will be cascaded, and the calculation formula is as follows:
[0081] (1)
[0082] The aggregated feature map F∈R H×W×(n×C) , it can be seen that F is expanded to n times in the channel dimension.
[0083] According to an embodiment of the present disclosure, by cascading the feature map to be aggregated in the previous layer and the adjacent aggregated feature map in the previous layer according to channels, feature information at different levels can be effectively integrated, so that low-level information can be fully utilized, while alleviating the problem of high-level targets being too coarse.
[0084] According to an embodiment of the present disclosure, when the aggregation module is the first aggregation module of the current level, the densely connected aggregation feature map can be used as the aggregation feature map of the current level.
[0085] According to an embodiment of the present disclosure, based on the densely connected aggregate feature map, obtaining the aggregate feature map of the current level includes, when the aggregate feature map of the previous level is between the first and the last part of the aggregate feature map sequence of the previous level, performing dimensionality reduction processing and downsampling processing on the densely connected aggregate feature map in sequence to obtain the aggregate feature map of the current level.
[0086] According to an embodiment of the present disclosure, when the aggregated feature map of the previous layer is between the first and last part of the aggregated feature map sequence of the previous layer, since the aggregated feature map at this time has gone through multiple feature special zones and the number of channels of the feature map is too many, the number of channels of the aggregation layer can be reduced through dimensionality reduction processing, and the size can be reduced through downsampling processing.
[0087] According to an embodiment of the present disclosure, a transition layer may be set in the aggregation module to perform dimensionality reduction and downsampling processing on the densely connected aggregated feature graph. The transition layer may include
[0088] It consists of a 1×1 convolutional layer and an average pooling layer. The convolutional layer is used to reduce the number of channels of the aggregated feature map, and the pooling layer is used to reduce the spatial size of the aggregated feature map. Assume that the aggregated feature map F is input in R H×W×C , H is the height of the aggregated feature map, W is the number of aggregated feature maps, C is the channel of each aggregated feature map, and the feature map output after convolution And the feature map after pooling output The formula is as follows:
[0089] (2)
[0090] (3)
[0091] Input feature map Entering the 1×1 convolutional layer, represents the compression factor, ×C represents the number of channels of the output feature map. , and then use average pooling to Reduce the size of the feature map, and output the feature map of the transition aggregation module ,in , , it can be seen halved in size.
[0092] Figure 5 The structure diagram of the aggregation module according to an embodiment of the present disclosure is schematically shown.
[0093] like Figure 5 As shown, the aggregation module may include a dense connection layer 510, a convolution layer 520, and a pooling layer 530. Two adjacent feature maps to be aggregated in the previous layer are first input to the dense connection layer 510 for dense connection to obtain a densely connected aggregated feature map, and then input to the convolution layer 520 for processing to reduce the channels of the densely connected aggregated feature map, and finally input to the pooling layer to reduce the size of the feature map.
[0094] According to an embodiment of the present disclosure, when the aggregated feature map of the previous layer is between the first and the last of the aggregated feature map sequence of the previous layer, the densely connected aggregated feature map is sequentially subjected to dimensionality reduction and downsampling processing. Through dimensionality reduction and downsampling processing, the computational workload and memory requirements of subsequent layers are reduced, thereby improving overall efficiency.
[0095] In some embodiments, when performing a convolution operation on an aggregated feature map, a convolution operation of magnitude can be used, including a depth separation convolution layer, a batch normalization layer, and a Relu activation function layer, to efficiently extract features from the input aggregated feature map. By separating the convolution operations of the channels, the feature information of each channel can be effectively extracted.
[0096] The formula for depthwise separable convolution is as follows:
[0097] (4)
[0098] is the cth channel of the output feature map at position The value at The size of the cth channel of the input feature map is local area, is the weight of c convolution kernels.
[0099] The batch normalization layer helps accelerate the convergence process of the model, reduce the gradient vanishing problem, and improve the generalization ability of the model. It normalizes the input of each batch to stabilize the input distribution, which is conducive to network training and convergence. The formula of batch normalization can be described as follows: Assume that the input feature map is , where N is the batch size, H is the height, W is the width, and C is the number of channels.
[0100] First, perform a sum operation on each channel to get the mean of each channel and variance , the formula is as follows:
[0101] (5)
[0102] (6)
[0103] Secondly, the features of each channel are normalized so that their mean is 0 and their variance is 1, where is a very small constant used to avoid the situation where the denominator is 0. The formula is as follows:
[0104] (7)
[0105] In order to increase the expressive power of the network, two learnable parameters γ and β can be introduced to scale the normalized features and add offsets, where and is the scaling factor and offset factor corresponding to channel c, which can be determined through training.
[0106] (8)
[0107] ReLU is a nonlinear activation function that can introduce nonlinear transformations into the network, which helps to enhance the network's expressive power. Its formula is as follows:
[0108] (9)
[0109] According to an embodiment of the present disclosure, based on the target feature and the reference feature map of the target object, identifying the target object and obtaining the identification result may include: matching the target feature map with the reference feature map to obtain a matching degree. The reference feature map is obtained by processing the reference finger vein image data of the target object in the same manner as that of obtaining the target feature map. Based on the matching degree, the identification result is obtained.
[0110] According to an embodiment of the present disclosure, the finger vein image data may be finger vein image data of a target object acquired in advance, for example, finger vein image data of the target object acquired during a registration or entry phase.
[0111] According to an embodiment of the present disclosure, the target feature map and the reference feature map are matched by matching the images based on the target feature map and the reference feature map. The target feature map and the reference feature map can be mapped into two feature vectors, and the matching is performed by calculating the cosine similarity of the two feature vectors. The formula is as follows:
[0112] (10)
[0113] Among them, a and b are two eigenvectors, a∙b is the dot product (inner product) of the two vectors, ‖a‖ and ‖b‖ represent the modulus of the two vectors. The value range of cosine similarity is [−1,1], where the closer the value is to 1, the more similar the two vectors are, the closer the value is to -1, the less similar the two vectors are, and the value of 0 indicates that the two vectors are orthogonal.
[0114] The present disclosure is not limited thereto, and the two feature maps may also be converted and input into a binary classification model, for example, using a Sigmoid activation function.
[0115] (11)
[0116] The Sigmoid function maps the input value to the interval (0,1), and the output can be interpreted as the probability of the positive class. When the output probability is greater than 0.5, the output is considered to be positive and the match is correct. To improve the matching confidence and security, the output threshold can be set to 0.8, for example.
[0117] In some embodiments, a twin network can be formed based on a skeleton network and a polymerization network, and two inputs are set, one input is the finger vein image data (i.e., the real-time finger vein image), and the other input is the reference finger vein image data entered into the database. The twin network is composed of two sub-networks composed of exactly the same skeleton network and polymerization network, and shares the same structure and parameters, and uses the same weight parameters and bias to process the finger vein image data and the reference finger vein image data.
[0118] The sample finger vein image data and sample object labels can be used to train the initial twin network to obtain the twin network. Under realistic conditions, the accuracy of finger vein image recognition may decrease due to the influence of factors such as environment, equipment and lighting conditions. Therefore, during training, the realistic environment can be simulated to randomly transform the input finger vein image in terms of contrast, brightness, saturation, horizontal / vertical offset, and rotation, as well as add random Gaussian noise to the image and randomly mask a certain position in the image. The sample finger vein image data can be expanded, and the expanded sample data and sample object labels can be used to train the initial twin network, thereby effectively improving the robustness of the twin network under various complex conditions and improving the accuracy and reliability in practical applications.
[0119] According to the embodiments of the present disclosure, the finger vein image data and the reference finger vein image data are processed in the same manner, thereby ensuring consistency in data processing, facilitating feature comparison between the output reference features and the target features, and improving recognition accuracy.
[0120] Figure 6 A schematic diagram of a network structure for performing multi-layer aggregation according to an embodiment of the present disclosure is shown schematically.
[0121] like Figure 6 As shown, the finger vein image data 610 and the reference finger vein data 620 are input into the twin network 630 together, and the same weight parameter 640 is used in the twin network 630 to obtain the target feature 660 and the reference feature 650 respectively. The target feature 660 and the reference feature 650 are matched to obtain the matching degree, and the recognition result is obtained according to the matching degree.
[0122] Based on the above object recognition method, the present disclosure also provides an object recognition device. Figure 7 The device is described in detail.
[0123] Figure 7 The structure block diagram of the object recognition device according to the embodiment of the present disclosure is schematically shown.
[0124] like Figure 7 As shown, the object recognition device 700 of this embodiment includes an extraction module 710 , an aggregation module 720 and a recognition module 730 .
[0125] The extraction module is used to perform feature extraction at different levels on the finger vein image data of the target object to obtain an initial feature map sequence representing different feature scales. The finger vein image data is collected under infrared irradiation, and the initial feature map sequence is obtained by sorting multiple initial feature maps obtained by feature extraction according to the feature scale. In one embodiment, the feature extraction module 710 can be used to perform the operation S210 described above, which will not be repeated here.
[0126] The processing module performs multi-level aggregation processing on the initial feature map sequence to obtain a target feature map. Among them, the aggregation processing of each level includes the following operations: feature preprocessing is performed on the aggregated feature map of the previous level to obtain a feature map to be aggregated with the same feature map size as the adjacent aggregated feature map of the previous level. The aggregated feature map of the previous level and the adjacent aggregated feature map of the previous level are respectively obtained through the aggregation processing of the previous level. Aggregation processing is performed on the feature map to be aggregated of the previous level and the adjacent aggregated feature map of the previous level to obtain the aggregated feature map of the current level. In one embodiment, the processing module 720 can be used to perform the operation S220 described above, which will not be repeated here.
[0127] The recognition module is used to recognize the target object based on the target feature map and the reference feature map of the target object to obtain a recognition result. In one embodiment, the recognition module 730 can be used to perform the operation S230 described above, which will not be repeated here.
[0128] According to an embodiment of the present disclosure, the processing module 720 includes a first processing unit.
[0129] The first processing unit is used to perform feature extraction processing and downsampling processing on the aggregated feature graph of the previous level in sequence when the aggregated feature graph of the previous level is at the first position in the aggregated feature graph sequence of the previous level. The aggregated feature graph sequence of the previous level is obtained by sorting multiple aggregated feature graphs of the previous level according to the feature scale.
[0130] According to an embodiment of the present disclosure, the processing module 720 further includes a second processing unit.
[0131] The second processing unit is used to downsample the aggregated feature map of the previous layer when the aggregated feature map of the previous layer is between the first and the last part of the aggregated feature map sequence of the previous layer.
[0132] According to an embodiment of the present disclosure, the processing module 720 further includes a third processing unit and a fourth processing unit.
[0133] A third processing unit is used for performing channel attention weighted processing on the adjacent aggregated feature map of the previous level when the adjacent aggregated feature map of the previous level is the initial feature map, so as to obtain the adjacent feature map to be aggregated of the previous level;
[0134] The fourth processing unit is used to aggregate the feature map to be aggregated at the previous level and the adjacent feature map to be aggregated at the previous level to obtain an aggregated feature map at the current level.
[0135] According to an embodiment of the present disclosure, the processing module 720 further includes a fifth processing unit and a sixth processing unit.
[0136] A fifth processing unit, configured to cascade the to-be-aggregated feature map of the previous layer and the adjacent aggregated feature map of the previous layer according to channels to obtain a densely connected aggregated feature map;
[0137] The sixth processing unit is used to obtain an aggregated feature map of the current level based on the densely connected aggregated feature map.
[0138] According to an embodiment of the present disclosure, the sixth processing unit includes a first processing sub-unit.
[0139] The first processing subunit is used to perform dimensionality reduction and downsampling processing on the densely connected aggregated feature map in sequence when the aggregated feature map of the previous level is between the first and the last of the aggregated feature map sequence of the previous level to obtain the aggregated feature map of the current level.
[0140] According to an embodiment of the present disclosure, the identification module 730 includes a matching unit and a determination unit.
[0141] The matching unit is used to match the target feature map with the reference feature map to obtain a matching degree, wherein the reference feature map is obtained by processing the reference finger vein image data of the target object in the same manner as that of the target feature map.
[0142] The determination unit is used to obtain a recognition result based on the matching degree.
[0143] According to an embodiment of the present disclosure, any multiple modules of the extraction module 710, the aggregation module 720 and the identification module 730 can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the extraction module 710, the aggregation module 720 and the identification module 730 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware and firmware or in any appropriate combination of any of them. Alternatively, at least one of the extraction module 710, the aggregation module 720 and the identification module 730 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding function can be performed.
[0144] Figure 8 A block diagram of an electronic device suitable for implementing the object method according to an embodiment of the present disclosure is schematically shown.
[0145] like Figure 8 As shown, the electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage part 808 to a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include an onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0146] In RAM803, various programs and data required for the operation of electronic device 800 are stored. Processor 801, ROM802 and RAM803 are connected to each other via bus 804. Processor 801 performs various operations of the method flow according to the embodiment of the present disclosure by executing the program in ROM802 and / or RAM803. It should be noted that the program can also be stored in one or more memories other than ROM802 and RAM803. Processor 801 can also perform various operations of the method flow according to the embodiment of the present disclosure by executing the program stored in one or more memories.
[0147] According to an embodiment of the present disclosure, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the input / output (I / O) interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed, so that the computer program read therefrom is installed into the storage portion 808 as needed.
[0148] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.
[0149] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM802 and / or RAM803 described above and / or one or more memories other than ROM802 and RAM803.
[0150] The embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the object recognition method provided by the embodiment of the present disclosure.
[0151] The above functions defined in the system / device of the embodiment of the present disclosure are executed when the computer program is executed by the processor 801. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0152] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 809, and / or installed from a removable medium 811. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0153] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, the above functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, means, module, unit, etc. described above can be implemented by a computer program module.
[0154] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect through the Internet).
[0155] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0156] It will be appreciated by those skilled in the art that the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways. All of these combinations and / or combinations fall within the scope of the present disclosure.
[0157] The embodiments of the present disclosure are described above. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments are described above, this does not mean that the measures in the various embodiments cannot be used in combination to advantage. Without departing from the scope of the present disclosure, those skilled in the art may make a variety of substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. An object recognition method, characterized in that: The method comprises: Performing feature extraction at different levels on the finger vein image data of the target object to obtain an initial feature map sequence representing different feature scales; wherein the finger vein image data is collected under infrared irradiation, and the initial feature map sequence is obtained by sorting a plurality of initial feature maps obtained by feature extraction according to the feature scale; The initial feature map sequence is subjected to multi-level aggregation processing to obtain a target feature map, wherein the aggregation processing at each level includes the following operations: Performing feature preprocessing on the aggregated feature map of the previous level to obtain a feature map to be aggregated having the same feature map size as that of an adjacent aggregated feature map of the previous level, wherein the aggregated feature map of the previous level and the adjacent aggregated feature map of the previous level are respectively obtained through the aggregation processing of the previous level; Aggregate the feature map to be aggregated at the previous level and the adjacent aggregated feature map at the previous level to obtain an aggregated feature map at the current level; Based on the target feature map and the reference feature map of the target object, the target object is identified to obtain a recognition result.
2. The method according to claim 1, wherein: The feature preprocessing of the aggregated feature map of the previous level includes: When the aggregated feature map of the previous level is at the first place in the aggregated feature map sequence of the previous level, feature extraction processing and downsampling processing are performed on the aggregated feature map of the previous level in sequence, wherein the aggregated feature map sequence of the previous level is obtained by sorting multiple aggregated feature maps of the previous level according to the feature scale.
3. The method according to claim 2, wherein: The performing feature preprocessing on the aggregated feature map of the previous level also includes: When the aggregated feature map of the previous level is between the first and last part of the aggregated feature map sequence of the previous level, the aggregated feature map of the previous level is downsampled.
4. The method according to claim 1, wherein: The step of aggregating the feature map to be aggregated at the previous level and the adjacent aggregated feature map at the previous level to obtain the aggregated feature map at the current level includes: When the adjacent aggregated feature map of the previous level is an initial feature map, performing channel attention weighted processing on the adjacent aggregated feature map of the previous level to obtain an adjacent feature map to be aggregated of the previous level; The feature map to be aggregated at the previous level and the adjacent feature map to be aggregated at the previous level are aggregated to obtain the aggregated feature map of the current level.
5. The method according to claim 1, wherein: The step of aggregating the feature map to be aggregated at the previous level and the adjacent aggregated feature map at the previous level to obtain the aggregated feature map at the current level further includes: Cascading the feature map to be aggregated at the previous level and the adjacent aggregated feature map at the previous level according to channels to obtain a densely connected aggregated feature map; Based on the densely connected aggregated feature map, the aggregated feature map of the current level is obtained.
6. The method according to claim 5, wherein: The step of obtaining the aggregated feature map of the current level based on the densely connected aggregated feature map includes: When the aggregate feature map of the previous level is between the first and last part of the aggregate feature map sequence of the previous level, the densely connected aggregate feature map is sequentially subjected to dimensionality reduction and downsampling to obtain an aggregate feature map of the current level.
7. The method according to any one of claims 1 to 6, characterized in that: The identifying the target object based on the target feature and the reference feature map of the target object to obtain an identification result includes: Matching the target feature map with the reference feature map to obtain a matching degree, wherein the reference feature map is obtained by processing the reference finger vein image data of the target object in the same manner as that of obtaining the target feature map; Based on the matching degree, a recognition result is obtained.
8. An object recognition device, characterized in that: The device comprises: An extraction module is used to perform feature extraction at different levels on the finger vein image data of the target object to obtain an initial feature map sequence representing different feature scales; wherein the finger vein image data is collected under infrared irradiation, and the initial feature map sequence is obtained by sorting a plurality of initial feature maps obtained by feature extraction according to the feature scale; The processing module performs multi-level aggregation processing on the initial feature map sequence to obtain a target feature map, wherein the aggregation processing at each level includes the following operations: Performing feature preprocessing on the aggregated feature map of the previous level to obtain a feature map to be aggregated having the same feature map size as that of an adjacent aggregated feature map of the previous level, wherein the aggregated feature map of the previous level and the adjacent aggregated feature map of the previous level are respectively obtained through the aggregation processing of the previous level; Aggregate the feature map to be aggregated at the previous level and the adjacent aggregated feature map at the previous level to obtain an aggregated feature map at the current level; The recognition module is used to recognize the target object based on the target feature map and the reference feature map of the target object to obtain a recognition result.
9. An electronic device, comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.