Feature Processing Method, Device, Product, Medium and Equipment

By dividing the image feature extraction process into global features and local features, and combining the two to generate classified indication features, the problem of feature information loss in the prior art is solved, and the feature extraction accuracy and efficiency are improved.

CN114972775BActive Publication Date: 2025-06-20JINGDONG TECH HLDG CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210635593.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-06
Publication Date
2025-06-20
Estimated Expiration
2042-06-06

AI Technical Summary

Technical Problem

The prior art reduces the amount of self-attention calculation through downsampling processing during image feature extraction, but this will lead to the loss of feature information, thereby reducing the accuracy of feature extraction.

Method used

The feature extraction process is divided into global feature extraction and local feature extraction, and local features are calculated based on global features, thereby generating accurate classification indicator features to avoid the loss of feature information.

Benefits of technology

The accuracy of image feature extraction is improved, the complexity of attention calculation is reduced, and global feature extraction and local feature extraction can be performed in parallel, improving the efficiency of overall feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972775B_ABST
    Figure CN114972775B_ABST
Patent Text Reader

Abstract

The present application provides a feature processing method, a feature processing device, a computer program product, a computer-readable storage medium and an electronic device, which relate to the field of computer technology. The method includes: obtaining a sample local feature and a sample global feature of an image to be classified; generating a reference global feature corresponding to the sample global feature; generating a reference local feature according to the reference global feature and the sample local feature; and determining a classification indication feature corresponding to the image to be classified based on the reference global feature and the reference local feature. In this way, the feature extraction process can be divided into global feature extraction and local feature extraction, and local features are calculated in combination with global features, so as to determine the classification indication feature corresponding to the image to be classified according to accurate global features and local features, which can avoid losing features in the process of extracting image features and improve the accuracy of image feature extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more particularly, to a feature processing method, a feature processing device, a computer program product, a computer-readable storage medium, and an electronic device. Background Art

[0002] In the field of artificial intelligence recognition such as image recognition and text recognition, it is usually necessary to use a network architecture based on Transformer blocks for feature extraction, so as to recognize images / texts based on the extracted features. A Transformer block usually includes multiple self-attention calculation modules. For feature extraction, it takes a long time. To solve this problem, existing solutions usually adopt a downsampling processing method to reduce the self-attention calculation amount. However, this method will lose a lot of feature information during the calculation process, thereby resulting in a low accuracy of the final feature extraction result.

[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0004] The purpose of the present application is to provide a feature processing method, a feature processing device, a computer program product, a computer-readable storage medium, and an electronic device, which can divide the feature extraction process into global feature extraction and local feature extraction, combine the global features to calculate the local features, and thus determine the classification indication features corresponding to the image to be classified according to the accurate global features and local features, so as to avoid losing features during the process of extracting image features and improve the accuracy of image feature extraction.

[0005] Other features and advantages of the present application will become apparent through the following detailed description, or be learned in part through the practice of the present application.

[0006] According to one aspect of the present application, a feature processing method is provided, and the method includes:

[0007] Obtain the sample local features and sample global features of the image to be classified;

[0008] Generate a reference global feature corresponding to the sample global feature;

[0009] Generate a reference local feature according to the reference global feature and the sample local feature;

[0010] Determine the classification indication feature corresponding to the image to be classified based on the reference global feature and the reference local feature.

[0011] In an exemplary embodiment of the present application, generating a reference global feature corresponding to a sample global feature includes:

[0012] Extracting a first intermediate global feature of the sample global feature based on a first global normalization network, a first global multi-head network, and a second global normalization network;

[0013] Obtaining a local normalization feature corresponding to the sample local feature, and inputting the first intermediate feature and the local normalization feature into a second global multi-head network, so that the second global multi-head network generates a second intermediate global feature;

[0014] Generating a feature corresponding to the second intermediate global feature based on a third global normalization network and a global feed-forward network as the reference global feature of the sample global feature;

[0015] Wherein, the first global normalization network, the second global normalization network, and the third global normalization network correspond to different network parameters; the first global multi-head network and the second global multi-head network correspond to different network parameters.

[0016] In an exemplary embodiment of the present application, generating a reference local feature according to the reference global feature and the sample local feature includes:

[0017] Inputting the local normalization feature and the reference global feature into a local multi-head network, so that the local multi-head network generates a first intermediate local feature;

[0018] Inputting the first intermediate local feature and the local normalization feature into a local normalization network, so that the local normalization network generates a second intermediate local feature;

[0019] Triggering a local feed-forward network to generate a reference local feature based on the second intermediate local feature and the first intermediate local feature.

[0020] In an exemplary embodiment of the present application, determining a classification indication feature corresponding to an image to be classified based on the reference global feature and the reference local feature includes:

[0021] Fusing the reference global feature and the reference local feature to obtain a feature to be split;

[0022] Splitting the feature to be split into a target global feature and a target local feature;

[0023] Determining a classification indication feature corresponding to the image to be classified according to the target global feature and the target local feature.

[0024] In an exemplary embodiment of the present application, fusing the reference global feature and the reference local feature to obtain a feature to be split includes:

[0025] Fuse the reference global feature and the reference local feature to obtain a first fusion result;

[0026] Perform layer normalization on the first fusion result to obtain a second fusion result;

[0027] Generate a self-attention fusion feature corresponding to the second fusion result;

[0028] Generate a feature to be split based on the self-attention fusion feature and the first fusion result.

[0029] In an exemplary embodiment of the present application, determining a classification indication feature corresponding to an image to be classified according to a target global feature and a target local feature includes:

[0030] Generate a first feature to be processed corresponding to the target global feature based on a global feature processing network;

[0031] Generate a second feature to be processed corresponding to the target local feature based on a local feature processing network;

[0032] Generate a classification indication feature corresponding to the image to be classified according to the first feature to be processed and the second feature to be processed.

[0033] In an exemplary embodiment of the present application, the global feature processing network includes a semantic normalization network and a semantic feed-forward network. Generating a first feature to be processed corresponding to the target global feature based on the global feature processing network includes:

[0034] Perform normalization processing on the target global feature through the semantic normalization network to obtain a semantic normalization result;

[0035] Generate a semantic comprehensive feature corresponding to the semantic normalization result through the semantic feed-forward network;

[0036] Fuse the semantic comprehensive feature and the target global feature to obtain a first feature to be processed corresponding to the target global feature.

[0037] In an exemplary embodiment of the present application, the local feature processing network includes a pixel normalization network and a pixel feed-forward network. Generating a second feature to be processed corresponding to the target local feature based on the local feature processing network includes:

[0038] Perform layer normalization processing on the target local feature through the pixel normalization network to obtain a pixel normalization result;

[0039] Generate a pixel comprehensive feature corresponding to the pixel normalization result through the pixel feed-forward network;

[0040] Fuse the pixel comprehensive feature and the target local feature to obtain a second feature to be processed corresponding to the target local feature.

[0041] In an exemplary embodiment of the present application, generating a classification indication feature corresponding to an image to be classified according to a first feature to be processed and a second feature to be processed includes:

[0042] Performing pooling processing on the first feature to be processed and the second feature to be processed to obtain a classification indication feature.

[0043] In an exemplary embodiment of the present application, after determining a classification indication feature corresponding to an image to be classified based on a reference global feature and a reference local feature, the above method further includes:

[0044] Determining the category corresponding to the image to be classified through the classification indication feature.

[0045] According to an aspect of the present application, there is provided a feature processing device, including:

[0046] A feature acquisition unit for acquiring a sample local feature and a sample global feature of an image to be classified;

[0047] A feature generation unit for generating a reference global feature corresponding to the sample global feature;

[0048] The feature generation unit is further configured to generate a reference local feature according to the reference global feature and the sample local feature;

[0049] A feature determination unit for determining a classification indication feature corresponding to the image to be classified based on the reference global feature and the reference local feature.

[0050] In an exemplary embodiment of the present application, the feature generation unit generating a reference global feature corresponding to the sample global feature includes:

[0051] Extracting a first intermediate global feature of the sample global feature based on a first global normalization network, a first global multi-head network, and a second global normalization network;

[0052] Obtaining a local normalization feature corresponding to the sample local feature, and inputting the first intermediate feature and the local normalization feature into a second global multi-head network so that the second global multi-head network generates a second intermediate global feature;

[0053] Generating a feature corresponding to the second intermediate global feature based on a third global normalization network and a global feed-forward network as the reference global feature of the sample global feature;

[0054] Wherein, the first global normalization network, the second global normalization network, and the third global normalization network correspond to different network parameters; the first global multi-head network and the second global multi-head network correspond to different network parameters.

[0055] In an exemplary embodiment of the present application, the feature generation unit generates a reference local feature according to a reference global feature and a sample local feature, including:

[0056] Input the local normalized feature and the reference global feature into a local multi-head network, so that the local multi-head network generates a first intermediate local feature;

[0057] Input the first intermediate local feature and the local normalized feature into a local normalization network, so that the local normalization network generates a second intermediate local feature;

[0058] Trigger the local feed-forward network to generate a reference local feature based on the second intermediate local feature and the first intermediate local feature.

[0059] In an exemplary embodiment of the present application, the feature determination unit determines a classification indication feature corresponding to the image to be classified based on the reference global feature and the reference local feature, including:

[0060] Fuse the reference global feature and the reference local feature to obtain a feature to be split;

[0061] Split the feature to be split into a target global feature and a target local feature;

[0062] Determine a classification indication feature corresponding to the image to be classified according to the target global feature and the target local feature.

[0063] In an exemplary embodiment of the present application, the feature determination unit fuses the reference global feature and the reference local feature to obtain a feature to be split, including:

[0064] Fuse the reference global feature and the reference local feature to obtain a first fusion result;

[0065] Perform layer normalization on the first fusion result to obtain a second fusion result;

[0066] Generate a self-attention fusion feature corresponding to the second fusion result;

[0067] Generate a feature to be split based on the self-attention fusion feature and the first fusion result.

[0068] In an exemplary embodiment of the present application, the feature determination unit determines a classification indication feature corresponding to the image to be classified according to the target global feature and the target local feature, including:

[0069] Generate a first feature to be processed corresponding to the target global feature based on a global feature processing network;

[0070] Generate a second feature to be processed corresponding to the target local feature based on a local feature processing network;

[0071] Generate a classification indication feature corresponding to the image to be classified according to the first feature to be processed and the second feature to be processed.

[0072] In an exemplary embodiment of the present application, the global feature processing network includes a semantic normalization network and a semantic feed-forward network. The feature determination unit generates a first feature to be processed corresponding to the target global feature based on the global feature processing network, including:

[0073] Perform normalization processing on the target global feature through the semantic normalization network to obtain a semantic normalization result;

[0074] Generate a semantic synthesis feature corresponding to the semantic normalization result through the semantic feed-forward network;

[0075] Fuse the semantic synthesis feature and the target global feature to obtain a first feature to be processed corresponding to the target global feature.

[0076] In an exemplary embodiment of the present application, the local feature processing network includes a pixel normalization network and a pixel feed-forward network. The feature determination unit generates a second feature to be processed corresponding to the target local feature based on the local feature processing network, including:

[0077] Perform layer normalization processing on the target local feature through the pixel normalization network to obtain a pixel normalization result;

[0078] Generate a pixel synthesis feature corresponding to the pixel normalization result through the pixel feed-forward network;

[0079] Fuse the pixel synthesis feature and the target local feature to obtain a second feature to be processed corresponding to the target local feature.

[0080] In an exemplary embodiment of the present application, the feature determination unit generates a classification indication feature corresponding to the image to be classified according to the first feature to be processed and the second feature to be processed, including:

[0081] Perform pooling processing on the first feature to be processed and the second feature to be processed to obtain a classification indication feature.

[0082] In an exemplary embodiment of the present application, the above device includes:

[0083] A category determination unit, configured to determine the category corresponding to the image to be classified through the classification indication feature after the feature determination unit determines the classification indication feature corresponding to the image to be classified based on the reference global feature and the reference local feature.

[0084] According to one aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various optional implementations described above.

[0085] According to one aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of the above is implemented.

[0086] According to one aspect of the present application, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the method described in any one of the above by executing the executable instructions.

[0087] The exemplary embodiments of the present application may have the following partial or all beneficial effects:

[0088] In the feature processing method provided in an exemplary implementation manner of the present application, a sample local feature and a sample global feature of an image to be classified can be obtained; a reference global feature corresponding to the sample global feature is generated; a reference local feature is generated according to the reference global feature and the sample local feature; and a classification indication feature corresponding to the image to be classified is determined based on the reference global feature and the reference local feature. In this way, the feature extraction process can be divided into global feature extraction and local feature extraction, and the local feature is calculated in combination with the global feature, so that the classification indication feature corresponding to the image to be classified is determined according to the accurate global feature and local feature, which can avoid losing features in the process of extracting image features and improve the accuracy of image feature extraction. In addition, based on global feature extraction and local feature extraction, it is helpful to reduce the computational amount of the attention mechanism on the extraction paths corresponding to global feature extraction and local feature extraction respectively, thereby reducing the complexity of attention calculation. In addition, global feature extraction and local feature extraction can be executed in parallel to improve the overall efficiency of feature extraction.

[0089] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Description of the Drawings

[0090] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0091] Figure 1 A schematic diagram showing an exemplary system architecture of a feature processing method and a feature processing device to which embodiments of the present application can be applied;

[0092] Figure 2 A flowchart schematically showing a feature processing method according to an embodiment of the present application;

[0093] Figure 3 A flowchart schematically showing a feature processing method according to another embodiment of the present application;

[0094] Figure 4 A schematic diagram showing the feature extraction network structure of the existing solution;

[0095] Figure 5 A network architecture diagram schematically showing a network architecture for implementing the feature processing method of the present application;

[0096] Figure 6 A network architecture diagram schematically showing another network architecture for implementing the feature processing method of the present application;

[0097] Figure 7 A network architecture diagram schematically showing yet another network architecture for implementing the feature processing method of the present application;

[0098] Figure 8 A block diagram schematically showing the structure of a feature processing device according to an embodiment of the present application;

[0099] Figure 9 A schematic diagram showing the structure of a computer system of an electronic device suitable for implementing embodiments of the present application. Detailed implementation manners

[0100] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will recognize that the technical solutions of the present application can be practiced without one or more of the specific details, or can be implemented using other methods, components, devices, steps, etc. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present application.

[0101] In addition, the accompanying drawings are only schematic illustrations of the present application and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0102] Please refer to Figure 1 , Figure 1 which shows a schematic diagram of the system architecture of an exemplary application environment where a feature processing method and a feature processing device according to embodiments of the present application can be applied. As Figure 1 shown, the system architecture 100 may include one or more of the terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105.

[0103] The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The terminal devices 101, 102, 103 may be devices that provide voice and / or data connectivity to users, handheld devices with wireless connection capabilities, or other processing devices connected to a wireless modem. The wireless terminal may communicate with one or more core networks via the RAN. The wireless terminal may be a user equipment (UE), a handheld terminal, a laptop computer, a subscriber unit, a cellular phone, a smart phone, a wireless data card, a personal digital assistant (PDA) computer, a tablet computer, a wireless modem, a handheld device, a laptop computer, a cordless phone, or a wireless local loop (WLL) station, a machine type communication (MTC) terminal, or other devices that can access the network. The terminal and the access network device communicate with each other using a certain air interface technology (for example, 3GPP access technology or non-3GPP access technology). It should be understood that Figure 1The numbers of the terminal devices, networks, and servers therein are merely illustrative. According to actual requirements, there can be any number of terminal devices, networks, and servers. For example, server 105 can be a server cluster composed of multiple servers, etc.

[0104] The path planning method for multi-node networking provided by the embodiments of the present application can be executed by server 105. Correspondingly, the path planning device for multi-node networking is generally set in server 105. However, those skilled in the art can easily understand that the path planning method for multi-node networking provided by the embodiments of the present application can also be executed by terminal devices 101, 102, or 103. Correspondingly, the path planning device for multi-node networking can also be set in terminal devices 101, 102, or 103. No special limitation is made in this exemplary embodiment. For example, in an exemplary embodiment, server 105 can obtain the sample local features and sample global features of the image to be classified; generate the reference global features corresponding to the sample global features; generate the reference local features according to the reference global features and the sample local features; and determine the classification indication features corresponding to the image to be classified based on the reference global features and the reference local features.

[0105] Please refer to Figure 2 , Figure 2 which schematically shows a flowchart of the feature processing method according to an embodiment of the present application. As Figure 2 shown, the feature processing method can include: step S210 to step S240.

[0106] Step S210: Obtain the sample local features and sample global features of the image to be classified.

[0107] Step S220: Generate the reference global features corresponding to the sample global features.

[0108] Step S230: Generate the reference local features according to the reference global features and the sample local features.

[0109] Step S240: Determine the classification indication features corresponding to the image to be classified based on the reference global features and the reference local features.

[0110] Next, the above steps of this exemplary embodiment will be described in more detail.

[0111] In the process of image feature extraction, network designs based on Convolutional Neural Network (CNN) and those based on Transformer blocks are usually used. CNN mainly uses stacked different convolutional kernels to extract features from local regions of pictures, and gradually expands the receptive field of the convolutional neural network in multiple stages through the method of downsampling with a pyramid structure, so as to achieve the extraction of global features of pictures. Transformer mainly relies on the self-attention mechanism between different image patches for feature fusion. Among them, CNN can only extract local information of pictures in the early stage of feature extraction and cannot directly process the global information of pictures. The computational cost of Transformer is large and the efficiency is low. Image feature extraction based on the structures of CNN and Transformer is prone to losing feature information and the efficiency is also relatively low.

[0112] To solve this problem, existing solutions use the PVT network for linear spatial reduction attention, and the downsampling process based on the PVT network reduces the spatial scale of keys and values; or, a local grouped self-attention layer is added before the PVT network to further enhance the feature representation through intra-region interaction. However, the above methods both need to rely on feature map downsampling and are still prone to causing excessive loss of global information of pictures in the early stage of the network.

[0113] Based on this, this application proposes that the feature extraction process can be divided into global feature extraction and local feature extraction, combining global features to calculate local features, so as to determine the classification indication features corresponding to the image to be classified according to accurate global and local features. This can avoid losing features in the process of extracting image features and can improve the accuracy of image feature extraction. In addition, based on global feature extraction and local feature extraction, it can help reduce the computational cost of the attention mechanism on the extraction paths corresponding to global feature extraction and local feature extraction respectively, thus reducing the complexity of attention calculation. In addition, global feature extraction and local feature extraction can be executed in parallel to improve the overall efficiency of feature extraction.

[0114] In step S210, the sample local features and sample global features of the image to be classified are obtained.

[0115] Specifically, the image to be classified can be a picture containing a commodity, a picture containing facial features, or a picture containing text, etc., which is not limited in the embodiments of this application. For example, when the image to be classified contains a commodity, the classification indication features extracted by the method shown in the embodiments of this application can be used to achieve more accurate commodity recognition. When this solution is applied to intelligent commodity sorting, accurate commodity sorting can be achieved through accurate commodity recognition.

[0116] In addition, the local features of the sample are used to characterize the local details of the image to be classified; the global features of the sample are used to characterize the global semantics of the image to be classified.

[0117] It should be noted that, optionally, the image to be classified can be the original image that needs to be recognized, or a certain image block in the picture. That is, before step S210, the picture can be divided into multiple image blocks. For each image block, the solution of the embodiment of the present application can be executed to obtain the classification indication features corresponding to each image block, and then the classification indication features of each image block are fused to obtain the classification indication features for characterizing the entire picture.

[0118] In step S220, a reference global feature corresponding to the sample global feature is generated.

[0119] Specifically, generating a reference global feature corresponding to the sample global feature is specifically implemented as: generating a reference global feature corresponding to the sample global feature based on the semantic-level global feature extraction path; where the semantic-level global feature extraction path includes the semantic structure block in the dual-stream structure block, and the dual-stream structure block also includes a pixel structure block. The pixel structure block belongs to the pixel-level local feature extraction path. The global and local feature extraction performed through the semantic-level global feature extraction path and the pixel-level local feature extraction path can simplify the self-attention calculation of each path, reduce the self-attention calculation amount, and improve the efficiency of feature extraction for the overall image.

[0120] In addition, the reference global feature and the sample global feature are different features. The reference global feature and the sample global feature can be represented as a feature vector or a feature matrix.

[0121] As an optional embodiment, generating a reference global feature corresponding to the sample global feature includes: extracting a first intermediate global feature of the sample global feature based on the first global normalization network, the first global multi-head network, and the second global normalization network; obtaining the local normalization feature corresponding to the sample local feature, and inputting the first intermediate feature and the local normalization feature into the second global multi-head network so that the second global multi-head network generates a second intermediate global feature; generating a feature corresponding to the second intermediate global feature based on the third global normalization network and the global feed-forward network as the reference global feature of the sample global feature; where the first global normalization network, the second global normalization network, and the third global normalization network correspond to different network parameters; the first global multi-head network and the second global multi-head network correspond to different network parameters. In this way, it is possible to extract the reference global feature of the sample global feature based on multiple global normalization networks, global multi-head networks, and global feed-forward networks to implement the multi-head attention calculation for the sample global feature, thereby improving the accuracy of the reference global feature.

[0122] Among them, the first intermediate global feature for extracting the sample global feature based on the first global normalization network, the first global multi-head network, and the second global normalization network includes: taking the sample global feature z l Input it into the first global normalization network Layer Norm in the semantic-level global feature extraction path, so that the first global normalization network is based on the expression Calculate the feature corresponding to the sample global feature z l Namely z l ∈R m*d , where l is used to represent the l-th two-stream structure block, R m*d Is used to represent the semantic features corresponding to the entire image to be recognized, n is used to represent the length of the feature sequence, and d is used to represent the word embedding size corresponding to each position in the sequence; furthermore, the feature Respectively input into each head self-attention network in the first global multi-head network Multi-Head Attention in the semantic-level global feature extraction path, and fuse the features generated by each head attention network to obtain the self-attention feature MHA Furthermore, the self-attention feature And the sample global feature z l Input into the second global normalization network Layer Norm in the semantic-level global feature extraction path, so that the second global normalization network is based on the expression Calculate the first intermediate global feature z l . Among them, the computational complexity corresponding to the semantic-level global feature extraction path can be expressed as O(nmd + m 2 d); where m is the number of semantic tokens (such as tokens).

[0123] Based on this, obtaining the local normalization feature corresponding to the sample local feature includes: inputting the sample local feature x l Into the normalization network Layer Norm in the pixel-level local feature extraction path, so that the normalization network is based on the expression Calculate the local normalization feature corresponding to the sample local feature x l Namely

[0124] Furthermore, input the first intermediate feature and the local normalization feature into the second global multi-head network, so that the second global multi-head network generates the second intermediate global feature, including: normalizing the first intermediate global feature z′ l To obtain the feature LN(z′ l ); The feature LN(z′ l ) And the local normalization feature Input the self-attention networks of each head in the second global multi-head network Multi-HeadAttention, and fuse the features generated by each head attention network to obtain the second intermediate global feature

[0125] Furthermore, generate a reference global feature corresponding to the second intermediate global feature as the sample global feature based on the third global normalization network and the global feed-forward network, including: taking the second intermediate global feature and the input first intermediate global feature z′ l into the third global normalization network Layer Norm, so that the third global normalization network calculates the feature based on the expression and performs normalization processing on the feature to obtain Taking as the input of the global feed-forward network Feed Forward, so that the global feed-forward network processes into Taking and the feature substitute them into the expression to calculate the reference global feature z of the sample global feature l+1 .

[0126] It should be noted that the above LayerNorm is used to perform layer normalization on the hidden layer, that is, to normalize the inputs of all neurons in a certain layer. Multi-Head Attention can be understood as a combination of multiple self-attention (Self-Attention).

[0127] Among them, the network parameters corresponding to the first global normalization network, the second global normalization network, and the third global normalization network may include weights, bias terms, etc., which are not limited in the embodiments of the present application. The same applies to the network parameters corresponding to the first global multi-head network and the second global multi-head network.

[0128] In step S230, generate a reference local feature according to the reference global feature and the sample local feature.

[0129] Specifically, the reference global feature and the reference local feature can be understood as intermediate features used to calculate the final classification indication feature.

[0130] As an alternative embodiment, generating reference local features based on reference global features and sample local features includes: inputting local normalized features and reference global features into a local multi-head network so that the local multi-head network generates first intermediate local features; inputting the first intermediate local features and local normalized features into a local normalization network so that the local normalization network generates second intermediate local features; triggering a local feed-forward network to generate reference local features based on the second intermediate local features and the first intermediate local features. In this way, reference local features corresponding to sample local features can be calculated based on reference global features. By combining the feature extraction of reference global features and sample local features, the loss of global features during feature extraction can be avoided. Combining reference global features and sample local features can also calculate reference local features with higher accuracy, and at the same time, the difficulty of extracting fine local features can be reduced.

[0131] Among them, inputting local normalized features and reference global features into a local multi-head network so that the local multi-head network generates first intermediate local features includes: based on the expression normalizing the reference global feature z l+1 to obtain the feature Inputting the feature and the local normalized feature into each head self-attention network (Self-Attention) in the local multi-head network Multi-Head Attention, and fusing the features generated by each head attention network to obtain the first intermediate local feature

[0132] Based on this, inputting the first intermediate local feature and local normalized features into a local normalization network so that the local normalization network generates second intermediate local features includes: inputting the first intermediate local feature and the local normalized feature x l into the local normalization network LayerNorm so that the local normalization network generates the second intermediate local feature x' based on the expression l .

[0133] Furthermore, triggering a local feed-forward network to generate reference local features based on the second intermediate local feature and the first intermediate local feature includes: normalizing the second intermediate local feature x' l to obtain Inputting into the local feed-forward network Feed Forward so that the local feed-forward network Feed Forward performs feature processing on to obtain Inputting and the second intermediate local feature x' l Substitute into the expression to calculate the reference local feature x l+1 .

[0134] In step S240, a classification indication feature corresponding to the image to be classified is determined based on the reference global feature and the reference local feature.

[0135] Specifically, the classification indication feature is used to describe the image to be classified in the form of a vector / matrix.

[0136] As an optional embodiment, determining the classification indication feature corresponding to the image to be classified based on the reference global feature and the reference local feature includes: fusing the reference global feature and the reference local feature to obtain a feature to be split; splitting the feature to be split into a target global feature and a target local feature; determining the classification indication feature corresponding to the image to be classified according to the target global feature and the target local feature. In this way, feature fusion of the reference global feature and the reference local feature can be realized to determine the feature to be split without missing feature information. Furthermore, the feature to be split can be split, and then global and local feature processing can be performed based on the splitting result, so as to obtain the classification indication feature, ensuring that the classification indication feature contains the local and global features that have not been lost and improving the accuracy of the classification indication feature.

[0137] As an optional embodiment, fusing the reference global feature and the reference local feature to obtain a feature to be split includes: fusing the reference global feature and the reference local feature to obtain a first fusion result; performing layer normalization processing on the first fusion result to obtain a second fusion result; generating a self-attention fusion feature corresponding to the second fusion result; generating the feature to be split based on the self-attention fusion feature and the first fusion result. In this way, the reference global feature and the reference local feature can be fused based on layer normalization processing and self-attention calculation, ensuring that the feature to be split obtained by fusion can restore the feature information of the entire image to be classified to the greatest extent, thereby ensuring the accuracy of subsequent feature extraction calculations and reducing distortion.

[0138] Specifically, fusing the reference global feature and the reference local feature to obtain a first fusion result includes: taking the reference global feature x l+1 and the reference local feature z l+1 as inputs to the fusion layer Concat, so that the fusion layer Concat performs tensor connection on the reference global feature x l+1 and the reference local feature z l+1 to obtain a first fusion result x l+1 ||z l+1 .

[0139] Based on this, perform layer normalization on the first fusion result to obtain the second fusion result, including: taking the first fusion result x l+1 ||z l+1 as the input to the normalization network Layer Norm, so that the normalization network Layer Norm is based on the expression to perform layer normalization on the first fusion result x l+1 ||z l+1 and obtain the second fusion result

[0140] Specifically, generate self-attention fusion features corresponding to the second fusion result, including: taking the second fusion result as the input to each head self-attention network (Self-Attention) in the multi-head network Multi-Head Attention, and fusing the features generated by each head attention network to obtain the self-attention fusion features

[0141] Specifically, generate the features to be split based on the self-attention fusion features and the first fusion result, including: taking the self-attention fusion features and the first fusion result x l+1 ||z l+1 as the input to the split layer Split, so that the split layer Split is based on the expression to determine the features to be split (x′ l+1 , z′ l+1 ). Among them, the features to be split (x′ l , z′ l ) include the target global feature z′ l+1 and the target local feature x′ l+1 .

[0142] As an optional embodiment, determine the classification indication feature corresponding to the image to be classified according to the target global feature and the target local feature, including: generating the first feature to be processed corresponding to the target global feature based on the global feature processing network; generating the second feature to be processed corresponding to the target local feature based on the local feature processing network; generating the classification indication feature corresponding to the image to be classified according to the first feature to be processed and the second feature to be processed. In this way, the first feature to be processed and the second feature to be processed can be obtained based on the fused local feature processing path and global feature processing path, and the classification indication feature for accurately representing the image to be classified can be determined based on the first feature to be processed and the second feature to be processed.

[0143] As an alternative embodiment, the global feature processing network includes a semantic normalization network and a semantic feed-forward network. Generating a first feature to be processed corresponding to the target global feature based on the global feature processing network includes: performing normalization processing on the target global feature through the semantic normalization network to obtain a semantic normalization result; generating a semantic comprehensive feature corresponding to the semantic normalization result through the semantic feed-forward network; and fusing the semantic comprehensive feature and the target global feature to obtain the first feature to be processed corresponding to the target global feature. In this way, the first feature to be processed can be obtained based on the processing of the semantic normalization network and the semantic feed-forward network, which can reduce the probability of feature loss during the global feature processing process.

[0144] Specifically, performing normalization processing on the target global feature through the semantic normalization network to obtain a semantic normalization result includes: inputting the target global feature z′ l+1 into the semantic normalization network LayerNorm, so that the semantic normalization network Layer Norm performs normalization processing on the target global feature z′ l+1 to obtain a semantic normalization result LN(z′ l+1 ).

[0145] Specifically, generating a semantic comprehensive feature corresponding to the semantic normalization result through the semantic feed-forward network includes: inputting the semantic normalization result LN(z′ l+1 ) into the semantic feed-forward network Feed Forward, so that the semantic feed-forward network FeedForward generates a semantic comprehensive feature FFN l+1 corresponding to the semantic normalization result LN(z′ Z )(LN(z′ l+1 ).

[0146] Specifically, fusing the semantic comprehensive feature and the target global feature to obtain the first feature to be processed corresponding to the target global feature includes: substituting the semantic comprehensive feature FFN Z (LN(z′ l+1 )) and the target global feature z′ l+1 into the expression z l+2 = FFN Z (LN(z′ l+1 )) + z′ l+1 to calculate the first feature to be processed z l+2 .

[0147] As an alternative embodiment, the local feature processing network includes a pixel normalization network and a pixel feed-forward network. Generating a second feature to be processed corresponding to the target local feature based on the local feature processing network includes: performing layer normalization processing on the target local feature through the pixel normalization network to obtain a pixel normalization result; generating a pixel comprehensive feature corresponding to the pixel normalization result through the pixel feed-forward network; and fusing the pixel comprehensive feature and the target local feature to obtain a second feature to be processed corresponding to the target local feature. In this way, the first feature to be processed can be obtained through the pixel normalization network and the pixel feed-forward network, reducing the probability of feature loss during the local feature processing.

[0148] Specifically, performing layer normalization processing on the target local feature through the pixel normalization network to obtain a pixel normalization result includes: inputting the target local feature x' l+1 into the pixel normalization network Layer Norm, so that the pixel normalization network Layer Norm performs normalization processing on the target local feature x' l+1 to obtain a pixel normalization result LN(x' l+1 ).

[0149] Specifically, generating a pixel comprehensive feature corresponding to the pixel normalization result through the pixel feed-forward network includes: inputting the pixel normalization result LN(x' l+1 ) into the pixel feed-forward network Feed Forward, so that the pixel feed-forward network FeedForward generates a pixel comprehensive feature FFN l+1 corresponding to the pixel normalization result LN(x' Z (LN(x' l+1 )).

[0150] Specifically, fusing the pixel comprehensive feature and the target local feature to obtain a second feature to be processed corresponding to the target local feature includes: substituting the pixel comprehensive feature FFN + (LN(x' l+1 )) and the target local feature x' l+1 into the expression x l+2 = FFN + (LN(x' l+1 )) + x' l+1 to calculate the second feature to be processed x l+2 corresponding to the target local feature.

[0151] As an alternative embodiment, generating a classification indication feature corresponding to the image to be classified according to the first feature to be processed and the second feature to be processed includes: performing pooling processing on the first feature to be processed and the second feature to be processed to obtain a classification indication feature. In this way, the classification indication feature can be obtained through pooling processing, improving the calculation efficiency of the classification indication feature.

[0152] Specifically, pooling the first feature to be processed and the second feature to be processed to obtain a classification indication feature includes: performing average pooling on the first feature to be processed and the second feature to be processed to obtain a classification indication feature; or, performing max pooling on the first feature to be processed and the second feature to be processed to obtain a classification indication feature.

[0153] In addition, optionally, after pooling the first feature to be processed and the second feature to be processed to obtain a classification indication feature, the above method may further include: calculating a loss function between the classification indication feature and the sample feature, and adjusting the parameters of the two-stream structure block according to the loss function.

[0154] As an alternative embodiment, after determining a classification indication feature corresponding to the image to be classified based on the reference global feature and the reference local feature, the above method further includes: determining the category corresponding to the image to be classified through the classification indication feature. In this way, the image category can be determined based on the classification indication feature used to accurately represent the image to be classified, improving the recognition accuracy of the image category.

[0155] For example, if the image to be classified is a product picture, the category may include multiple types (such as daily necessities, pet supplies, kitchen supplies, etc.). The category corresponding to the image to be classified may belong to one of the above multiple types.

[0156] Among them, determining the category corresponding to the image to be classified through the classification indication feature includes: generating a classification sequence through the classification indication feature. The classification sequence includes multiple probability values, each probability value corresponding to a category, and each probability value being used to represent the probability that the image to be classified belongs to that category; determining the category corresponding to the maximum probability value in the classification sequence as the category corresponding to the image to be classified.

[0157] Please refer to Figure 3 , Figure 3 which schematically shows a flowchart of a feature processing method according to another embodiment of the present application. As Figure 3 shown, the feature processing method may include: step S310 to step S334.

[0158] Step S310: Obtain the sample local feature and the sample global feature of the image to be classified.

[0159] Step S312: Extract a first intermediate global feature of the sample global feature based on the first global normalization network, the first global multi-head network, and the second global normalization network.

[0160] Step S314: Obtain the locally normalized features corresponding to the sample local features, and input the first intermediate feature and the locally normalized features into the second global multi-head network, so that the second global multi-head network generates the second intermediate global feature.

[0161] Step S316: Generate the features corresponding to the second intermediate global feature based on the third global normalization network and the global feed-forward network as the reference global features of the sample global features.

[0162] Step S318: Input the locally normalized features and the reference global features into the local multi-head network, so that the local multi-head network generates the first intermediate local feature.

[0163] Step S320: Input the first intermediate local feature and the locally normalized features into the local normalization network, so that the local normalization network generates the second intermediate local feature.

[0164] Step S322: Trigger the local feed-forward network to generate the reference local features based on the second intermediate local feature and the first intermediate local feature.

[0165] Step S324: Fuse the reference global features and the reference local features to obtain the first fusion result, and perform layer normalization on the first fusion result to obtain the second fusion result, and then generate the self-attention fusion feature corresponding to the second fusion result.

[0166] Step S326: Generate the feature to be split based on the self-attention fusion feature and the first fusion result, and split the feature to be split into the target global feature and the target local feature.

[0167] Step S328: Perform normalization processing on the target global feature through the semantic normalization network to obtain the semantic normalization result, and generate the semantic comprehensive feature corresponding to the semantic normalization result through the semantic feed-forward network, and then fuse the semantic comprehensive feature and the target global feature to obtain the first feature to be processed corresponding to the target global feature.

[0168] Step S330: Perform layer normalization on the target local feature through the pixel normalization network to obtain the pixel normalization result, and generate the pixel comprehensive feature corresponding to the pixel normalization result through the pixel feed-forward network, and then fuse the pixel comprehensive feature and the target local feature to obtain the second feature to be processed corresponding to the target local feature.

[0169] Step S332: Perform pooling processing on the first feature to be processed and the second feature to be processed to obtain the classification indication feature.

[0170] Step S334: Determine the category of the image to be classified through the classification indication feature.

[0171] It should be noted that steps S310 to S334 correspond to the respective steps and their embodiments shown in Figure 2 . For the specific implementation manners of steps S310 to S334, please refer to Figure 2 for the respective steps and their embodiments shown therein, which will not be elaborated herein.

[0172] It can be seen that by implementing the method shown in Figure 3 , the feature extraction process can be divided into global feature extraction and local feature extraction, and local features are calculated in combination with global features, so as to determine the classification indication features corresponding to the image to be classified according to accurate global features and local features. This can avoid losing features during the process of extracting image features and can improve the accuracy of image feature extraction. In addition, based on global feature extraction and local feature extraction, it is helpful to reduce the computational amount of the attention mechanism on the extraction paths corresponding to global feature extraction and local feature extraction respectively, thereby reducing the complexity of attention calculation. In addition, global feature extraction and local feature extraction can be executed in parallel to improve the overall efficiency of feature extraction.

[0173] Please refer to Figure 4 , Figure 4 which shows a schematic diagram of the feature extraction network structure of the existing solution. As shown in Figure 4 , the feature extraction network of the existing solution may include: a normalization network Layer Norm410, a multi-head attention network Multi-HeadAttention420, a normalization network Layer Norm430, and a feed-forward neural network Feed Forward440.

[0174] Specifically, as shown in Figure 4 , the traditional feature extraction network needs to obtain the features of the image to be classified and input the features x of the image to be classified l into the normalization network Layer Norm410, so that the normalization network Layer Norm410 calculates the normalized feature based on the expression . Furthermore, the normalized feature is input into the multi-head attention network Multi-HeadAttention420, and the multi-head attention network Multi-Head Attention420 can calculate the corresponding multi-head self-attention feature . Among them, for each , its self-attention feature can be calculated according to the following expression, MHA(q,k,v) = Concat(head1, ……, head h )W o , h, W, and d respectively represent the height, width, and number of channels of the image features to be classified; v is a vector representing the input image features to be classified, and q and k are feature vectors for calculating the Attention weights; is the weight matrix, and d h is the dimension of each head; then, based on the normalization network Layer Norm430, x′ is generated l The corresponding normalized feature LN(x′ l ), and based on the feed-forward neural network FeedForward440, the feed-forward feature FFN(LN(x′ l )) corresponding to LN(x′ l ) is calculated; furthermore, FFN(LN(x′ l )) and the image features x to be classified l can be input into the expression x l+1 = FFN(LN(x′ l )) + x′ l to determine the class feature x l+1 .

[0175] Among them, the multi-head attention network Multi-Head Attention420 has a computational complexity of O(n l d) for x 2 , which is quadratic to the number of image patch features. It can be seen that this design will generate a large amount of computation when processing high-resolution inputs. To reduce the computational load of the multi-head attention network Multi-Head Attention420, the prior art has expanded the downsampling range in the image patch feature pooling operation to generate fewer tokens. However, this is likely to lose the global features of the image. To solve this problem, this application proposes the architecture as Figure 5 shown.

[0176] Please refer to Figure 5 , Figure 5 which schematically shows a network architecture diagram for implementing the feature processing method of this application. As Figure 5 shown, the network architecture (Dual-ViT) for implementing the feature processing method of this application includes: a semantic-level global feature extraction path and a pixel-level local feature extraction path.

[0177] Among them, the semantic-level global feature extraction path includes: the first global normalization network Layer Norm510, the first global multi-head network Multi-Head Attention511, the second global normalization network Layer Norm512, the second global multi-head network Multi-Head Attention513, the third global normalization network Layer Norm514, and the global feed-forward network Feed Forward515.

[0178] The pixel-level local feature extraction path includes: the local normalization network Layer Norm520, the local multi-head network Multi-Head Attention521, the local normalization network Layer Norm522, and the local feed-forward network FeedForward523.

[0179] Specifically, the sample local feature x of the image to be classified can be obtained l and the sample global feature z l , and the sample local feature x l is input into the pixel-level local feature extraction path, and the sample global feature z l is input into the semantic-level global feature extraction path.

[0180] Furthermore, the local normalization network Layer Norm520 in the pixel-level local feature extraction path is used to calculate the local normalization feature corresponding to the sample local feature x based on the expression l

[0181] Furthermore, the first global normalization network Layer Norm510 in the semantic-level global feature extraction path is used to calculate the feature corresponding to the sample global feature z based on the expression l Each head self-attention network in the first global multi-head network Multi-HeadAttention511 is used to generate its respective corresponding self-attention feature and fuse them to obtain the self-attention feature The second global normalization network Layer Norm512 is used to calculate the first intermediate global feature z′ based on the expression l . Furthermore, the first intermediate global feature z′ l can be normalized to obtain the feature LN(z′ l ), and the feature LN(z′ l ) and the local normalization feature Input the self-attention networks of each head in the second global multi-head network Multi-Head Attention513, and fuse the features generated by each head attention network to obtain the second intermediate global feature The third global normalization network Layer Norm514 is used to generate based on the second intermediate global feature and the input first intermediate global feature z′ l Generate The global feed-forward network Feed Forward515 is used to process Into Furthermore, And the feature Substitute into the expression To calculate the reference global feature z of the sample global feature l+1 .

[0182] In addition, the pixel-level local feature extraction path can be based on the expression Normalize the reference global feature z l+1 To obtain the feature And the feature And the local normalization feature Input the self-attention networks (Self-Attention) of each head in the local multi-head network Multi-Head Attention521, and fuse the features generated by each head attention network to obtain the first intermediate local feature The local normalization network Layer Norm522 is used to generate the second intermediate local feature x′ based on the first intermediate local feature And the local normalization feature x l Generate the second intermediate local feature x′ l , Normalize the second intermediate local feature x′ l To get LN(x′ l ), input LN(x′ l ) into the local feed-forward network Feed Forward523, and the local feed-forward network Feed Forward523 is used to process the features of LN(x′ l ) to obtain FFN(LN(x′ l )); furthermore, FFN(LN(x′ l )) and the second intermediate local feature x′ l Can be substituted into the expression x l+1 =FFN(LN(x′ l ))+x′ l , to calculate the reference local feature x l+1 .

[0183] Furthermore, the reference global feature z l+1 and the reference local feature x l+1 can be used as the input of the next network architecture (Dual-ViT). In the actual application process of this application, multiple network architectures (Dual-ViT) can be spliced to achieve multiple extractions of features, so as to improve the expression ability of features. Optionally, the network architecture (Dual-ViT) can also be spliced with Figure 6 the feature merging and extraction path shown to achieve the fusion and new-stage splitting of the reference global feature z l+1 and the reference local feature x l+1 , which can also improve the expression ability of features.

[0184] Please refer to Figure 6 , Figure 6 which schematically shows another network architecture diagram for implementing the feature processing method of this application. As Figure 6 shown, the network architecture for implementing the feature processing method of this application includes: a semantic-level global feature extraction path, a pixel-level local feature extraction path, a feature merging and extraction path, and a pooling layer 680.

[0185] Among them, the semantic-level global feature extraction path includes: the first global normalization network Layer Norm610, the first global multi-head network Multi-Head Attention611, the second global normalization network Layer Norm612, the second global multi-head network Multi-Head Attention613, the third global normalization network Layer Norm614, and the global feed-forward network Feed Forward615.

[0186] The pixel-level local feature extraction path includes: the local normalization network Layer Norm620, the local multi-head network Multi-Head Attention621, the local normalization network Layer Norm622, and the local feed-forward network FeedForward623.

[0187] The feature merging and extraction path includes: the connection layer Concat630, the normalization network Layer Norm640, the multi-head attention network Multi-Head Attention650, the feature splitting layer Split660, the semantic normalization network Layer Norm671, the semantic feed-forward network Feed Forward672, the pixel normalization network Layer Norm681, and the pixel feed-forward network FeedForward682.

[0188] Among them, each network in the semantic-level global feature extraction path and the pixel-level local feature extraction path is consistent with the networks shown in Figure 5 For the specific application methods of each network in the semantic-level global feature extraction path and the pixel-level local feature extraction path, please refer to Figure 5 for the description, which will not be elaborated here.

[0189] The connection layer Concat630 is used to perform tensor connection on the reference global feature z l+1 and the reference local feature x l+1 after receiving them, to obtain the first fusion result x l+1 ||z l+1 l+1 ||z l+1 .

[0190] The normalization network Layer Norm640 is used to perform layer normalization on the first fusion result x ||z l+1 ||z l+1 based on the expression to obtain the second fusion result

[0191] The multi-head attention network Multi-Head Attention650 is used to input the second fusion result into each head self-attention network (Self-Attention) in the multi-head network Multi-Head Attention, and fuse the features generated by each head attention network to obtain the self-attention fusion feature

[0192] The feature splitting layer Split660 is used to determine the features to be split (x′ , z′ l+1 , z′ l+1 ) based on the expression. Among them, the features to be split (x′ l , z′ l ) include the target global feature z′ l+1 and the target local feature x′ l+1 .

[0193] The semantic normalization network Layer Norm671 is used to perform normalization on the target global feature z′ l+1 to obtain the semantic normalization result LN(z′ l+1 ).

[0194] The semantic feed-forward network Feed Forward672 is used to generate the semantic comprehensive feature corresponding to the semantic normalization result LN(z′ l+1 ) ​

[0195] Furthermore, the feature merging and extraction path can also combine the semantic comprehensive feature and the target global feature z' l+1 and substitute them into the expression z l+2 = FFN Z (LN(z' l+1 )) + z' l+1 to calculate the first feature to be processed z corresponding to the target global feature l+2 .

[0196] The pixel normalization network Layer Norm681 is used to normalize the target local feature x' l+1 to obtain the pixel normalization result LN(x' l+1 ).

[0197] The pixel feed-forward network Feed Forward682 is used to generate the pixel comprehensive feature FFN corresponding to the pixel normalization result LN(x' l+1 ) Z (LN(x' l+1 ).

[0198] Furthermore, the feature merging and extraction path can also combine the pixel comprehensive feature FFN X (LN(x' l+1 )) and the target local feature x' l+1 and substitute them into the expression x l+2 = FFN X (LN(x' l+1 )) + x' l+1 to calculate the second feature to be processed x corresponding to the target local feature l+2 .

[0199] Furthermore, the pooling layer 680 can perform pooling on the first feature to be processed z l+2 and the second feature to be processed x l+2 to obtain the classification indication feature

[0200] It can be seen that in this way, the feature extraction process can be divided into global feature extraction and local feature extraction, combining global features to calculate local features, so as to determine the classification indication feature corresponding to the image to be classified according to accurate global features and local features. This can avoid losing features in the process of extracting image features and improve the accuracy of image feature extraction. In addition, based on global feature extraction and local feature extraction, it helps to reduce the computational complexity of the attention mechanism on the extraction paths corresponding to global feature extraction and local feature extraction respectively, thereby reducing the complexity of attention calculation. In addition, global feature extraction and local feature extraction can be executed in parallel to improve the efficiency of overall feature extraction

[0201] Please refer to Figure 7 , Figure 7 which schematically shows another network architecture diagram for implementing the feature processing method of the present application. As Figure 7 shown, the network architecture for implementing the feature processing method of the present application includes: a pixel-level local feature extraction path 710, ……, a pixel-level local feature extraction path 711, a semantic-level global feature extraction path 720, ……, a semantic-level global feature extraction path 721, a feature merging and extraction path 730, ……, a feature merging and extraction path 731. Among them, each pixel-level local feature extraction path, semantic-level global feature extraction path, and feature merging and extraction path is used to implement the steps as Figure 5 and 6 shown, which will not be elaborated here.

[0202] Specifically, in the present application, it may include multiple stages. Taking one stage as an example, it may include one or more pixel-level local feature extraction paths (which can also be understood as network blocks), and one or more semantic-level global feature extraction paths (which can also be understood as network blocks); or, it includes one or more feature merging and extraction paths (which can also be understood as network blocks). The number of paths in each stage can be customized according to the size of the image to be recognized. It should be noted that the number of pixel-level local feature extraction paths, semantic-level global feature extraction paths, and feature merging and extraction paths is not limited.

[0203] In a network architecture including multiple stages, the parameters corresponding to different paths can be different.

[0204] For example, the first stage includes a pixel-level local feature extraction path and a semantic-level global feature extraction path. The parameters corresponding to the first stage may include the dimensionality expansion rate of the pixel-level local feature extraction path the dimensionality expansion rate of the semantic-level global feature extraction path the number of attention heads HD1 = 2 in the multi-head attention mechanism, the number of feature channels C1 = 64, the feature resolution of the first stage the number of network blocks in the first stage (e.g., 3).

[0205] The second stage includes a pixel-level local feature extraction path and a semantic-level global feature extraction path. The parameters corresponding to the second stage may include the dimensionality expansion rate of the pixel-level local feature extraction path the dimensionality expansion rate of the semantic-level global feature extraction path the number of attention heads HD2 = 4 in the multi-head attention mechanism, the number of feature channels C2 = 128, the feature resolution of the second stage the number of network blocks in the second stage (e.g., 4).

[0206] The third stage includes a feature merging and extraction path. The parameters corresponding to the third stage may include the dimensional expansion rate of the pixel-level local feature extraction path The dimensional expansion rate of the semantic-level global feature extraction path The number of attention heads HD2 = 10 in the multi-head attention mechanism, the number of feature channels C3 = 320, and the feature resolution of the third stage The number of network blocks in the third stage (e.g., 6).

[0207] The fourth stage includes a feature merging and extraction path. The parameters corresponding to the fourth stage may include the dimensional expansion rate of the pixel-level local feature extraction path The dimensional expansion rate of the semantic-level global feature extraction path The number of attention heads HD4 = 14 in the multi-head attention mechanism, the number of feature channels C4 = 448, and the feature resolution of the fourth stage The number of network blocks in the fourth stage (e.g., 3).

[0208] Please refer to Figure 8 , Figure 8 which schematically shows a structural block diagram of a feature processing device according to an embodiment of the present application. The feature processing device 800 corresponds to Figure 2 the method shown, as Figure 8 shown, the feature processing device 800 includes:

[0209] A feature acquisition unit 801, configured to acquire a sample local feature and a sample global feature of an image to be classified;

[0210] A feature generation unit 802, configured to generate a reference global feature corresponding to the sample global feature;

[0211] The feature generation unit 802 is further configured to generate a reference local feature according to the reference global feature and the sample local feature;

[0212] A feature determination unit 803, configured to determine a classification indication feature corresponding to the image to be classified based on the reference global feature and the reference local feature.

[0213] It can be seen that implementing Figure 8The device shown can divide the feature extraction process into global feature extraction and local feature extraction, combine global features to calculate local features, and thus determine classification indication features corresponding to the image to be classified based on accurate global and local features. This can avoid losing features during the process of extracting image features and improve the accuracy of image feature extraction. In addition, based on global feature extraction and local feature extraction, it can help reduce the computational complexity of the attention mechanism on the extraction paths corresponding to global feature extraction and local feature extraction respectively, thereby reducing the complexity of attention calculation. In addition, global feature extraction and local feature extraction can be executed in parallel to improve the efficiency of overall feature extraction.

[0214] In an exemplary embodiment of the present application, the feature generation unit 802 generates a reference global feature corresponding to the sample global feature, including:

[0215] Extracting a first intermediate global feature of the sample global feature based on the first global normalization network, the first global multi-head network, and the second global normalization network;

[0216] Obtaining the local normalization feature corresponding to the sample local feature, and inputting the first intermediate feature and the local normalization feature into the second global multi-head network, so that the second global multi-head network generates a second intermediate global feature;

[0217] Generating a feature corresponding to the second intermediate global feature based on the third global normalization network and the global feed-forward network as the reference global feature of the sample global feature;

[0218] Wherein, the first global normalization network, the second global normalization network, and the third global normalization network correspond to different network parameters; the first global multi-head network and the second global multi-head network correspond to different network parameters.

[0219] It can be seen that implementing this optional embodiment can achieve extracting the reference global feature of the sample global feature based on multiple global normalization networks, global multi-head networks, and global feed-forward networks to implement the multi-head attention calculation of the sample global feature, thereby improving the accuracy of the reference global feature.

[0220] In an exemplary embodiment of the present application, the feature generation unit 802 generates a reference local feature according to the reference global feature and the sample local feature, including:

[0221] Inputting the local normalization feature and the reference global feature into the local multi-head network, so that the local multi-head network generates a first intermediate local feature;

[0222] Inputting the first intermediate local feature and the local normalization feature into the local normalization network, so that the local normalization network generates a second intermediate local feature;

[0223] The trigger local feedforward network generates a reference local feature based on the second intermediate local feature and the first intermediate local feature.

[0224] It can be seen that by implementing this optional embodiment, a reference local feature corresponding to the sample local feature can be calculated based on the reference global feature. Through the feature extraction by combining the reference global feature and the sample local feature, the loss of the global feature during the feature extraction process can be avoided. Combining the reference global feature and the sample local feature can also calculate a reference local feature with higher accuracy, and at the same time, the difficulty of extracting fine local features can be reduced.

[0225] In an exemplary embodiment of the present application, the feature determination unit 803 determines a classification indication feature corresponding to the image to be classified based on the reference global feature and the reference local feature, including:

[0226] Fuse the reference global feature and the reference local feature to obtain a feature to be split;

[0227] Split the feature to be split into a target global feature and a target local feature;

[0228] Determine a classification indication feature corresponding to the image to be classified according to the target global feature and the target local feature.

[0229] It can be seen that by implementing this optional embodiment, the feature fusion of the reference global feature and the reference local feature can be realized to determine a feature to be split without missing feature information. Furthermore, the feature to be split can be split, and then global and local feature processing can be performed based on the splitting result, so as to obtain a classification indication feature, ensuring that the classification indication feature contains local features and global features that have not been lost, and improving the accuracy of the classification indication feature.

[0230] In an exemplary embodiment of the present application, the feature determination unit 803 fuses the reference global feature and the reference local feature to obtain a feature to be split, including:

[0231] Fuse the reference global feature and the reference local feature to obtain a first fusion result;

[0232] Perform layer normalization processing on the first fusion result to obtain a second fusion result;

[0233] Generate a self-attention fusion feature corresponding to the second fusion result;

[0234] Generate a feature to be split based on the self-attention fusion feature and the first fusion result.

[0235] It can be seen that by implementing this optional embodiment, the reference global feature and the reference local feature can be fused based on layer normalization processing and self-attention calculation, so as to ensure that the feature to be split obtained by fusion can restore the feature information of the entire image to be classified to the greatest extent, thereby ensuring the accuracy of subsequent feature extraction calculations and reducing distortion.

[0236] In an exemplary embodiment of the present application, the feature determination unit 803 determines a classification indication feature corresponding to the image to be classified according to the target global feature and the target local feature, including:

[0237] Generating a first feature to be processed corresponding to the target global feature based on the global feature processing network;

[0238] Generating a second feature to be processed corresponding to the target local feature based on the local feature processing network;

[0239] Generating a classification indication feature corresponding to the image to be classified according to the first feature to be processed and the second feature to be processed.

[0240] It can be seen that by implementing this optional embodiment, the first feature to be processed and the second feature to be processed can be obtained based on the fused local feature processing path and global feature processing path, and the classification indication feature for accurately representing the image to be classified can be determined based on the first feature to be processed and the second feature to be processed.

[0241] In an exemplary embodiment of the present application, the global feature processing network includes a semantic normalization network and a semantic feedforward network. The feature determination unit 803 generates a first feature to be processed corresponding to the target global feature based on the global feature processing network, including:

[0242] Performing normalization processing on the target global feature through the semantic normalization network to obtain a semantic normalization result;

[0243] Generating a semantic comprehensive feature corresponding to the semantic normalization result through the semantic feedforward network;

[0244] Fusing the semantic comprehensive feature and the target global feature to obtain a first feature to be processed corresponding to the target global feature.

[0245] It can be seen that by implementing this optional embodiment, the first feature to be processed can be obtained based on the semantic normalization network and the semantic feedforward network, and the probability of feature loss in the global feature processing process can be reduced.

[0246] In an exemplary embodiment of the present application, the local feature processing network includes a pixel normalization network and a pixel feedforward network. The feature determination unit 803 generates a second feature to be processed corresponding to the target local feature based on the local feature processing network, including:

[0247] The target local features are subjected to layer normalization processing through a pixel normalization network to obtain a pixel normalization result;

[0248] A pixel comprehensive feature corresponding to the pixel normalization result is generated through a pixel feed-forward network;

[0249] The pixel comprehensive feature and the target local feature are fused to obtain a second feature to be processed corresponding to the target local feature.

[0250] It can be seen that by implementing this optional embodiment, the first feature to be processed can be obtained based on the pixel normalization network and the pixel feed-forward network, which can reduce the probability of feature loss during the local feature processing.

[0251] In an exemplary embodiment of the present application, the feature determination unit 803 generates a classification indication feature corresponding to the image to be classified according to the first feature to be processed and the second feature to be processed, including:

[0252] The first feature to be processed and the second feature to be processed are subjected to pooling processing to obtain a classification indication feature.

[0253] It can be seen that by implementing this optional embodiment, the classification indication feature can be obtained through pooling processing, which improves the calculation efficiency of the classification indication feature.

[0254] In an exemplary embodiment of the present application, the above device includes:

[0255] A category determination unit, configured to determine the category corresponding to the image to be classified through the classification indication feature after the feature determination unit 803 determines the classification indication feature corresponding to the image to be classified based on the reference global feature and the reference local feature.

[0256] It can be seen that by implementing this optional embodiment, the image category can be determined based on the classification indication feature for accurately representing the image to be classified, which improves the recognition accuracy of the image category.

[0257] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0258] Since each functional module of the task scheduling device in the exemplary embodiment of the present application corresponds to the steps of the exemplary embodiment of the above task scheduling method, for the details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the above task scheduling method of the present application.

[0259] Please refer to Figure 9 , Figure 9 which shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application.

[0260] It should be noted that Figure 9 the computer system 900 of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0261] As Figure 9 shown, the computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage section 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for system operation are also stored. The CPU 901, ROM 902, and RAM 903 are connected to each other via a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.

[0262] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that a computer program read from it can be installed into the storage section 908 as needed.

[0263] Specifically, according to the embodiments of the present application, the process described with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by the central processing unit (CPU) 901, various functions defined in the method and device of the present application are executed.

[0264] As another aspect, the present application also provides a computer-readable medium. This computer-readable medium can be included in the electronic device described in the above embodiments; or it can exist separately without being assembled into the electronic device. When the one or more programs carried by the above computer-readable medium are executed by an electronic device, the electronic device implements the method described in the above embodiments.

[0265] It should be noted that the computer-readable medium shown in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0266] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0267] The units involved in the embodiments described in the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.

[0268] Those skilled in the art will readily conceive of other implementations of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include known common knowledge or conventional technical means in the art not disclosed herein. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present application are pointed out by the foregoing claims.

Claims

1. A feature processing method, characterized in that, Including: Obtaining sample local features and sample global features of the image to be classified; Extracting first intermediate global features of the sample global features based on a first global normalization network, a first global multi-head network, and a second global normalization network; Obtaining local normalization features corresponding to the sample local features, and inputting the first intermediate global features and the local normalization features into a second global multi-head network, so that the second global multi-head network generates second intermediate global features; Generating features corresponding to the second intermediate global features based on a third global normalization network and a global feed-forward network as reference global features of the sample global features; wherein, the first global normalization network, the second global normalization network, and the third global normalization network correspond to different network parameters; the first global multi-head network and the second global multi-head network correspond to different network parameters; Generating reference local features according to the reference global features and the sample local features; Determining classification indication features corresponding to the image to be classified based on the reference global features and the reference local features.

2. The method according to claim 1, characterized in that, Generating reference local features according to the reference global features and the sample local features, including: Inputting the local normalization features and the reference global features into a local multi-head network, so that the local multi-head network generates first intermediate local features; Inputting the first intermediate local features and the local normalization features into a local normalization network, so that the local normalization network generates second intermediate local features; Triggering a local feed-forward network to generate reference local features based on the second intermediate local features and the first intermediate local features.

3. The method according to claim 1, characterized in that, Determining classification indication features corresponding to the image to be classified based on the reference global features and the reference local features, including: Fusing the reference global features and the reference local features to obtain features to be split; Splitting the features to be split into target global features and target local features; Determining classification indication features corresponding to the image to be classified according to the target global features and the target local features.

4. The method according to claim 3, characterized in that, Fusing the reference global features and the reference local features to obtain features to be split, including: Fusing the reference global features and the reference local features to obtain a first fusion result; Performing layer normalization processing on the first fusion result to obtain a second fusion result; Generating self-attention fusion features corresponding to the second fusion result; Generating the features to be split based on the self-attention fusion features and the first fusion result.

5. The method according to claim 3, characterized in that, Determining classification indication features corresponding to the image to be classified according to the target global features and the target local features, including: Generating first features to be processed corresponding to the target global features based on a global feature processing network; Generating second features to be processed corresponding to the target local features based on a local feature processing network; Generating classification indication features corresponding to the image to be classified according to the first features to be processed and the second features to be processed.

6. The method according to claim 5, characterized in that, The global feature processing network includes a semantic normalization network and a semantic feedforward network. Generating a first feature to be processed corresponding to the target global feature based on the global feature processing network includes: Performing normalization processing on the target global feature through the semantic normalization network to obtain a semantic normalization result; Generating a semantic comprehensive feature corresponding to the semantic normalization result through the semantic feedforward network; Fusing the semantic comprehensive feature and the target global feature to obtain a first feature to be processed corresponding to the target global feature.

7. The method according to claim 5, characterized in that, The local feature processing network includes a pixel normalization network and a pixel feedforward network. Generating a second feature to be processed corresponding to the target local feature based on the local feature processing network includes: Performing layer normalization processing on the target local feature through the pixel normalization network to obtain a pixel normalization result; Generating a pixel comprehensive feature corresponding to the pixel normalization result through the pixel feedforward network; Fusing the pixel comprehensive feature and the target local feature to obtain a second feature to be processed corresponding to the target local feature.

8. The method according to claim 5, characterized in that, Generating a classification indication feature corresponding to the image to be classified according to the first feature to be processed and the second feature to be processed includes: Performing pooling processing on the first feature to be processed and the second feature to be processed to obtain a classification indication feature.

9. The method according to claim 1, wherein, After determining the classification indication feature corresponding to the image to be classified based on the reference global feature and the reference local feature, the method further includes: Determining the category corresponding to the image to be classified through the classification indication feature.

10. A feature processing device, wherein, Including: A feature acquisition unit for acquiring a sample local feature and a sample global feature of the image to be classified; A feature generation unit for extracting a first intermediate global feature of the sample global feature based on a first global normalization network, a first global multi-head network, and a second global normalization network; Obtaining a local normalization feature corresponding to the sample local feature, and inputting the first intermediate global feature and the local normalization feature into a second global multi-head network so that the second global multi-head network generates a second intermediate global feature; Generating a feature corresponding to the second intermediate global feature based on a third global normalization network and a global feedforward network as the reference global feature of the sample global feature; wherein, the first global normalization network, the second global normalization network, and the third global normalization network correspond to different network parameters; the first global multi-head network and the second global multi-head network correspond to different network parameters; The feature generation unit is further configured to generate a reference local feature according to the reference global feature and the sample local feature; A feature determination unit for determining a classification indication feature corresponding to the image to be classified based on the reference global feature and the reference local feature.

11. A computer program product, comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-9.

12. A computer-readable storage medium, on which a computer program is stored, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-9.

13. An electronic device, wherein, Including: A processor; And A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the method according to any one of claims 1-9 by executing the executable instructions.

Citation Information

Patent Citations

  • Image processing method and device, equipment, storage medium and computer program product

    CN113642585A

  • Target re-identification method, terminal equipment and computer readable storage medium

    CN114419408A