Image feature processing method, device, product, medium and equipment
Feature downsampling and reconstruction are performed through discrete wavelet transformation and inverse transformation, combined with the multi-head self-attention feature, the information loss problem caused by feature pooling is solved, and the accuracy of image feature extraction is improved.
Patent Information
- Application Number
- CN202210635612.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-06
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-06-06
AI Technical Summary
The prior art reduces the feature dimension through feature pooling in image recognition, resulting in loss of feature information, thereby reducing the final feature extraction accuracy.
Discrete wavelet transformation is used for reversible feature downsampling, and discrete wavelet inverse transformation is used to expand the receptive field, reconstruct the feature map, and combine the reconstruction feature map and the multi-head self-attention feature to generate classification indicator features with high accuracy.
Through discrete wavelet transformation and inverse transformation, the loss of feature information is avoided, the expression ability of classification indicator features is enhanced, and the feature extraction accuracy is improved.
Smart Images

Figure CN114972897B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image feature processing method, an image feature processing device, a computer program product, a computer-readable storage medium, and an electronic device. Background Art
[0002] In the field of artificial intelligence recognition, such as image recognition and text recognition, it is usually necessary to use a network architecture based on the Transformer block for feature extraction, so as to recognize the image / text based on the extracted features. The Transformer block usually needs to use feature pooling to reduce the feature dimension, so as to achieve the purpose of reducing the amount of calculation. However, this method usually loses a lot of feature information, which leads to low accuracy of the final feature extraction result.
[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute an existing solution known to ordinary technicians in the field. Summary of the invention
[0004] The purpose of the present application is to provide an image feature processing method, an image feature processing device, a computer program product, a computer-readable storage medium and an electronic device, which can perform reversible feature downsampling through discrete wavelet transform, and use discrete wavelet inverse transform to expand the receptive field to obtain a reconstructed feature map of the downsampled feature map, so that a higher-precision classification indication feature can be obtained by combining the reconstructed feature map and multi-head self-attention, thereby improving the feature extraction accuracy compared to the prior art.
[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.
[0006] According to one aspect of the present application, a method for processing image features is provided, the method comprising:
[0007] Generate a downsampled feature map corresponding to the image to be identified based on discrete wavelet transform;
[0008] Generate multi-head self-attention features corresponding to the image to be recognized according to the downsampled feature map;
[0009] The downsampled feature map is transformed into a reconstructed feature map based on inverse discrete wavelet transform;
[0010] A classification indication feature corresponding to the image to be identified is generated based on the reconstructed feature map and the multi-head self-attention features.
[0011] In an exemplary embodiment of the present application, generating a downsampled feature map corresponding to an image to be identified based on discrete wavelet transform includes:
[0012] Obtaining the feature space corresponding to the image to be identified;
[0013] A downsampled feature map corresponding to the feature space is generated based on discrete wavelet transform.
[0014] In an exemplary embodiment of the present application, obtaining a feature space corresponding to an image to be identified includes:
[0015] Generate a feature map to be identified corresponding to the image to be identified;
[0016] A first linear transformation is performed on the feature map to be identified using preset parameters to obtain a feature space.
[0017] In an exemplary embodiment of the present application, generating a feature map to be identified corresponding to an image to be identified includes:
[0018] Perform multi-layer convolution processing on the image to be identified to obtain an image block sequence of the image to be identified;
[0019] Generate a feature map to be identified based on the image block sequence.
[0020] In an exemplary embodiment of the present application, generating a downsampled feature map corresponding to a feature space based on discrete wavelet transform includes:
[0021] Decompose the feature space into a set of wavelet sub-band features through discrete wavelet transform;
[0022] The wavelet sub-band feature set is fused into the target feature;
[0023] The target features are convolved to obtain a downsampled feature map.
[0024] In an exemplary embodiment of the present application, the wavelet sub-band feature set includes low-frequency band row features, high-frequency band row features, low-frequency band column features, and high-frequency band column features. The feature space is decomposed into the wavelet sub-band feature set by discrete wavelet transform, including:
[0025] Based on the low-pass filter in discrete wavelet transform, the low-frequency band row features corresponding to the feature space are obtained;
[0026] Based on the high-pass filter in discrete wavelet transform, high-frequency band row features corresponding to the feature space are obtained;
[0027] Obtain low-frequency band column features corresponding to the feature space based on the low-pass filter in discrete wavelet transform;
[0028] The high-frequency band column features corresponding to the feature space are obtained based on the high-pass filter in discrete wavelet transform.
[0029] In an exemplary embodiment of the present application, generating a multi-head self-attention feature corresponding to an image to be recognized according to a downsampled feature map includes:
[0030] Perform a second linear transformation on the feature map to be identified to obtain the query feature;
[0031] Perform the third linear transformation on the downsampled feature map to obtain key features and value features;
[0032] Generate multi-head self-attention features based on query features, key features, and value features.
[0033] In an exemplary embodiment of the present application, a multi-head self-attention feature is generated according to a query feature, a key feature, and a value feature, including:
[0034] According to the preset division rules, the query features, key features and value features are respectively divided into query sub-feature sets, key sub-feature sets and value sub-feature sets;
[0035] Each feature in the query sub-feature set is input into the corresponding self-attention network, each feature in the key sub-feature set is input into the corresponding self-attention network, and each feature in the value sub-feature set is input into the corresponding self-attention network, so that each head self-attention network generates a corresponding self-attention feature and obtains a multi-head self-attention feature.
[0036] In an exemplary embodiment of the present application, generating a classification indication feature corresponding to the image to be identified according to the reconstructed feature map and the multi-head self-attention feature includes:
[0037] Fusion of multi-head self-attention features and reconstructed feature maps to obtain comprehensive target features;
[0038] Generate classification indicator features based on the target comprehensive features.
[0039] In an exemplary embodiment of the present application, generating a classification indication feature according to a target comprehensive feature includes:
[0040] Perform discrete wavelet transform on the comprehensive features of the target to obtain a new downsampled feature map;
[0041] Generate a new multi-head self-attention feature corresponding to the image to be recognized according to the new downsampled feature map;
[0042] The new downsampled feature map is transformed into a new reconstructed feature map by inverse discrete wavelet transform;
[0043] A classification indication feature corresponding to the image to be identified is generated according to the new reconstructed feature map and the new multi-head self-attention feature.
[0044] In an exemplary embodiment of the present application, the method further includes:
[0045] If the current feature processing stage is not the final feature processing stage, the classification indication feature is input into the next feature processing stage.
[0046] In an exemplary embodiment of the present application, after generating the classification indication feature of the image to be identified according to the target comprehensive feature, the method further includes:
[0047] Determine the category of the image to be identified based on the classification indication features.
[0048] In an exemplary embodiment of the present application, determining the category of the image to be identified according to the classification indication feature includes:
[0049] The classification indicator feature is extracted through a feedforward neural network to obtain a new classification indicator feature;
[0050] Determine the category of the image to be identified based on the new classification indication features.
[0051] According to one aspect of the present application, there is provided an image feature processing device, comprising:
[0052] A feature sampling unit, used for generating a down-sampled feature map corresponding to the image to be identified based on discrete wavelet transform;
[0053] A feature generation unit, used for generating a multi-head self-attention feature corresponding to the image to be recognized according to the downsampled feature map;
[0054] A feature transformation unit, used for transforming the downsampled feature map into a reconstructed feature map based on an inverse discrete wavelet transform;
[0055] The feature generation unit is also used to generate classification indication features corresponding to the image to be identified based on the reconstructed feature map and the multi-head self-attention features.
[0056] In an exemplary embodiment of the present application, the feature sampling unit generates a downsampled feature map corresponding to the image to be identified based on discrete wavelet transform, including:
[0057] Obtaining the feature space corresponding to the image to be identified;
[0058] A downsampled feature map corresponding to the feature space is generated based on discrete wavelet transform.
[0059] In an exemplary embodiment of the present application, the feature sampling unit obtains a feature space corresponding to the image to be identified, including:
[0060] Generate a feature map to be identified corresponding to the image to be identified;
[0061] A first linear transformation is performed on the feature map to be identified using preset parameters to obtain a feature space.
[0062] In an exemplary embodiment of the present application, the feature sampling unit generates a feature map to be identified corresponding to the image to be identified, including:
[0063] Perform multi-layer convolution processing on the image to be identified to obtain an image block sequence of the image to be identified;
[0064] Generate a feature map to be identified based on the image block sequence.
[0065] In an exemplary embodiment of the present application, the feature sampling unit generates a downsampled feature map corresponding to the feature space based on discrete wavelet transform, including:
[0066] Decompose the feature space into a set of wavelet sub-band features through discrete wavelet transform;
[0067] The wavelet sub-band feature set is fused into the target feature;
[0068] The target features are convolved to obtain a downsampled feature map.
[0069] In an exemplary embodiment of the present application, the wavelet sub-band feature set includes low-frequency band row features, high-frequency band row features, low-frequency band column features, and high-frequency band column features. The feature sampling unit decomposes the feature space into a wavelet sub-band feature set by discrete wavelet transform, including:
[0070] Based on the low-pass filter in discrete wavelet transform, the low-frequency band row features corresponding to the feature space are obtained;
[0071] Based on the high-pass filter in discrete wavelet transform, high-frequency band row features corresponding to the feature space are obtained;
[0072] Obtain low-frequency band column features corresponding to the feature space based on the low-pass filter in discrete wavelet transform;
[0073] The high-frequency band column features corresponding to the feature space are obtained based on the high-pass filter in discrete wavelet transform.
[0074] In an exemplary embodiment of the present application, the feature generation unit generates a multi-head self-attention feature corresponding to the image to be recognized according to the downsampled feature map, including:
[0075] Perform a second linear transformation on the feature map to be identified to obtain the query feature;
[0076] Perform the third linear transformation on the downsampled feature map to obtain key features and value features;
[0077] Generate multi-head self-attention features based on query features, key features, and value features.
[0078] In an exemplary embodiment of the present application, the feature generation unit generates a multi-head self-attention feature according to the query feature, the key feature and the value feature, including:
[0079] According to the preset division rules, the query features, key features and value features are respectively divided into query sub-feature sets, key sub-feature sets and value sub-feature sets;
[0080] Each feature in the query sub-feature set is input into the corresponding self-attention network, each feature in the key sub-feature set is input into the corresponding self-attention network, and each feature in the value sub-feature set is input into the corresponding self-attention network, so that each head self-attention network generates a corresponding self-attention feature and obtains a multi-head self-attention feature.
[0081] In an exemplary embodiment of the present application, the feature generation unit generates a classification indication feature corresponding to the image to be identified according to the reconstructed feature map and the multi-head self-attention feature, including:
[0082] Fusion of multi-head self-attention features and reconstructed feature maps to obtain comprehensive target features;
[0083] Generate classification indicator features based on the target comprehensive features.
[0084] In an exemplary embodiment of the present application, the feature generation unit generates a classification indication feature according to the target comprehensive feature, including:
[0085] Perform discrete wavelet transform on the comprehensive features of the target to obtain a new downsampled feature map;
[0086] Generate a new multi-head self-attention feature corresponding to the image to be recognized according to the new downsampled feature map;
[0087] The new downsampled feature map is transformed into a new reconstructed feature map by inverse discrete wavelet transform;
[0088] A classification indication feature corresponding to the image to be identified is generated according to the new reconstructed feature map and the new multi-head self-attention feature.
[0089] In an exemplary embodiment of the present application, the above-mentioned device further includes:
[0090] The feature input unit is used to input the classification indication feature into the next feature processing stage if the current feature processing stage is not the final feature processing stage.
[0091] In an exemplary embodiment of the present application, the above-mentioned device further includes:
[0092] The image recognition unit is used to determine the category to which the image to be recognized belongs according to the classification indication feature after the feature generation unit generates the classification indication feature of the image to be recognized according to the target comprehensive feature.
[0093] In an exemplary embodiment of the present application, the image recognition unit determines the category of the image to be recognized according to the classification indication feature, including:
[0094] The classification indicator feature is extracted through a feedforward neural network to obtain a new classification indicator feature;
[0095] Determine the category of the image to be identified based on the new classification indication features.
[0096] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above-mentioned various optional implementations.
[0097] According to one aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, any one of the above methods is implemented.
[0098] According to one aspect of the present application, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any one of the above methods by executing the executable instructions.
[0099] The exemplary embodiments of the present application may have some or all of the following beneficial effects:
[0100] In the image feature processing method provided in an example embodiment of the present application, a downsampled feature map corresponding to the image to be identified can be generated based on discrete wavelet transform; a multi-head self-attention feature corresponding to the image to be identified can be generated based on the downsampled feature map; the downsampled feature map is transformed into a reconstructed feature map based on inverse discrete wavelet transform; and a classification indication feature corresponding to the image to be identified is generated based on the reconstructed feature map and the multi-head self-attention feature. In this way, reversible feature downsampling can be performed through discrete wavelet transform, and the receptive field can be expanded using inverse discrete wavelet transform to obtain a reconstructed feature map of the downsampled feature map, so that a classification indication feature with higher accuracy can be obtained by combining the reconstructed feature map and multi-head self-attention, which improves the feature extraction accuracy compared to the prior art. In addition, feature information loss can be avoided, and the expression ability of the classification indication feature can be enhanced.
[0101] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] The drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0103] Figure 1 A schematic diagram showing an exemplary system architecture of an image feature processing method and an image feature processing device to which the embodiments of the present application can be applied;
[0104] Figure 2 A flowchart of an image feature processing method according to an embodiment of the present application is schematically shown;
[0105] Figure 3 A flowchart of an image feature processing method according to another embodiment of the present application is schematically shown;
[0106] Figure 4 A network architecture diagram for implementing the image feature processing method of the present application is schematically shown;
[0107] Figure 5 Another network architecture diagram for implementing the image feature processing method of the present application is schematically shown;
[0108] Figure 6 The structure block diagram of an image feature processing device according to an embodiment of the present application is schematically shown;
[0109] Figure 7 The structure diagram of a computer system suitable for implementing an electronic device of an embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0110] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as being limited to the examples set forth herein; on the contrary, these embodiments are provided so that the present application will be more comprehensive and complete, and the concept of the example embodiments will be fully conveyed to those skilled in the art. The described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical solutions of the present application may be practiced while omitting one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present application.
[0111] In addition, the accompanying drawings are only schematic illustrations of the present application and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0112] See also Figure 1 , Figure 1 The following is a schematic diagram showing a system architecture of an exemplary application environment in which an image feature processing method and an image feature processing device according to an embodiment of the present application can be applied. Figure 1 As shown, the system architecture 100 may include one or more of terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105.
[0113] The network 104 may include various connection types, such as wired, wireless communication links or fiber optic cables, etc. The terminal devices 101, 102, and 103 may be devices that provide voice and / or data connectivity to users, handheld devices with wireless connection functions, or other processing devices connected to wireless modems. The wireless terminal may communicate with one or more core networks via the RAN. The wireless terminal may be a user equipment (UE), a handheld terminal, a laptop, a subscriber unit, a cellular phone, a smart phone, a wireless data card, a personal digital assistant (PDA) computer, a tablet computer, a wireless modem, a handheld device (handheld), a laptop computer, a cordless phone or a wireless local loop (WLL) station, a machine type communication (MTC) terminal, or other device that can access the network. The terminal and the access network device communicate with each other using a certain air interface technology (for example, 3GPP access technology or non-3GPP access technology). It should be understood that Figure 1The number of terminal devices, networks and servers in the embodiment is only for illustration. According to the implementation requirements, there may be any number of terminal devices, networks and servers. For example, the server 105 may be a server cluster composed of multiple servers.
[0114] The path planning method for multi-node networking provided in the embodiment of the present application can be executed by the server 105, and accordingly, the path planning device for multi-node networking is generally arranged in the server 105. However, it is easy for those skilled in the art to understand that the path planning method for multi-node networking provided in the embodiment of the present application can also be executed by the terminal device 101, 102 or 103, and accordingly, the path planning device for multi-node networking can also be arranged in the terminal device 101, 102 or 103, which is not particularly limited in this exemplary embodiment. For example, in an exemplary embodiment, the server 105 can generate a down-sampled feature map corresponding to the image to be identified based on a discrete wavelet transform; generate a multi-head self-attention feature corresponding to the image to be identified based on the down-sampled feature map; transform the down-sampled feature map into a reconstructed feature map based on an inverse discrete wavelet transform; and generate a classification indication feature corresponding to the image to be identified based on the reconstructed feature map and the multi-head self-attention feature.
[0115] See also Figure 2 , Figure 2 The following schematically shows a flow chart of an image feature processing method according to an embodiment of the present application. Figure 2 As shown, the image feature processing method may include: steps S210 to S240.
[0116] Step S210: Generate a downsampled feature map corresponding to the image to be identified based on discrete wavelet transform.
[0117] Step S220: Generate a multi-head self-attention feature corresponding to the image to be identified based on the downsampled feature map.
[0118] Step S230: transforming the downsampled feature map into a reconstructed feature map based on inverse discrete wavelet transform.
[0119] Step S240: Generate classification indication features corresponding to the image to be identified based on the reconstructed feature map and the multi-head self-attention features.
[0120] The above method can be applied to visual tasks such as image classification, object detection and semantic segmentation, and is not limited to the embodiments of the present application.
[0121] Implementation Figure 2The method shown can perform reversible feature downsampling through discrete wavelet transform, and use inverse discrete wavelet transform to expand the receptive field to obtain a reconstructed feature map of the downsampled feature map, so that a high-precision classification indicator feature can be obtained by combining the reconstructed feature map and multi-head self-attention, which improves the feature extraction accuracy compared to the existing technology. In addition, it can also avoid feature information loss and enhance the expression ability of classification indicator features.
[0122] Next, the above steps of this exemplary embodiment are described in more detail.
[0123] In step S210, a down-sampled feature map corresponding to the image to be recognized is generated based on discrete wavelet transform.
[0124] Specifically, the Discrete Wavelet Transform (DWT) is obtained by discretizing the scale and displacement of the continuous wavelet transform according to the power of 2, so it is also called the binary wavelet transform (wavelettransform, WT). Compared with the short-time Fourier transform, WT can process signals through adaptive window sizes, without the need to process signals according to a fixed window size, and can provide a "time-frequency" window that changes with frequency. WT can make up for the complement of Fourier decomposition on non-stationary time series. By replacing the sine and cosine waves of Fourier decomposition with a set of attenuated orthogonal bases, it can better express the mutations and non-stationary parts in the sequence.
[0125] In addition, the image to be identified may be a picture containing a product, a picture containing facial features, or a picture containing text, etc., which is not limited in the embodiments of the present application. For example, when the image to be identified contains a product, the classification indication feature extracted by the method shown in the embodiments of the present application can be used to achieve more accurate product identification. When the solution is applied to intelligent product sorting, accurate product sorting can be achieved through accurate product identification. In addition, the present application does not limit the size of the image to be identified.
[0126] Optionally, before generating a downsampled feature map corresponding to the image to be identified based on discrete wavelet transform, the above method may also include the following steps: obtaining size information of the image to be identified, and if the parameters in the size information do not meet preset conditions, adjusting the parameters in the size information until the parameters in the size information (e.g., 224×224) meet the preset conditions; wherein the preset conditions can be used to limit the value range of one or more parameters in the size information.
[0127] As an optional embodiment, generating a downsampled feature map corresponding to the image to be identified based on discrete wavelet transform includes: obtaining a feature space corresponding to the image to be identified; generating a downsampled feature map corresponding to the feature space based on discrete wavelet transform. In this way, the feature space corresponding to the image to be identified can be based on the downsampled feature map, and the feature space can accurately represent the image to be identified, which is conducive to improving the accuracy of the downsampled feature map.
[0128] Specifically, feature space can be used as a feature representation method for images to be recognized.
[0129] As an optional embodiment, obtaining a feature space corresponding to the image to be identified includes: generating a feature map to be identified corresponding to the image to be identified; performing a first linear transformation on the feature map to be identified using preset parameters to obtain a feature space. In this way, an effective feature space can be obtained based on the linear transformation, thereby improving the feature extraction efficiency.
[0130] Specifically, the feature map to be identified can be expressed as Wherein, H, w, and D are the height, width, and dimension of the feature map to be identified X, respectively. In addition, the feature map to be identified can be used to represent the image to be identified by a vector / matrix. In addition, the preset parameters may include one or more parameters, which are not limited in the embodiment of the present application. For example, the preset parameters may be represented as a parameter matrix Based on this, the first linear transformation of the feature map to be identified is performed by preset parameters to obtain the feature space, including Perform the first linear transformation on the feature map X to be identified, and obtain the feature space
[0131] As an optional embodiment, generating a feature map to be identified corresponding to the image to be identified includes: performing multi-layer convolution processing on the image to be identified to obtain an image block sequence of the image to be identified; and generating the feature map to be identified based on the image block sequence. In this way, the image to be identified can be decomposed, and the feature map to be identified is formed by the image block sequence obtained by the decomposition, which can improve the processing efficiency of the image to be identified.
[0132] Specifically, multi-layer convolution processing is performed on the image to be recognized to obtain an image block sequence of the image to be recognized, including: performing convolution processing on the image to be recognized in sequence based on multiple convolution layers to obtain an image block sequence of the image to be recognized; wherein the multiple convolution layers may correspond to different convolution kernels or to the same convolution kernel, which is not limited in the embodiments of the present application.
[0133] In addition, generating a feature map to be identified according to the image block sequence includes: splicing each image block in the image block sequence according to a predetermined order between the image blocks, so as to obtain the feature map to be identified.
[0134] As an optional embodiment, a downsampled feature map corresponding to a feature space is generated based on a discrete wavelet transform, including: decomposing the feature space into a wavelet sub-band feature set by a discrete wavelet transform; fusing the wavelet sub-band feature set into a target feature; and performing convolution processing on the target feature to obtain a downsampled feature map. In this way, a reversible downsampled feature map can be obtained based on a discrete wavelet transform, which can reduce the amount of calculation and improve the utilization of computing resources on the one hand; on the other hand, it can avoid losing the high-frequency part of the image (such as the texture detail part) and avoid adverse effects on the translation equivariance of the recognition network.
[0135] Specifically, the wavelet sub-band feature set is fused into the target feature, including: High frequency band characteristics Low frequency band features High frequency band features Concatenate features according to dimensions to obtain target features
[0136] In addition, the target features are convolved to obtain a downsampled feature map, including: Perform convolution processing to obtain the downsampled feature map X c .
[0137] As an optional embodiment, the wavelet sub-band feature set includes low-frequency band row features, high-frequency band row features, low-frequency band column features, and high-frequency band column features. The feature space is decomposed into a wavelet sub-band feature set by discrete wavelet transform, including: obtaining low-frequency band row features corresponding to the feature space based on the low-pass filter in the discrete wavelet transform; obtaining high-frequency band row features corresponding to the feature space based on the high-pass filter in the discrete wavelet transform; obtaining low-frequency band column features corresponding to the feature space based on the low-pass filter in the discrete wavelet transform; and obtaining high-frequency band column features corresponding to the feature space based on the high-pass filter in the discrete wavelet transform. In this way, the acquisition of low-frequency band row features, high-frequency band row features, low-frequency band column features, and high-frequency band column features can be achieved, and the features of the image to be identified from four aspects can be obtained, which can reduce the spatial complexity and avoid the loss of detail features, so as to enhance the feature expression ability and generalization ability of the model when applied to the network model.
[0138] Specifically, based on the low-pass filter in the discrete wavelet transform, the low-frequency band row features corresponding to the feature space are obtained, including: based on the low-pass filter in the discrete wavelet transform Perform row encoding on the feature space to obtain low-frequency band row features Among them, the low frequency band features Used to characterize contour features in images.
[0139] Based on the high-pass filter in discrete wavelet transform, high-frequency band row features corresponding to the feature space are obtained, including: based on the high-pass filter in discrete wavelet transform Perform row encoding on the feature space to obtain high-frequency band row features
[0140] Based on the low-pass filter in discrete wavelet transform, the low-frequency band column features corresponding to the feature space are obtained, including: based on the low-pass filter in discrete wavelet transform Column encoding of feature space to obtain low-frequency band column features
[0141] Based on the high-pass filter in discrete wavelet transform, the high-frequency band column features corresponding to the feature space are obtained, including: based on the high-pass filter in discrete wavelet transform Column encoding of feature space to obtain high-frequency band column features Among them, the high-frequency band features Low frequency band features High frequency band features Used to jointly characterize the detail features in the image.
[0142] In step S220, a multi-head self-attention feature corresponding to the image to be identified is generated according to the downsampled feature map.
[0143] Specifically, the multi-head self-attention features include the self-attention features output by each head of the self-attention network.
[0144] As an optional embodiment, a multi-head self-attention feature corresponding to an image to be identified is generated based on a downsampled feature map, including: performing a second linear transformation on the feature map to be identified to obtain a query feature; performing a third linear transformation on the downsampled feature map to obtain a key feature and a value feature; and generating a multi-head self-attention feature based on the query feature, the key feature, and the value feature. In this way, the query feature, the key feature, and the value feature can be obtained based on multiple linear transformations, and self-attention calculation can be performed based on the query feature, the key feature, and the value feature, thereby obtaining a multi-head self-attention feature that can be used to generate a classification indication feature, which is beneficial to improving the accuracy of the classification indication feature.
[0145] Specifically, performing a second linear transformation on the identification feature map to obtain the query feature includes: performing a second linear transformation on the identification feature map to obtain the query feature Performing a third linear transformation on the downsampled feature map to obtain a key feature and a value feature, including: performing a third linear transformation on the downsampled feature map to obtain a key feature Sum value characteristics Among them, W q , W k and W vExpressed as a parameter matrix, n=H×W, m=H / 2×W / 2.
[0146] It should be noted that the first linear transformation, the second linear transformation, and the third linear transformation mentioned in the present application may correspond to different linear transformation algorithms or may correspond to the same linear transformation algorithm, and the embodiments of the present application are not limited thereto.
[0147] As an optional embodiment, a multi-head self-attention feature is generated according to the query feature, key feature and value feature, including: performing feature division on the query feature, key feature and value feature respectively according to the preset division rule to obtain a query sub-feature set, a key sub-feature set and a value sub-feature set; inputting each feature in the query sub-feature set into the corresponding self-attention network, and inputting each feature in the key sub-feature set into the corresponding self-attention network, and inputting each feature in the value sub-feature set into the corresponding self-attention network, so that each head self-attention network generates a corresponding self-attention feature to obtain a multi-head self-attention feature. In this way, multi-head self-attention calculation can be realized, and more accurate classification indication features can be extracted based on the calculated multi-head self-attention features, which is conducive to further improving the accuracy of the classification indication features.
[0148] Specifically, according to the preset division rules, the query features, the key features and the value features are respectively divided into query sub-feature sets, key sub-feature sets and value sub-feature sets, including: dividing the query features into Divide into N dimensions h The query sub-feature set of the query sub-feature, the key feature Divide into N dimensions h The key sub-feature set of the key sub-feature, the value feature Divide into N dimensions h The value sub-feature set of the value sub-feature. Among them, N h It represents the number of heads of the self-attention network.
[0149] Based on this, each feature in the query sub-feature set is input into the corresponding self-attention network, and each feature in the key sub-feature set is input into the corresponding self-attention network, and each feature in the value sub-feature set is input into the corresponding self-attention network, so that each head self-attention network generates a corresponding self-attention feature, and a multi-head self-attention feature is obtained, including: the feature corresponding to the jth head in the query sub-feature set Input the jth self-attention network and convert the feature corresponding to the jth head in the key sub-feature set Input the jth self-attention network and take the feature corresponding to the jth head in the sub-feature set Input the jth self-attention network so that the jth self-attention network is based on the expression Calculate the jth self-attention feature head j , each attention feature head % It can form a multi-head self-attention feature; the number of self-attention networks can be multiple, % is expressed as a positive integer, D h The dimension used to represent the self-attention network for each head.
[0150] In step S230, the downsampled feature map is transformed into a reconstructed feature map based on inverse discrete wavelet transform.
[0151] Specifically, the downsampled feature map is transformed into a reconstructed feature map based on an inverse discrete wavelet transform, comprising: transforming the downsampled feature map into a reconstructed feature map based on an inverse discrete wavelet transform (IDWT); Transformed into reconstructed feature map X r , X r have All details of the website.
[0152] In step S240, a classification indication feature corresponding to the image to be identified is generated based on the reconstructed feature map and the multi-head self-attention feature.
[0153] As an optional embodiment, generating a classification indication feature corresponding to the image to be identified based on the reconstructed feature map and the multi-head self-attention feature includes: fusing the multi-head self-attention feature and the reconstructed feature map to obtain a target comprehensive feature; generating a classification indication feature based on the target comprehensive feature. In this way, the expression ability of the target comprehensive feature can be improved to generate a more accurate classification indication feature.
[0154] Specifically, the multi-head self-attention features and the reconstructed feature map are integrated to obtain the target comprehensive features, including: The multi-head self-attention features and the reconstructed feature maps are concatenated according to the dimensions, and then the linear transformation matrix parameters are used. right Perform linear transformation to obtain the target comprehensive features
[0155] Then, the classification indicator feature is generated according to the target comprehensive feature, including: w (Q,K,V,X r ) into the expression WaveletsBlock(X) = MultiHead w (XW q ,X c W k ,X c W v,X r ) to calculate the classification indicator feature WaveletsBlock(X).
[0156] As an optional embodiment, generating classification indication features according to the target comprehensive features includes: performing discrete wavelet transform on the target comprehensive features to obtain a new down-sampled feature map; generating a new multi-head self-attention feature corresponding to the image to be identified according to the new down-sampled feature map; transforming the new down-sampled feature map into a new reconstructed feature map by inverse discrete wavelet transform; generating a classification indication feature corresponding to the image to be identified according to the new reconstructed feature map and the new multi-head self-attention feature. In this way, multiple rounds of discrete wavelet transforms can be implemented for the features, thereby optimizing the expressiveness of the features.
[0157] Specifically, the above steps may be executed by the next Wavelets network block, and there is a preset triggering sequence between the next Wavelets network block and the current Wavelets network block.
[0158] As an optional embodiment, the method further includes: if the current feature processing stage is not the final feature processing stage, inputting the classification indication feature into the next feature processing stage. In this way, multi-stage feature processing can be achieved, which is conducive to achieving the effectiveness of the final classification indication feature when applied to image recognition, and helps to more accurately achieve image recognition.
[0159] Specifically, the current feature processing stage may include multiple network blocks, each of which may be used to execute the technical solution of the embodiment of the present application, and the network structure of the next feature processing stage may be the same as or different from that of the current feature processing stage. The application scenario mentioned in this application may include multiple stages, and the embodiment of this application does not limit the number of stages.
[0160] In addition, optionally, the above method may also include: if the current feature processing stage is the final feature processing stage, extracting the classification indication feature through a feedforward neural network to obtain a new classification indication feature; and determining the category to which the image to be identified belongs based on the new classification indication feature.
[0161] As an optional embodiment, after generating the classification indication feature of the image to be identified according to the target comprehensive feature, the method further includes: determining the category to which the image to be identified belongs according to the classification indication feature. In this way, image recognition based on the classification indication feature can be realized, which is conducive to improving the accuracy of image recognition.
[0162] For example, if the image to be identified is a product image, the categories may include multiple categories (such as daily necessities, pet supplies, kitchen supplies, etc.), and the category corresponding to the image to be identified may belong to one of the above multiple categories.
[0163] As an optional embodiment, determining the category of the image to be identified based on the classification indication feature includes: extracting the classification indication feature through a feedforward neural network to obtain a new classification indication feature; and determining the category of the image to be identified based on the new classification indication feature. In this way, the new classification indication feature after feature extraction by the feedforward neural network can be used to implement image recognition, which is conducive to improving the accuracy and efficiency of image recognition.
[0164] Specifically, a feedforward neural network (FFN) is a neural network in which neurons are arranged in layers. Each neuron is only connected to the neurons in the previous layer. It receives the output of the previous layer and outputs it to the next layer. There is no feedback between the layers. The classification indicator feature is extracted by the feedforward neural network. The new classification indicator feature obtained is different from the elements contained in the classification indicator feature before feature extraction. The new classification indicator feature can be used for more accurate image recognition.
[0165] Among them, determining the category of the image to be identified based on the new classification indication feature includes: generating a classification sequence through the new classification indication feature, the classification sequence includes multiple probability values, each probability value corresponds to a category, and each probability value is used to characterize the probability that the image to be identified belongs to the category; determining the category corresponding to the largest probability value in the classification sequence as the category corresponding to the image to be identified.
[0166] See also Figure 3 , Figure 3 The following schematically shows a flow chart of an image feature processing method according to another embodiment of the present application. Figure 3 As shown, the image feature processing method includes steps S310 to S338.
[0167] Step S310: performing multi-layer convolution processing on the image to be identified to obtain an image block sequence of the image to be identified.
[0168] Step S312: Generate a feature map to be identified according to the image block sequence.
[0169] Step S314: performing a first linear transformation on the feature map to be identified using preset parameters to obtain a feature space.
[0170] Step S316: Based on the low-pass filter in the discrete wavelet transform, obtain the low-frequency band row features corresponding to the feature space; based on the high-pass filter in the discrete wavelet transform, obtain the high-frequency band row features corresponding to the feature space; based on the low-pass filter in the discrete wavelet transform, obtain the low-frequency band column features corresponding to the feature space; based on the high-pass filter in the discrete wavelet transform, obtain the high-frequency band column features corresponding to the feature space.
[0171] Step S318: Fusing the wavelet sub-band feature set into the target feature.
[0172] Step S320: Perform convolution processing on the target features to obtain a downsampled feature map.
[0173] Step S322: Perform a second linear transformation on the feature graph to be identified to obtain a query feature.
[0174] Step S324: Perform a third linear transformation on the downsampled feature map to obtain key features and value features.
[0175] Step S326: feature division is performed on the query feature, key feature and value feature respectively according to the preset division rule to obtain a query sub-feature set, a key sub-feature set and a value sub-feature set.
[0176] Step S328: Input each feature in the query sub-feature set into the corresponding self-attention network, input each feature in the key sub-feature set into the corresponding self-attention network, and input each feature in the value sub-feature set into the corresponding self-attention network, so that each head self-attention network generates a corresponding self-attention feature to obtain a multi-head self-attention feature.
[0177] Step S330: transforming the downsampled feature map into a reconstructed feature map based on inverse discrete wavelet transform.
[0178] Step S332: Fuse the multi-head self-attention features and reconstruct the feature map to obtain the target comprehensive features.
[0179] Step S334: Generate classification indication features based on the target comprehensive features.
[0180] Step S336: extracting the classification indication feature through a feedforward neural network to obtain a new classification indication feature.
[0181] Step S338: Determine the category to which the image to be identified belongs according to the new classification indication feature.
[0182] It should be noted that steps S310 to S338 are Figure 2 The steps and their embodiments shown correspond to each other. For the specific implementation of steps S310 to S338, please refer to Figure 2 The steps and embodiments shown are not described in detail here.
[0183] It can be seen that implementation Figure 3The method shown can perform reversible feature downsampling through discrete wavelet transform, and use inverse discrete wavelet transform to expand the receptive field to obtain a reconstructed feature map of the downsampled feature map, so that a high-precision classification indicator feature can be obtained by combining the reconstructed feature map and multi-head self-attention, which improves the feature extraction accuracy compared to the existing technology. In addition, it can also avoid feature information loss and enhance the expression ability of classification indicator features.
[0184] See also Figure 4 , Figure 4 A network architecture diagram for implementing the image feature processing method of the present application is schematically shown. Figure 4 As shown, the network architecture for implementing the image feature processing method can be called a Wavelets network block 400, and the Wavelets network block 400 can include: Linear410, Linear420, DWT430, Linear440, Linear450, IDWT460, Softmax470, and Linear480.
[0185] Specifically, Linear420 is used to perform a first linear transformation on the feature map to be identified through preset parameters to obtain a feature space. DWT430 is used to obtain the low-frequency band row features corresponding to the feature space based on the low-pass filter in the discrete wavelet transform; obtain the high-frequency band row features corresponding to the feature space based on the high-pass filter in the discrete wavelet transform; obtain the low-frequency band column features corresponding to the feature space based on the low-pass filter in the discrete wavelet transform; obtain the high-frequency band column features corresponding to the feature space based on the high-pass filter in the discrete wavelet transform. Linear410 is used to perform a second linear transformation on the feature map to be identified to obtain a query feature. Linear440 is used to perform a third linear transformation on the downsampled feature map to obtain a key feature. Linear450 is used to perform a third linear transformation on the downsampled feature map to obtain a value feature.
[0186] Furthermore, the Wavelets network block 400 can concatenate the query features and the key features based on the channel dimension. Softmax 470 is used to normalize the concatenated query features and key features. The Wavelets network block 400 can also concatenate the normalized results and the value features based on the channel dimension to obtain the features to be divided.
[0187] Furthermore, the Wavelets network block 400 can perform feature division on the query features, key features and value features in each feature to be divided according to the preset division rules, and obtain a query sub-feature set, a key sub-feature set and a value sub-feature set, and input each feature in the query sub-feature set into the corresponding self-attention network, and input each feature in the key sub-feature set into the corresponding self-attention network, and input each feature in the value sub-feature set into the corresponding self-attention network, so that each head self-attention network generates a corresponding self-attention feature to obtain a multi-head self-attention feature.
[0188] Furthermore, IDWT460 is used to transform the downsampled feature map into a reconstructed feature map based on inverse discrete wavelet transform. Linear480 is used to fuse the multi-head self-attention features and the reconstructed feature map to obtain the target comprehensive features, and generate classification indication features based on the target comprehensive features.
[0189] It can be seen that implementation Figure 4 The network structure shown can perform reversible feature downsampling through discrete wavelet transform, and use inverse discrete wavelet transform to expand the receptive field to obtain the reconstructed feature map of the downsampled feature map, so that the reconstructed feature map and multi-head self-attention can be combined to obtain a classification indicator feature with higher accuracy, which improves the feature extraction accuracy compared to the existing technology. In addition, it can also avoid the loss of feature information and enhance the expression ability of the classification indicator feature.
[0190] See also Figure 5 , Figure 5 Another network architecture diagram for implementing the image feature processing method of the present application is schematically shown. Figure 5 As shown, the network architecture for implementing the image feature processing method may include: a first stage 510, a second stage 520, ..., an Nth stage 530, where N is a positive integer. The first stage 510 includes a Wavelets network block 511, a feedforward neural network 512, ..., a Wavelets network block 513, and a feedforward neural network 514; the second stage 520 includes a Wavelets network block 521, a feedforward neural network 522, ..., a Wavelets network block 523, and a feedforward neural network 524; the Nth stage 530 includes a Wavelets network block 531, a feedforward neural network 532, ..., a Wavelets network block 533, and a feedforward neural network 534.
[0191] Specifically, the internal structure of each Wavelets network block in the first stage 510, the second stage 520, ..., and the Nth stage 530 is as follows: Figure 4 For the specific steps of each Wavelets network block, please refer to Figure 4In addition, each feedforward neural network can be used for feature extraction, so as to facilitate the transmission of the obtained features to the next Wavelets network block in the same stage or the Wavelets network block in the next stage, so as to realize multi-stage processing of image features, thereby helping to improve the accuracy of image recognition.
[0192] In addition, the first stage 510, the second stage 520, ..., the Nth stage 530 may correspond to different feature resolutions, for example, the first stage corresponds to a feature resolution The second stage corresponds to feature resolution The third stage corresponds to feature resolution The fourth stage corresponds to feature resolution
[0193] See also Figure 6 , Figure 6 The structure block diagram of the image feature processing device according to an embodiment of the present application is schematically shown. Figure 2 The method shown corresponds to Figure 6 As shown, the image feature processing device 600 includes:
[0194] A feature sampling unit 601 is used to generate a down-sampled feature map corresponding to the image to be identified based on discrete wavelet transform;
[0195] A feature generation unit 602, configured to generate a multi-head self-attention feature corresponding to the image to be recognized according to the downsampled feature map;
[0196] A feature transformation unit 603, configured to transform the downsampled feature map into a reconstructed feature map based on an inverse discrete wavelet transform;
[0197] The feature generation unit 602 is also used to generate classification indication features corresponding to the image to be identified based on the reconstructed feature map and the multi-head self-attention features.
[0198] It can be seen that implementation Figure 6 The device shown can perform reversible feature downsampling through discrete wavelet transform, and use inverse discrete wavelet transform to expand the receptive field to obtain a reconstructed feature map of the downsampled feature map, so that a high-precision classification indicator feature can be obtained by combining the reconstructed feature map and multi-head self-attention, which improves the feature extraction accuracy compared to the existing technology. In addition, it can also avoid feature information loss and enhance the expression ability of the classification indicator feature.
[0199] In an exemplary embodiment of the present application, the feature sampling unit 601 generates a downsampled feature map corresponding to the image to be identified based on discrete wavelet transform, including:
[0200] Obtaining the feature space corresponding to the image to be identified;
[0201] A downsampled feature map corresponding to the feature space is generated based on discrete wavelet transform.
[0202] It can be seen that by implementing this optional embodiment, the feature space corresponding to the image to be identified can be based on the downsampled feature map, and the feature space can accurately characterize the image to be identified, which is conducive to improving the accuracy of the downsampled feature map.
[0203] In an exemplary embodiment of the present application, the feature sampling unit 601 obtains a feature space corresponding to the image to be identified, including:
[0204] Generate a feature map to be identified corresponding to the image to be identified;
[0205] A first linear transformation is performed on the feature map to be identified using preset parameters to obtain a feature space.
[0206] It can be seen that by implementing this optional embodiment, an effective feature space can be obtained based on linear transformation, thereby improving feature extraction efficiency.
[0207] In an exemplary embodiment of the present application, the feature sampling unit 601 generates a feature map to be identified corresponding to the image to be identified, including:
[0208] Perform multi-layer convolution processing on the image to be identified to obtain an image block sequence of the image to be identified;
[0209] Generate a feature map to be identified based on the image block sequence.
[0210] It can be seen that by implementing this optional embodiment, the image to be identified can be decomposed, and the image block sequence obtained by the decomposition can be used to form a feature map to be identified, which can improve the processing efficiency of the image to be identified.
[0211] In an exemplary embodiment of the present application, the feature sampling unit 601 generates a downsampled feature map corresponding to the feature space based on discrete wavelet transform, including:
[0212] Decompose the feature space into a set of wavelet sub-band features through discrete wavelet transform;
[0213] The wavelet sub-band feature set is fused into the target feature;
[0214] The target features are convolved to obtain a downsampled feature map.
[0215] It can be seen that by implementing this optional embodiment, a reversible down-sampled feature map can be obtained based on discrete wavelet transform, which can reduce the amount of calculation and improve the utilization of computing resources on the one hand; on the other hand, it can avoid the loss of high-frequency parts of the image (such as texture details) and avoid adverse effects on the translation and equivariance of the recognition network.
[0216] In an exemplary embodiment of the present application, the wavelet sub-band feature set includes low-frequency band row features, high-frequency band row features, low-frequency band column features, and high-frequency band column features. The feature sampling unit 601 decomposes the feature space into a wavelet sub-band feature set by discrete wavelet transform, including:
[0217] Based on the low-pass filter in discrete wavelet transform, the low-frequency band row features corresponding to the feature space are obtained;
[0218] Based on the high-pass filter in discrete wavelet transform, high-frequency band row features corresponding to the feature space are obtained;
[0219] Obtain low-frequency band column features corresponding to the feature space based on the low-pass filter in discrete wavelet transform;
[0220] The high-frequency band column features corresponding to the feature space are obtained based on the high-pass filter in discrete wavelet transform.
[0221] It can be seen that by implementing this optional embodiment, it is possible to obtain low-frequency band row features, high-frequency band row features, low-frequency band column features, and high-frequency band column features, and obtain features that describe the image to be identified from four aspects, which can reduce spatial complexity and avoid the loss of detail features, so as to enhance the feature expression ability and generalization ability of the model when applied to the network model.
[0222] In an exemplary embodiment of the present application, the feature generation unit 602 generates a multi-head self-attention feature corresponding to the image to be recognized according to the downsampled feature map, including:
[0223] Perform a second linear transformation on the feature map to be identified to obtain the query feature;
[0224] Perform the third linear transformation on the downsampled feature map to obtain key features and value features;
[0225] Generate multi-head self-attention features based on query features, key features, and value features.
[0226] It can be seen that by implementing this optional embodiment, query features, key features and value features can be obtained based on multiple linear transformations. Self-attention calculation can be performed based on the query features, key features and value features, thereby obtaining multi-head self-attention features that can be used to generate classification indication features, which is beneficial to improving the accuracy of the classification indication features.
[0227] In an exemplary embodiment of the present application, the feature generation unit 602 generates a multi-head self-attention feature according to the query feature, the key feature and the value feature, including:
[0228] According to the preset division rules, the query features, key features and value features are respectively divided into query sub-feature sets, key sub-feature sets and value sub-feature sets;
[0229] Each feature in the query sub-feature set is input into the corresponding self-attention network, each feature in the key sub-feature set is input into the corresponding self-attention network, and each feature in the value sub-feature set is input into the corresponding self-attention network, so that each head self-attention network generates a corresponding self-attention feature and obtains a multi-head self-attention feature.
[0230] It can be seen that the implementation of this optional embodiment can realize multi-head self-attention calculation, and more accurate classification indication features can be extracted based on the calculated multi-head self-attention features, which is conducive to further improving the accuracy of the classification indication features.
[0231] In an exemplary embodiment of the present application, the feature generation unit 602 generates a classification indication feature corresponding to the image to be identified according to the reconstructed feature map and the multi-head self-attention feature, including:
[0232] Fusion of multi-head self-attention features and reconstructed feature maps to obtain comprehensive target features;
[0233] Generate classification indicator features based on the target comprehensive features.
[0234] It can be seen that implementing this optional embodiment can improve the expression ability of the target comprehensive features to generate more accurate classification indication features.
[0235] In an exemplary embodiment of the present application, the feature generation unit 602 generates a classification indication feature according to the target comprehensive feature, including:
[0236] Perform discrete wavelet transform on the comprehensive features of the target to obtain a new downsampled feature map;
[0237] Generate a new multi-head self-attention feature corresponding to the image to be recognized according to the new downsampled feature map;
[0238] The new downsampled feature map is transformed into a new reconstructed feature map by inverse discrete wavelet transform;
[0239] A classification indication feature corresponding to the image to be identified is generated according to the new reconstructed feature map and the new multi-head self-attention feature.
[0240] It can be seen that by implementing this optional embodiment, multiple rounds of discrete wavelet transforms can be performed on the features, thereby optimizing the expressiveness of the features.
[0241] In an exemplary embodiment of the present application, the above-mentioned device further includes:
[0242] The feature input unit is used to input the classification indication feature into the next feature processing stage if the current feature processing stage is not the final feature processing stage.
[0243] It can be seen that the implementation of this optional embodiment can realize multi-stage feature processing, which is conducive to realizing the effectiveness of the final classification indication feature when applied to image recognition, and helps to realize more accurate recognition of the image.
[0244] In an exemplary embodiment of the present application, the above-mentioned device further includes:
[0245] The image recognition unit is used to determine the category to which the image to be recognized belongs according to the classification indication feature after the feature generation unit 602 generates the classification indication feature of the image to be recognized according to the target comprehensive feature.
[0246] It can be seen that implementing this optional embodiment can realize image recognition based on classification indication features, which is conducive to improving the accuracy of image recognition.
[0247] In an exemplary embodiment of the present application, the image recognition unit determines the category of the image to be recognized according to the classification indication feature, including:
[0248] The classification indicator feature is extracted through a feedforward neural network to obtain a new classification indicator feature;
[0249] Determine the category of the image to be identified based on the new classification indication features.
[0250] It can be seen that by implementing this optional embodiment, image recognition can be achieved using new classification indication features after feature extraction using a feedforward neural network, which is beneficial to improving the accuracy and efficiency of image recognition.
[0251] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.
[0252] Since the various functional modules of the image feature processing device of the exemplary embodiment of the present application correspond to the steps of the exemplary embodiment of the above-mentioned image feature processing method, for details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the above-mentioned image feature processing method of the present application.
[0253] See also Figure 7 , Figure 7 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown.
[0254] It should be noted that Figure 7 The computer system 700 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0255] like Figure 7 As shown, computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage part 708 into a random access memory (RAM) 703. Various programs and data required for system operation are also stored in RAM 703. CPU 701, ROM 702 and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.
[0256] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed, so that a computer program read therefrom is installed into the storage section 708 as needed.
[0257] In particular, according to an embodiment of the present application, the process described with reference to the flowchart above can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 709, and / or installed from a removable medium 711. When the computer program is executed by a central processing unit (CPU) 701, various functions defined in the method and apparatus of the present application are executed.
[0258] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiment.
[0259] It should be noted that the computer-readable medium shown in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0260] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0261] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.
[0262] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary technical means in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the aforementioned claims.
Claims
1. A method for processing image features, It is characterized in that include: Performing multi-layer convolution processing on the image to be identified to obtain an image block sequence of the image to be identified, generating a feature map to be identified corresponding to the image to be identified according to the image block sequence, performing a first linear transformation on the feature map to be identified by using preset parameters to obtain a feature space corresponding to the image to be identified, and generating a down-sampled feature map corresponding to the feature space based on discrete wavelet transform; Generating a multi-head self-attention feature corresponding to the image to be recognized according to the downsampled feature map; Transforming the downsampled feature map into a reconstructed feature map based on an inverse discrete wavelet transform; A classification indication feature corresponding to the image to be identified is generated according to the reconstructed feature map and the multi-head self-attention feature.
2. The method according to claim 1, It is characterized in that Generating a downsampled feature map corresponding to the feature space based on discrete wavelet transform, comprising: Decomposing the feature space into a set of wavelet sub-band features by discrete wavelet transform; fusing the wavelet sub-band feature set into a target feature; The target feature is subjected to convolution processing to obtain the down-sampled feature map.
3. The method according to claim 2, It is characterized in that The wavelet sub-band feature set includes low-frequency band row features, high-frequency band row features, low-frequency band column features, and high-frequency band column features. The feature space is decomposed into the wavelet sub-band feature set by discrete wavelet transform, including: Acquire the low-frequency band row feature corresponding to the feature space based on a low-pass filter in discrete wavelet transform; Acquire the high-frequency band row feature corresponding to the feature space based on a high-pass filter in discrete wavelet transform; Acquire the low-frequency band column features corresponding to the feature space based on a low-pass filter in discrete wavelet transform; The high-frequency band column features corresponding to the feature space are obtained based on a high-pass filter in discrete wavelet transform.
4. The method according to claim 1, It is characterized in that Generating a multi-head self-attention feature corresponding to the image to be recognized according to the downsampled feature map, comprising: Performing a second linear transformation on the feature graph to be identified to obtain a query feature; Performing a third linear transformation on the downsampled feature map to obtain a key feature and a value feature; The multi-head self-attention feature is generated according to the query feature, the key feature and the value feature.
5. The method according to claim 4, It is characterized in that Generating the multi-head self-attention feature according to the query feature, the key feature, and the value feature includes: According to a preset division rule, the query feature, the key feature and the value feature are respectively divided into features to obtain a query sub-feature set, a key sub-feature set and a value sub-feature set; Each feature in the query sub-feature set is input into the corresponding self-attention network, each feature in the key sub-feature set is input into the corresponding self-attention network, and each feature in the value sub-feature set is input into the corresponding self-attention network, so that each head self-attention network generates a corresponding self-attention feature to obtain the multi-head self-attention feature.
6. The method according to claim 1, It is characterized in that Generating a classification indication feature corresponding to the image to be identified according to the reconstructed feature map and the multi-head self-attention feature, comprising: Fusing the multi-head self-attention features and the reconstructed feature map to obtain a comprehensive target feature; The classification indication feature is generated according to the target comprehensive feature.
7. The method according to claim 6, It is characterized in that Generating the classification indication feature according to the target comprehensive feature includes: Performing discrete wavelet transform on the target comprehensive features to obtain a new down-sampled feature map; Generating a new multi-head self-attention feature corresponding to the image to be recognized according to the new down-sampled feature map; Transforming the new down-sampled feature map into a new reconstructed feature map by inverse discrete wavelet transform; A classification indication feature corresponding to the image to be identified is generated according to the new reconstructed feature map and the new multi-head self-attention feature.
8. The method according to claim 6, It is characterized in that The method further comprises: If the current feature processing stage is not the final feature processing stage, the classification indication feature is input into the next feature processing stage.
9. The method according to claim 6, It is characterized in that After generating the classification indication feature of the image to be identified according to the target comprehensive feature, the method further includes: The category to which the image to be identified belongs is determined according to the classification indication feature.
10. The method according to claim 9, It is characterized in that Determining the category to which the image to be identified belongs according to the classification indication feature includes: Extracting the classification indication feature through a feedforward neural network to obtain a new classification indication feature; The category to which the image to be identified belongs is determined according to the new classification indication feature.
11. An image feature processing device, It is characterized in that include: A feature sampling unit, configured to perform multi-layer convolution processing on an image to be identified to obtain an image block sequence of the image to be identified, generate a feature map to be identified corresponding to the image to be identified according to the image block sequence, perform a first linear transformation on the feature map to be identified by using preset parameters to obtain a feature space corresponding to the image to be identified, and generate a down-sampled feature map corresponding to the feature space based on discrete wavelet transform; A feature generation unit, configured to generate a multi-head self-attention feature corresponding to the image to be recognized according to the downsampled feature map; A feature conversion unit, used for converting the down-sampled feature map into a reconstructed feature map based on an inverse discrete wavelet transform; The feature generation unit is also used to generate a classification indication feature corresponding to the image to be identified based on the reconstructed feature map and the multi-head self-attention feature.
12. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
13. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
14. An electronic device, It is characterized in that include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1-10 by executing the executable instructions.
Citation Information
Patent Citations
Method, system and equipment for removing rain from rain map based on selective mechanism and attention mechanism
CN112862875A
Image semantic segmentation method and device, equipment and storage medium
CN113807354A