Adaptive attention feature fusion image extraction method, system, device and product

By performing channel and spatial attention operation on the image, and fusing it with local feature maps, adaptive adjustment parameters are determined based on image feature information, the problem of poor accuracy in image feature extraction in the prior art is solved, and the feature extraction ability of small-sized targets is improved.

CN120163989APending Publication Date: 2025-06-17PEOPLE'S INSURANCE COMPANY OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510213556.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art has problems with poor accuracy when extracting image features, especially when the effective area size of the detection target is small, the apparent area features are similar, and the difference between classes is not obvious.

Method used

By obtaining the original image and feature map of the detection target, channel attention operation and spatial attention operation are performed, channel attention feature map and spatial attention feature map are extracted, and the multi-channel local feature map is combined for fusion, and adaptive adjustment parameters are determined based on image feature information to realize adaptive attention feature fusion of feature maps.

Benefits of technology

The feature extraction capability of detection targets with small effective area sizes and similar features is improved, and the accuracy of image feature extraction is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163989A_ABST
    Figure CN120163989A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an adaptive attention feature fusion image extraction method, system, device and product. The method comprises the steps of obtaining an original image of a detection target and a corresponding feature image; performing channel attention operation on the feature map to obtain a channel attention feature map; performing space attention operation on the channel attention feature map to obtain a space attention feature map; extracting multi-channel local features of the feature map to obtain a supplementary local feature map; determining adaptive adjustment parameters according to the image feature information of the original image; the image feature information represents image features of the original image; according to the adaptive adjustment parameters, fusing the spatial attention feature map and the supplementary local feature map to obtain an adaptive attention feature fusion map; target object information is generated and output by processing the self-adaptive attention feature fusion image; the problem of low detection accuracy caused by the existing technical scheme is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a method, system, device, and product for extracting an adaptive attention feature fusion map. Background Art

[0002] With the development of artificial intelligence technology, great progress has been made in the research of vision detection based on deep learning.

[0003] In the existing technology solutions, when performing vision detection on an image, the attention mechanism is often used to obtain the key features in the image, and then the vision detection result is obtained.

[0004] However, the existing technology solutions have the problem of poor accuracy in extracting image features. Summary of the Invention

[0005] Embodiments of this application provide a method, system, device, and product for extracting an adaptive attention feature fusion map, which are used to solve the problem of poor accuracy in extracting image features.

[0006] In a first aspect, an embodiment of this application provides a method for extracting an adaptive attention feature fusion map, including: obtaining an original image of a detection target and a corresponding feature map; performing channel attention operation on the feature map to obtain a channel attention feature map; performing spatial attention operation on the channel attention feature map to obtain a spatial attention feature map; extracting multi-channel local features from the feature map to obtain a supplementary local feature map; determining an adaptive adjustment parameter according to the image feature information of the original image, where the image feature information represents the image features of the original image; and fusing the spatial attention feature map and the supplementary local feature map according to the adaptive adjustment parameter to obtain an adaptive attention feature fusion map.

[0007] In a possible implementation manner, the extracting multi-channel local features from the feature map to obtain a supplementary local feature map includes: based on a convolutional kernel with a receptive field of , extracting multi-channel local features from the feature map to obtain the supplementary local feature map, where N is a positive integer.

[0008] In a possible implementation manner, the determining an adaptive adjustment parameter according to the image feature information of the original image includes: obtaining the application scenario information of the detection target, where the application scenario information represents the application scenario features corresponding to the detection target; and determining the adaptive adjustment parameter according to the application scenario information and the image feature information.

[0009] In a possible implementation manner, determining the adaptive adjustment parameter according to the application scenario information and the image feature information includes: determining an initial adjustment parameter according to the application scenario information and the image feature information; updating the initial adjustment parameter according to the gradient optimization result of the loss function to obtain the adaptive adjustment parameter.

[0010] In a possible implementation manner, determining the initial adjustment parameter according to the application scenario information and the image feature information includes: generating a first complexity index according to the application scenario information; generating a second complexity index according to the image feature information; determining the initial adjustment parameter according to the first complexity index and the second complexity index.

[0011] In a second aspect, an embodiment of the present application provides an adaptive attention feature fusion map extraction system, and the system includes a preprocessing module and an adaptive attention feature fusion module;

[0012] The preprocessing module is used to obtain the original image of the detection target and the corresponding feature map;

[0013] The adaptive attention feature fusion module is used to perform channel attention operation on the feature map to obtain a channel attention feature map; perform spatial attention operation on the channel attention feature map to obtain a spatial attention feature map; extract multi-channel local features from the feature map to obtain a supplementary local feature map; determine an adaptive adjustment parameter according to the image feature information of the original image; the image feature information represents the image features of the original image; fuse the spatial attention feature map and the supplementary local feature map according to the adaptive adjustment parameter to obtain an adaptive attention feature fusion map.

[0014] In a third aspect, an embodiment of the present application provides an adaptive attention feature fusion map extraction device, including:

[0015] An acquisition unit is used to acquire the original image of the detection target and the corresponding feature map;

[0016] A processing unit is used to perform channel attention operation on the feature map to obtain a channel attention feature map; perform spatial attention operation on the channel attention feature map to obtain a spatial attention feature map; extract multi-channel local features from the feature map to obtain a supplementary local feature map; determine an adaptive adjustment parameter according to the image feature information of the original image; the image feature information represents the image features of the original image;

[0017] A generation unit is used to fuse the spatial attention feature map and the supplementary local feature map according to the adaptive adjustment parameter to obtain an adaptive attention feature fusion map.

[0018] In a possible implementation manner, when the processing unit extracts multi-channel local features from the feature map to obtain a supplementary local feature map, it is specifically configured to: based on a convolution kernel with a receptive field of , extract multi-channel local features from the feature map to obtain the supplementary local feature map; where N is a positive integer.

[0019] In a possible implementation manner, when the processing unit determines an adaptive adjustment parameter according to the image feature information of the original image, it is specifically configured to: obtain the application scenario information of the detection target, where the application scenario information characterizes the application scenario features corresponding to the detection target; determine the adaptive adjustment parameter according to the application scenario information and the image feature information.

[0020] In a possible implementation manner, when the processing unit determines the adaptive adjustment parameter according to the application scenario information and the image feature information, it is specifically configured to: determine an initial adjustment parameter according to the application scenario information and the image feature information; update the initial adjustment parameter according to the gradient optimization result of the loss function to obtain the adaptive adjustment parameter.

[0021] In a possible implementation manner, when the processing unit determines an initial adjustment parameter according to the application scenario information and the image feature information, it is specifically configured to: generate a first complexity index according to the application scenario information; generate a second complexity index according to the image feature information; determine the initial adjustment parameter according to the first complexity index and the second complexity index.

[0022] Fourthly, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0023] The memory stores computer execution instructions;

[0024] The processor executes the computer execution instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementation manners of the first aspect.

[0025] Fifthly, an embodiment of the present application provides a computer-readable storage medium, where computer execution instructions are stored in the computer-readable storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the above first aspect and / or various possible implementation manners of the first aspect.

[0026] Sixthly, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above first aspect and / or various possible implementation manners of the first aspect.

[0027] The adaptive attention feature fusion map extraction method, system, device and product provided by the embodiments of the present application obtain the original image of the detection target and the corresponding feature map; perform channel attention operation on the feature map to obtain a channel attention feature map; perform spatial attention operation on the channel attention feature map to obtain a spatial attention feature map; extract multi-channel local features from the feature map to obtain a supplementary local feature map; determine an adaptive adjustment parameter according to the image feature information of the original image, where the image feature information represents the image features of the original image; and fuse the spatial attention feature map and the supplementary local feature map according to the adaptive adjustment parameter to obtain an adaptive attention feature fusion map. On the basis of performing global feature processing and local feature processing on the feature map, an adaptive adjustment parameter corresponding thereto is determined based on the image feature information of the original image, and then the spatial attention feature map and the supplementary local feature map are fused based on the adaptive adjustment parameter to obtain an adaptive attention feature fusion map. That is, by supplementing the spatial attention features with local features, the feature extraction ability for detection targets with small effective region sizes, similar features in each apparent region, and insignificant inter-class feature differences is improved, and the problem of poor accuracy in image feature extraction in the existing technical solutions is solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0029] Figure 1 It is a schematic diagram of the scenario of the adaptive attention feature fusion map extraction method provided by the present application;

[0030] Figure 2 It is a flowchart of the adaptive attention feature fusion map extraction method provided by an embodiment of the present application;

[0031] Figure 3 For Figure 2 It is a schematic diagram of the specific implementation steps of step S105 in the illustrated embodiment;

[0032] Figure 4 For Figure 3 It is a schematic diagram of the specific implementation steps of step S1052 in the illustrated embodiment;

[0033] Figure 5 It is a structural diagram of the module for generating an adaptive attention feature fusion map provided by the present application;

[0034] Figure 6 It is a schematic diagram of the system structure of an adaptive attention feature fusion map extraction system provided by the present application;

[0035] Figure 7 Schematic structural diagram of an adaptive attention feature fusion map extraction device provided by an embodiment of the present application;

[0036] Figure 8 Schematic structural diagram of an electronic device provided by the present application.

[0037] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Specific Embodiments

[0038] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0039] In the technical solution of the present application, the processing of user personal information and data involved, such as collection, storage, use, processing, transmission, provision, and disclosure, all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0040] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant region, and corresponding operation entrances are provided for the user to choose to authorize or refuse.

[0041] The application scenarios of the embodiments of the present application will be explained below:

[0042] Figure 1 Scenario schematic diagram of the adaptive attention feature fusion map extraction method provided by the present application, as Figure 1As shown in the figure, the specific application scenario of this application is a scenario of object recognition for small-sized images with similar features. The execution subject of the method provided in the embodiments of this application can be an electronic control unit, a terminal device, or a server. Exemplarily, the scenario of object recognition for small-sized images with similar features is, for example, a scenario of recognizing a cow face picture and determining the identity information of the cow. Taking the terminal device as an example for the execution subject, that is, when recognizing the identity information of a cow, since the size of the effective area of the cow face in the cow face picture is small, the apparent area features between each cow face are similar, and the inter-class feature differences are not obvious. Therefore, after the user captures the face image of the cow through the terminal device, the terminal device performs feature extraction and feature fusion on the cow face image based on the adaptive attention feature fusion map extraction method provided in this application, and then obtains the corresponding adaptive attention feature fusion map, and further obtains the recognition result based on the adaptive attention feature fusion map, that is, determines the identity information of the cow.

[0043] In the prior art, when performing visual detection on an image, the attention mechanism is often used to obtain the key features in the image. However, the attention mechanism adopted in the prior art is more suitable for detection scenarios when the target area is large, the apparent area features are obvious, and the inter-class feature differences are large, and its detection result is relatively accurate; combined with the above scenario, when the size of the effective area of the detection target is small, the features of each apparent area are similar, and the inter-class feature differences are not obvious, there will be a problem of poor accuracy in image feature extraction.

[0044] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0045] Figure 2 It is a flowchart of the adaptive attention feature fusion map extraction method provided in an embodiment of this application. As Figure 2 shown, the execution subject of the adaptive attention feature fusion map extraction method provided in this embodiment can be an electronic control unit, a terminal device, or a server. Exemplarily, this embodiment uses the terminal device as the execution subject of the method in this embodiment for description. The adaptive attention feature fusion map extraction method provided in this embodiment includes the following steps:

[0046] Step S101, obtain the original image of the detection target and the corresponding feature map.

[0047] Step S102, perform channel attention operation on the feature map to obtain a channel attention feature map.

[0048] Exemplarily, the channel attention operation uses a 3D arrangement to preserve the three-dimensional information of the feature map and amplifies the dependencies across the dimensional channel space through a perceptron with a double hidden layer; specifically, for example, the input feature corresponding to the feature map is , perform a channel attention operation on the feature map, and the corresponding mapping representation obtained is , and the output feature corresponding to the channel attention feature map is calculated by formula (1),

[0049] (1)

[0050] where is the per-pixel dot product.

[0051] Step S103: Perform a spatial attention operation on the channel attention feature map to obtain a spatial attention feature map.

[0052] Exemplarily, the spatial attention operation uses two convolutional layers to fuse the spatial information of the channel attention feature map, deletes the max pooling operation to preserve the spatial feature information, and then uses group convolution with channel shuffling to improve the processing efficiency; specifically, for example, the output feature corresponding to the channel attention feature map is , perform a spatial attention operation on the channel attention feature map, and the corresponding mapping representation obtained is , and the output feature corresponding to the spatial attention feature map is calculated by formula (2),

[0053] (2)

[0054] where is the per-pixel multiplication.

[0055] In the steps of this embodiment, since deleting the max pooling operation will cause a significant increase in the number of parameters of the spatial attention sub-module, an increase in the computational amount, and thus a decrease in the feature extraction efficiency of the attention mechanism, group convolution with channel shuffling is used to improve the processing efficiency.

[0056] Step S104: Extract multi-channel local features from the feature map to obtain a supplementary local feature map.

[0057] Exemplarily, extract multi-channel local features from the feature map to obtain a supplementary local feature map and the corresponding output feature; for example, extract multi-channel local features from the feature map through a local feature descriptor, and then obtain a supplementary local feature map.

[0058] In a possible implementation manner, the specific implementation steps of step S104 include based on a receptive field of The convolutional kernel extracts multi-channel local features from the feature map to obtain a supplementary local feature map, where N is a positive integer; specifically, for example, when N is 7, based on the receptive field being The convolutional kernel extracts multi-channel local features from the feature map, and the input feature corresponding to the feature map is , and the mapping of the convolutional kernel to the feature map is expressed as , and the output feature corresponding to the supplementary local feature map is calculated by formula (3),

[0059] (3)

[0060] Step S105: Determine an adaptive adjustment parameter according to the image feature information of the original image; the image feature information characterizes the image features of the original image.

[0061] Step S106: Fuse the spatial attention feature map and the supplementary local feature map according to the adaptive adjustment parameter to obtain an adaptive attention feature fusion map.

[0062] Exemplarily, the image feature information is used to indicate the image features of the original image, that is, to indicate the image complexity of the original image. Furthermore, in the case of complex image content and large noise interference, local features may be more effective; while in the case of simple image content and clear targets, global features may be more reliable. Then, the corresponding adaptive adjustment parameter is determined; and further, according to the adaptive adjustment parameter, the spatial attention feature map and the supplementary local feature map are fused to obtain an adaptive attention feature fusion map, where the spatial attention feature map corresponds to the global feature removed, and the supplementary local feature map corresponds to the local feature. Specifically, for example, based on the image features of the original image, determine the adaptive adjustment parameter corresponding to the spatial attention feature map , the adaptive adjustment parameter corresponding to the supplementary local feature map , and then the output feature corresponding to the adaptive attention feature fusion map is calculated by formula (4),

[0063] (4)

[0064] Among them, is the output feature corresponding to the supplementary local feature map; is the output feature corresponding to the spatial attention feature map; is the pixel-by-pixel addition.

[0065] In a possible implementation manner, Figure 3 is Figure 2 a schematic diagram of the specific implementation steps of step S105 in the embodiment shown, as Figure 3As shown, the specific implementation steps of step S105 include:

[0066] Step S1051: Obtain the application scenario information of the detection target, where the application scenario information represents the application scenario features corresponding to the detection target.

[0067] Step S1052: Determine the adaptive adjustment parameters according to the application scenario information and the image feature information.

[0068] Exemplarily, the application scenario information is used to indicate the application scenario features corresponding to the detection target. Since different application scenarios have different requirements for the accuracy and robustness of the detection target recognition, correspondingly, the requirements for global features and local features are also different. Specifically, for example, in autonomous driving, high requirements are placed on the real-time performance and accuracy of the detection target, and it may be necessary to comprehensively consider global and local features; while in medical image analysis, the accurate recognition of the target is more important, so more reliance on local features is required. Furthermore, according to the application scenario information and the image feature information, the corresponding adaptive adjustment parameters are determined. Further still, the spatial attention feature map and the supplementary local feature map can be fused according to the adaptive adjustment parameters to obtain the adaptive attention feature fusion map.

[0069] In another possible implementation manner, Figure 4 For Figure 3 a schematic diagram of the specific implementation steps of step S1052 in the illustrated embodiment, as Figure 4 shown, the specific implementation steps of step S1052 include:

[0070] Step S10521: Determine the initial adjustment parameters according to the application scenario information and the image feature information.

[0071] Step S10522: Update the initial adjustment parameters according to the gradient optimization result of the loss function to obtain the adaptive adjustment parameters.

[0072] Exemplarily, according to the application scenario features of the detection target and the image features of the original image, the corresponding initial adjustment parameters are determined. Furthermore, on the basis of determining the initial adjustment parameters, the initial adjustment parameters are updated through the gradient optimization result of the loss function to obtain the adaptive adjustment parameters, which simplifies the optimization process of randomly determining the adjustment parameters and then determining the adaptive adjustment parameters through the loss function, and improves the determination speed of the adaptive adjustment parameters.

[0073] In a possible implementation manner, Figure 5 is a structural diagram of the module for generating the adaptive attention feature fusion map provided by this application. As Figure 5 shown, based on the channel attention sub-module corresponding to the channel attention operation, the input features of the feature map Process to obtain the output features corresponding to the channel attention feature map ; then, based on the spatial attention sub-module corresponding to the spatial attention operation, process the output features to obtain the output features corresponding to the spatial attention feature map ; extract multi-channel local features from the input feature F1 of the feature map based on the local feature information complementation module to obtain the output features corresponding to the supplementary local feature map ; further, the adaptive multi-scale feature fusion module is based on the adaptive adjustment parameter corresponding to the spatial attention feature map , the adaptive adjustment parameter corresponding to the supplementary local feature map , fuse the spatial attention feature map and the supplementary local feature map to obtain an adaptive attention feature fusion map, that is, obtain the corresponding output features .

[0074] Further, in another possible implementation manner, the specific implementation manner of step S10521 includes: generating a first complexity index according to the application scenario information; generating a second complexity index according to the image feature information; determining an initial adjustment parameter according to the first complexity index and the second complexity index. Specifically, for example, based on the data of the complexity of the preset application scenario features, quantify the complexity of the application scenario features corresponding to the application scenario information to generate a first complexity index; based on the data of the complexity of the preset image features, quantify the complexity of the image features of the original image corresponding to the image feature information to generate a second complexity index; then, according to the first complexity index and the second complexity index, the initial adjustment parameter can be determined. For example, by performing addition, or multiplication, or weighted calculation on the first complexity index and the second complexity index, the initial adjustment parameter is calculated.

[0075] In this embodiment, the original image of the detection target and the corresponding feature map are obtained; channel attention operation is performed on the feature map to obtain a channel attention feature map; spatial attention operation is performed on the channel attention feature map to obtain a spatial attention feature map; multi-channel local features are extracted from the feature map to obtain a supplementary local feature map; an adaptive adjustment parameter is determined according to the image feature information of the original image, where the image feature information represents the image features of the original image; and the spatial attention feature map and the supplementary local feature map are fused according to the adaptive adjustment parameter to obtain an adaptive attention feature fusion map. Based on the global feature processing and local feature processing of the feature map, the corresponding adaptive adjustment parameter is determined based on the image feature information of the original image, and then the spatial attention feature map and the supplementary local feature map are fused based on the adaptive adjustment parameter to obtain an adaptive attention feature fusion map. That is, by supplementing the spatial attention feature with the local feature, the feature extraction ability for detection targets with small effective region sizes, similar apparent region features among individual apparent regions, and insignificant inter-class feature differences is improved, and the problem of poor accuracy in image feature extraction existing in the prior art solutions is solved.

[0076] Furthermore, the method provided by the embodiments of the present application can be applied to Figure 1 the application scenario of detecting a cow face image to determine the identity of a cow as shown in the figure. This solves the problem that it is difficult to determine the identity of a cow due to the small size of the effective region of the cow face in the cow face picture, the similar apparent region features among individual cow faces, and the insignificant inter-class feature differences, and further leads to the inability to determine whether the detected and recognized cow is an insured cow. It can be understood that the present application does not specifically limit the detected and recognized object (detection target). For example, the detection target can also be a sheep (sheep face image) or a pig (pig face image).

[0077] The present application provides an adaptive attention feature fusion map extraction system, which includes a preprocessing module and an adaptive attention feature fusion module. The preprocessing module is used to obtain the original image of the detection target and the corresponding feature map. The adaptive attention feature fusion module is used to perform channel attention operation on the feature map to obtain a channel attention feature map; perform spatial attention operation on the channel attention feature map to obtain a spatial attention feature map; extract multi-channel local features from the feature map to obtain a supplementary local feature map; determine an adaptive adjustment parameter according to the image feature information of the original image, where the image feature information represents the image features of the original image; and fuse the spatial attention feature map and the supplementary local feature map according to the adaptive adjustment parameter to obtain an adaptive attention feature fusion map.

[0078] Exemplarily, Figure 6 is a schematic structural diagram of an adaptive attention feature fusion map extraction system provided by the present application, as shown in Figure 6As shown in (a), the adaptive attention feature fusion map extraction system consists of a preprocessing module and an adaptive attention feature fusion module; the adaptive attention feature fusion map extraction system at least includes: the preprocessing module is at least determined by an input module, a focus module, a CBS module, and a C3 module; the input module is used to indicate the input of the original image of the detection target, the focus module is used to indicate the focusing process on the original image, the CBS module is used to indicate the downsampled feature map, the width and height are halved, and the channel depth is increased to further extract deeper features; the C3 module is used to further comprehensively extract deep features, prevent feature loss, and increase the fusion of multi-scale features. Among them, as Figure 6 shown in (b), the C3 module is at least determined by 3 CBS modules, 1 Concat module, and 1 module. Further, the Bottleneck module uses residual connection. There are two CBS modules in the Bottleneck. The first CBS module is convolution, which reduces the number of channels to half of the original. The second CBS module is convolution, which doubles the number of channels. First, dimensionality reduction is beneficial for the convolution kernel to better understand the feature information, and then dimensionality increase is beneficial for extracting more detailed features; the adaptive attention feature fusion module (Auto Mixed Global Attention Mechanism, abbreviated as AMGAM module) is used to perform adaptive attention feature fusion processing on the feature map to output the adaptive attention feature fusion map.

[0079] Figure 7 FIG. is a schematic structural diagram of an adaptive attention feature fusion map extraction device provided by an embodiment of the present application. As Figure 7 shown, the adaptive attention feature fusion map extraction device 3 provided in this embodiment includes:

[0080] An acquisition unit 31, configured to acquire the original image of the detection target and the corresponding feature map;

[0081] A processing unit 32, configured to perform channel attention operation on the feature map to obtain a channel attention feature map; perform spatial attention operation on the channel attention feature map to obtain a spatial attention feature map; extract multi-channel local features from the feature map to obtain a supplementary local feature map; determine an adaptive adjustment parameter according to the image feature information of the original image; the image feature information represents the image features of the original image;

[0082] A generation unit 33, configured to fuse the spatial attention feature map and the supplementary local feature map according to the adaptive adjustment parameter to obtain an adaptive attention feature fusion map.

[0083] In a possible implementation manner, when the processing unit 32 extracts multi-channel local features from the feature map to obtain a supplementary local feature map, it is specifically configured to: based on a convolution kernel with a receptive field of extract multi-channel local features from the feature map to obtain a supplementary local feature map; N is a positive integer.

[0084] In a possible implementation manner, when the processing unit 32 determines an adaptive adjustment parameter according to the image feature information of the original image, it is specifically configured to: obtain the application scenario information of the detection target, where the application scenario information represents the application scenario features corresponding to the detection target; determine the adaptive adjustment parameter according to the application scenario information and the image feature information.

[0085] In a possible implementation manner, when the processing unit 32 determines an adaptive adjustment parameter according to the application scenario information and the image feature information, it is specifically configured to: determine an initial adjustment parameter according to the application scenario information and the image feature information; update the initial adjustment parameter according to the gradient optimization result of the loss function to obtain the adaptive adjustment parameter.

[0086] In a possible implementation manner, when the processing unit 32 determines an initial adjustment parameter according to the application scenario information and the image feature information, it is specifically configured to: generate a first complexity index according to the application scenario information; generate a second complexity index according to the image feature information; determine the initial adjustment parameter according to the first complexity index and the second complexity index.

[0087] Wherein, the acquisition unit 31, the processing unit 32, and the generation unit 33 are connected in sequence. The adaptive attention feature fusion map extraction device 3 provided in this embodiment can execute the technical solutions of the method embodiments as Figures 2 - 5 shown in any one of them. The implementation principle and technical effects are similar, and will not be elaborated here.

[0088] Figure 8 This is a schematic structural diagram of the electronic device provided in this application. As Figure 8 shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the device 50 further includes a communication component 503. Among them, the processor 501, the memory 502, and the communication component 503 are connected through a bus 504.

[0089] In a specific implementation process, at least one processor 501 executes the computer execution instructions stored in the memory 502, so that at least one processor 501 executes the above method.

[0090] For the specific implementation process of the processor 501, reference can be made to the above method embodiments. The implementation principle and technical effects are similar, and will not be elaborated here in this embodiment.

[0091] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU for short), or may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor, or implemented by the combination of the hardware and software modules in the processor.

[0092] The memory may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.

[0093] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0094] This application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0095] This application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the above method is implemented.

[0096] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disk. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0097] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be part of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0098] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the couplings or direct couplings or communication connections shown or discussed among each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0099] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0100] Furthermore, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0101] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0102] Those of ordinary skill in the art will understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the aforementioned storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0103] Finally, it should be noted that: After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other implementation manners of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed by the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. An adaptive attention feature fusion graph extraction method, characterized in that: The method comprises: Get the original image of the detection target and the corresponding feature map; Performing a channel attention operation on the feature map to obtain a channel attention feature map; Performing a spatial attention operation on the channel attention feature map to obtain a spatial attention feature map; Extracting multi-channel local features from the feature map to obtain a supplementary local feature map; Determining an adaptive adjustment parameter according to image feature information of the original image, wherein the image feature information represents image features of the original image; According to the adaptive adjustment parameters, the spatial attention feature map and the supplementary local feature map are fused to obtain an adaptive attention feature fusion map.

2. The method according to claim 1, characterized in that The step of extracting multi-channel local features from the feature map to obtain a supplementary local feature map includes: Based on the receptive field A convolution kernel is used to extract multi-channel local features from the feature map to obtain the supplementary local feature map; N is a positive integer.

3. The method according to claim 1, characterized in that The step of determining the adaptive adjustment parameter according to the image feature information of the original image includes: Acquire application scenario information of the detection target, where the application scenario information represents application scenario characteristics corresponding to the detection target; The adaptive adjustment parameter is determined according to the application scenario information and the image feature information.

4. The method according to claim 3, characterized in that The step of determining the adaptive adjustment parameter according to the application scenario information and the image feature information includes: Determining initial adjustment parameters according to the application scenario information and the image feature information; The initial adjustment parameters are updated according to the gradient optimization result of the loss function to obtain the adaptive adjustment parameters.

5. The method according to claim 4, characterized in that The determining of initial adjustment parameters according to the application scenario information and the image feature information includes: Generate a first complexity index according to the application scenario information; generating a second complexity index according to the image feature information; The initial adjustment parameter is determined according to the first complexity index and the second complexity index.

6. An adaptive attention feature fusion graph extraction system, characterized in that: The system includes a preprocessing module and an adaptive attention feature fusion module; The preprocessing module is used to obtain the original image of the detection target and the corresponding feature map; The adaptive attention feature fusion module is used to perform a channel attention operation on the feature map to obtain a channel attention feature map; perform a spatial attention operation on the channel attention feature map to obtain a spatial attention feature map; Extract multi-channel local features from the feature map to obtain a supplementary local feature map; determine adaptive adjustment parameters based on the image feature information of the original image; the image feature information characterizes the image features of the original image; and fuse the spatial attention feature map and the supplementary local feature map based on the adaptive adjustment parameters to obtain an adaptive attention feature fusion map.

7. An adaptive attention feature fusion map extraction device, characterized in that: include: An acquisition unit, used to acquire an original image of a detection target and a corresponding feature map; A processing unit, configured to perform a channel attention operation on the feature map to obtain a channel attention feature map; Performing a spatial attention operation on the channel attention feature map to obtain a spatial attention feature map; Extracting multi-channel local features from the feature map to obtain a supplementary local feature map; determining an adaptive adjustment parameter according to image feature information of the original image; the image feature information represents the image features of the original image; A generating unit is used to fuse the spatial attention feature map and the supplementary local feature map according to the adaptive adjustment parameters to obtain an adaptive attention feature fusion map.

8. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 5.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 5 when executed by a processor.

10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 5 when being executed by a processor.