Multi-scale attention feature extraction method and device, electronic equipment and storage medium

By adopting a multi-scale attention feature extraction method in online live broadcast, the features of different scales are extracted and attention scores are calculated, and misclassification problems caused by rough classification of illegal content and different target sizes in the existing technology are solved, and a higher accuracy of feature extraction and classification accuracy are achieved.

CN120088500APending Publication Date: 2025-06-03GUANGZHOU HUYA INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510137441.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The classification of illegal content in the existing technology in online live broadcast is too rough, which cannot meet the business's needs for detailed category classification, and the model is difficult to flexibly control costs and benefits. In the identification of violations, the target size difference is significant, resulting in the impact of identification accuracy.

Method used

The multi-scale attention feature extraction method is used to obtain the first feature map of the target image, and extract features of different scales using convolutional layer and global pooling operations, calculate the attention score of the channel feature vector, and finally fusion is obtained to obtain the channel attention multi-scale feature map.

Benefits of technology

By combining multi-scale features with channel attention information, the model's ability to capture key features of the image is enhanced, the model's feature extraction accuracy and classification accuracy are improved, and the misclassification problem caused by large differences in target size is effectively solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088500A_ABST
    Figure CN120088500A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, and provides a multi-scale attention feature extraction method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a first feature map of a target image; inputting the first feature map into a serial convolution layer for convolution processing to obtain a second feature map; the second feature map comprises a plurality of scale features; performing global pooling operation on each scale feature of the second feature map to obtain a channel feature vector; performing convolution processing on the channel feature vector to obtain a convolution feature vector, calculating an attention score of the obtained convolution feature vector, and obtaining a channel-to-attention score graph; and fusing the second feature map with the channel attention score map to obtain a channel attention multi-scale feature map. Multi-scale information in an image is gradually and deeply captured through serial convolution, and a channel attention mechanism is fused, so that the model can more accurately identify and classify targets of different scales when facing a scene with large target size difference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and more particularly, to a multi-scale attention feature extraction method, apparatus, electronic device, and storage medium. Background Art

[0002] Currently, the content security issue in webcasting is becoming increasingly prominent. Therefore, AI technology is used to accurately classify the illegal content in webcasting to achieve effective content security control. Fine-grained classification aims to divide the illegal behaviors into more detailed categories, providing more flexible and accurate control means for the platform, so as to effectively address the content security challenges. According to the understanding of the business, various illegal behavior categories can be divided into multiple specific categories.

[0003] In the existing technical solutions, the classification of various illegal contents is too rough to meet the business requirements for detailed category classification, and the thresholds are difficult to adjust, so it is difficult to flexibly control the costs and benefits therein. At the same time, in the classification tasks of various illegal behaviors, the sizes of the recognition targets vary significantly. For example, in a certain type of illegal behavior dataset, the corresponding target often occupies a large pixel area in the image; while in another illegal behavior dataset, the pixel proportion of the corresponding target is relatively small. In addition, during live broadcasts, scenarios such as split screens and small-window live broadcasts usually occur, and these illegal behaviors appear in different scales in the image. If the model only focuses on the features at a single specific scale, the model is likely to miss the targets at other scales, thus affecting the recognition accuracy. Summary of the Invention

[0004] The present invention aims to overcome at least one defect (shortcoming) of the above-mentioned prior art, and provides a multi-scale attention feature extraction method, apparatus, electronic device, and storage medium for improving the accuracy of fine-grained classification of the model.

[0005] According to the first aspect of the present application, a multi-scale attention feature extraction method is provided, and the method includes:

[0006] Obtain a first feature map of a target image;

[0007] Input the first feature map into a serial convolutional layer for convolutional processing to obtain a second feature map; the second feature map includes several scale features;

[0008] Perform global pooling operations on each of the scale features of the second feature map to obtain channel feature vectors;

[0009] Perform convolutional processing on the channel feature vectors to obtain convolutional feature vectors, calculate the attention scores of the convolutional feature vectors, and obtain a channel attention score map;

[0010] Fuse the second feature map with the channel attention score map to obtain a channel attention multi-scale feature map.

[0011] By obtaining the first feature map of the target image, using convolutional layers and global pooling operations to extract features of different scales, then calculating the attention scores of the channel feature vectors, and finally fusing to obtain a channel attention multi-scale feature map. By combining multi-scale features with channel attention information, the model's ability to capture key features of the image is enhanced, thereby improving the accuracy of the model's feature extraction.

[0012] Optionally, the serial convolutional layer includes a first convolutional layer and a second convolutional layer connected in series; the first convolutional layer is used to adjust the number of channels of the first feature map; the second convolutional layer is used to extract several scale features of the first feature map after adjusting the number of channels.

[0013] Through the first convolution and the second convolution connected in series, as the convolutional layer deepens, the image regions concerned by each layer of convolution are also continuously expanding, enabling the model to capture larger-scale information, thereby enhancing the model's ability to capture image details and its ability to integrate global information, and further improving the accuracy of the model in image classification tasks.

[0014] Optionally, the convolution kernel size of the first convolutional layer is 1x1, and the convolution kernel size of the second convolutional layer is 3x3.

[0015] Adjust the number of channels through the first convolutional layer with a convolution kernel size of 1x1 to reduce the high computation brought by subsequent serial convolutions. Gradually expand the receptive field by stacking the second convolutional layer with a convolution kernel size of 3x3, thereby capturing features of different scales in the target image.

[0016] Optionally, performing a global pooling operation on each scale feature of the second feature map to obtain a channel feature vector specifically includes:

[0017] Perform a global average pooling operation on each scale feature to obtain a first feature vector;

[0018] Perform a global maximum pooling operation on each scale feature to obtain a second feature vector;

[0019] Add the first feature vector and the second feature vector to obtain the channel feature vector corresponding to each scale feature.

[0020] By performing a global average pooling operation and a global maximum pooling operation on each scale feature, the information within the channel can be effectively aggregated.

[0021] Optionally, the specific process of performing convolutional processing on the channel feature vector to obtain a convolutional feature vector includes:

[0022] Apply the channel feature vector corresponding to each of the said scale features to a third convolutional layer for convolution operations; input the output of the third convolutional layer into a fourth convolutional layer for convolution operations to obtain the corresponding convolutional feature vector.

[0023] Perform convolution operations on the said channel feature vector through a fifth convolutional layer to fuse the information between channels. To maintain the compatibility of the feature vector with the original network or subsequent processing modules, perform convolution operations through a sixth convolutional layer to adjust the number of channels of the fused feature vector to the original quantity, thereby achieving the consistency of the channel dimension, which lays the foundation for calculating the attention score of each channel in the subsequent calculation.

[0024] Optionally, the attention score of the said channel feature vector is calculated through a softmax activation function.

[0025] Optionally, the fusion of the said second feature map and the channel attention score map to obtain a channel attention multi-scale feature map includes:

[0026] Multiply each of the said scale features of the second feature map by the elements of the corresponding channel attention score map to generate the corresponding channel attention multi-scale feature map, and the shape of the channel attention multi-scale feature map is the same as the shape of the first feature map.

[0027] According to the second aspect of the present application, there is provided a multi-scale attention feature extraction device, the device includes:

[0028] A first feature map acquisition module for acquiring a first feature map of a target image;

[0029] A second feature map acquisition module for inputting the first feature map into a serial convolutional layer for convolution processing to obtain a second feature map; the second feature map includes several scale features;

[0030] A channel feature vector acquisition module for performing global pooling operations on each of the said scale features of the second feature map to obtain channel feature vectors;

[0031] An attention score map acquisition module for performing convolution processing on the said channel feature vector to obtain a convolutional feature vector, calculating the attention score of the convolutional feature vector to obtain a channel attention score map;

[0032] A fusion module for fusing the second feature map and the channel attention score map to obtain a channel attention multi-scale feature map.

[0033] According to the third aspect of the present application, there is provided an electronic device, including:

[0034] A memory for storing one or more computer programs;

[0035] A processor, which, when the one or more computer programs are executed by the processor, implements the multi-scale attention feature extraction method described in the first aspect above.

[0036] According to the fourth aspect of the present application, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the multi-scale attention feature extraction method described in the first aspect when executed.

[0037] Based on any of the above aspects, the multi-scale attention feature extraction method, apparatus, electronic device and storage medium provided by the embodiments of the present application have the following beneficial effects:

[0038] 1. By obtaining the first feature map of the target image, using convolutional layers and global pooling operations to extract features of different scales, then calculating the attention scores of the channel feature vectors, and finally fusing to obtain the channel attention multi-scale feature map. By combining multi-scale features with channel attention information, the model's ability to capture key features of the image is enhanced, thereby improving the accuracy of the model's feature extraction and the accuracy of model classification.

[0039] 1. By serially convolving to gradually capture more delicate and rich multi-scale information in the image, the model's perception ability of image details is enhanced, enabling the model to more effectively understand and analyze target features at different scales. By introducing the channel attention mechanism, when the model faces scenarios with significant differences in target sizes and large classification difficulties, the model can more accurately identify and classify targets at different scales, thus effectively solving the problem of misclassification caused by large differences in target sizes.

[0040] 2. The multi-scale attention feature extraction method provided by the present application has high modularity and flexibility, can be integrated into other backbone networks or basic modules, and by fusing multi-scale features and channel attention, enhances the model's classification and recognition ability for complex image content. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0042] Figure 1 It is a schematic application scenario diagram of the multi-scale attention feature extraction method provided in this embodiment.

[0043] Figure 2 It is a schematic structural diagram of the model of the multi-scale attention feature extraction method provided in this embodiment.

[0044] Figure 3 It is a schematic diagram of the receptive field change of the serial convolution receptive field of the multi-scale attention feature extraction method provided in this embodiment.

[0045] Figure 4 It is a specific flowchart of the multi-scale attention feature extraction method provided in this embodiment.

[0046] Figure 5 It is a schematic structural diagram of the multi-scale attention feature extraction device provided in this embodiment.

[0047] Figure 6 It is a schematic structural diagram of the electronic device provided in this embodiment. Detailed implementation manners

[0048] The accompanying drawings of this application are only for illustrative purposes and should not be construed as a limitation to this application. To better illustrate the following embodiments, some components in the drawings are omitted, enlarged or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0049] In order to enable those skilled in the art to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0050] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above accompanying drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0051] In the classification task of live broadcast scenarios, there is usually a large difference in the scales of the targets. When traditional convolutional stacking models process targets with large scale differences, it is often difficult to balance global features and detailed features, thus limiting the performance of the models. For the recognition of lying live broadcast behavior, the recognition network pays more attention to capturing the overall features of the person and their surrounding environment, so the attention area of the model is usually large. In contrast, the recognition of smoking behavior requires the model to focus more on key local features such as the hands and mouth, which have a relatively small pixel ratio. Given the obvious differences in scale features between lying live broadcast and smoking behaviors, if a single-scale feature is used to predict both lying live broadcast and smoking behaviors simultaneously, the model often fails to achieve its best performance.

[0052] This embodiment provides a technical solution that can solve the above problems. The following will combine the accompanying drawings to elaborate on the specific implementation of the present application in detail.

[0053] Exemplarily, it is a schematic diagram of an application scenario of a multi-scale attention feature extraction method provided by an embodiment of the present application. As Figure 1 shown, the application scenario at least includes a server 100 and a terminal 200 that can communicate with the server 100. The server 100 has an image processing function and can also have a data transmission function for video streams and audio streams; the terminal 200 has a streaming media playback function and can also have an image processing function.

[0054] It can be understood that the server 100 can be an independent electronic device or a cluster composed of multiple electronic devices; the terminal 200 can be a smart phone terminal, a personal computer, a tablet computer, a vehicle-mounted terminal, etc., but is not limited thereto.

[0055] In an implementable manner, the server 100 and the terminal 200 can respectively execute the multi-scale attention feature extraction method provided by the embodiment of the present application. Alternatively, part of the multi-scale attention feature extraction method provided by the embodiment of the present application is executed in the server 100, and part is executed in the terminal 200.

[0056] As Figure 4 shown, this embodiment provides a multi-scale attention feature extraction method, which can include the following steps:

[0057] S1. Obtain a first feature map of the target image.

[0058] In this embodiment, the corresponding feature map can be extracted from the target image through the convolutional layer or pooling layer of the convolutional neural network. As Figure 2 shown, the dimensions of the first feature map include C, H, and W, where C represents the number of feature map channels, and H and W respectively represent the height and width of the feature map.

[0059] S2. Input the first feature map into a serial convolutional layer for convolutional processing to obtain a second feature map; the second feature map includes several scale features.

[0060] Specifically, the serial convolutional layer includes a first convolutional layer and a second convolutional layer connected in series; the first convolutional layer is used to adjust the number of channels of the feature map; the second convolutional layer is used to extract several scale features of the first feature map after adjusting the number of channels.

[0061] As Figure 3 shown, when the serial convolutional layer performs convolutional processing, as the convolutional layers are stacked, the receptive field range of the deep convolutional operation on the original feature map gradually expands, so the deep features can contain information on a larger scale.

[0062] In a specific implementation, the serial convolutional layer may include a convolutional layer with a kernel size of 1x1 and three convolutional layers with a kernel size of 3x3. Among them, the convolutional layer with a kernel size of 1x1 is mainly used to adjust the number of channels of the first feature map and reduce the computational complexity of the subsequent serial 3x3 convolutional layer operations. By performing convolutional processing on the first feature map through three convolutional layers with a kernel size of 3x3, multi-scale feature information can be further extracted from the adjusted first feature map. In this embodiment, by continuously passing the first feature map of C x H x W through multiple convolutional layers connected in series for convolutional processing, a second feature map of 4x C / / 4x H x W can be obtained, where C / / 4 represents the integer result of dividing the number of channels C of the first feature map by 4; 4x C / / 4 represents 4 different scale feature representations, and each scale feature has C / / 4 channels.

[0063] Through the processing of the serial convolutional layer, the model can effectively capture information at different scales from the input feature map and provide multiple scale feature representations for subsequent classification tasks.

[0064] S3. Perform global pooling operations on each of the scale features of the second feature map to obtain channel feature vectors.

[0065] Performing global pooling operations on each of the scale features of the second feature map to obtain channel feature vectors specifically includes:

[0066] Performing global average pooling operations on each of the scale features to obtain a first feature vector;

[0067] Performing global maximum pooling operations on each of the scale features to obtain a second feature vector;

[0068] Adding the first feature vector and the second feature vector to obtain the channel feature vector corresponding to each of the scale features.

[0069] Through global average pooling and global maximum pooling operations, two feature vectors describing channel characteristics are respectively generated, thereby effectively aggregating the information within the channels. By adding the two generated feature vectors, a channel feature vector corresponding to each of the scale features is obtained, and this channel feature vector can be used as a description of the channel information.

[0070] S4. Perform a convolution operation on the channel feature vector to obtain a convolution feature vector, and calculate the attention score of the obtained convolution feature vector to obtain a channel attention score map.

[0071] Apply the channel feature vector corresponding to each of the scale features to a third convolutional layer for a convolution operation; input the output of the third convolutional layer into a fourth convolutional layer for a convolution operation to obtain a corresponding convolution feature vector.

[0072] In the specific implementation process, the features of each scale global pooling can be respectively convolved using a third convolutional layer with a kernel size of 1x1, so as to fuse the relevant information between channels. Then, a fourth convolutional layer with a kernel size of 1×1 is used again for a convolution operation to restore the number of channels of the fused feature vector to the original number, so as to ensure that the channel dimension of the fused channel feature vector is consistent with the channel dimension of the first feature map, laying a foundation for calculating the attention score of each channel subsequently.

[0073] In this embodiment, the softmax activation function can be used to calculate the attention score of the channel feature vector, so as to obtain the attention score of each channel feature vector. This attention score represents the importance of the feature vector, that is, the higher the score, the more important the feature vector.

[0074] S5. Fuse the second feature map and the channel attention score map to obtain a channel attention multi-scale feature map.

[0075] Specifically, multiply each element of each scale feature map by the corresponding channel attention score of each element to generate a channel attention multi-scale feature map, and the shape of the channel attention multi-scale feature map is the same as the shape of the first feature map. By multiplying the second feature map and the channel attention score map element by element, the corresponding attention score is used to weight each channel feature, so that each scale feature map of the channel attention multi-scale feature map incorporates the weight information of the channels, thereby generating a multi-scale feature map with channel attention.

[0076] The multi-scale attention feature extraction method provided in this embodiment captures more delicate and rich multi-scale information in the image step by step through serial convolution, enhancing the model's perception ability of image details, enabling the model to more effectively understand and analyze target features at different scales. By introducing the channel attention mechanism, when the model faces scenarios with significant differences in target sizes and large classification difficulties, the model can more accurately identify and classify targets at different scales, thus effectively solving the problem of misclassification caused by large differences in target sizes.

[0077] Next, a specific embodiment is combined to illustrate the technical solution of the multi-scale attention feature extraction method provided in this application.

[0078] As Figure 5 shown, the embodiment of this application also provides a multi-scale attention feature extraction device 210. Optionally, the multi-scale attention feature extraction device 210 may include:

[0079] A first feature map acquisition module 211, configured to acquire a first feature map of a target image;

[0080] A second feature map acquisition module 212, configured to input the first feature map into a serial convolution layer for convolution processing to obtain a second feature map; the second feature map includes several scale features;

[0081] A channel feature vector acquisition module 213, configured to perform global pooling operations on each of the scale features of the second feature map to obtain channel feature vectors;

[0082] An attention score map acquisition module 214, configured to perform convolution processing on the channel feature vectors to obtain convolution feature vectors, calculate the attention scores of the convolution feature vectors, and obtain a channel attention score map;

[0083] A fusion module 215, configured to fuse the second feature map with the channel attention score map to obtain a channel attention multi-scale feature map.

[0084] It can be understood that the above device embodiment and the above method embodiment can correspond to each other. Similar descriptions of the device embodiment can refer to the method embodiment. To avoid repetition, it will not be elaborated here. The multi-scale attention feature extraction device provided in the embodiment of this application can execute the multi-scale attention feature extraction method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the execution method. The functional modules of the multi-scale attention feature extraction device can be implemented in the form of hardware, can be implemented by instructions in the form of software, and can also be implemented by a combination of hardware and software modules.

[0085] Specifically, each step of the method embodiment of the present application can be completed by the hardware integrated logic circuit and / or software instructions in the processor, and the steps of the multi-scale attention feature extraction method in combination with the embodiment of the present application can be directly embodied as a hardware encoding processor for execution, or a combination of hardware and software modules in the encoding processor for execution. Optionally, the software module can be located in a random access memory, a read-only memory, a programmable read-only memory, a flash memory, an electrically erasable programmable memory, a register, or other storage media. The storage medium is located in the memory, and the processor reads the information in the memory, and completes the steps in the above method embodiment in combination with its hardware.

[0086] The present application embodiment provides an electronic device 310, whose structure is as follows: Figure 6 The electronic device 310 may be Figure 1 The server 100 or the terminal 200 is shown.

[0087] like Figure 6 As shown, the electronic device 310 includes a memory 311, a processor 312, a communication module 313 and an input / output interface 314, etc. Optionally, the memory 311, the processor 312, the communication module 313 and the input / output interface 314 can be connected and communicated through a bus 315.

[0088] The memory 311 is used to store one or more computer programs and transmit the code of the computer program to the processor 312; when the one or more computer programs are executed by the processor 312, the multi-scale attention feature extraction method in the embodiment of the present application is implemented.

[0089] Optionally, the electronic device 310 can be connected to a network via a communication module 313 to communicate with other devices, such as a terminal or a server, through the network to achieve data interaction. The electronic device 310 can be various forms of digital computers, such as desktop computers, servers, workstations, mainframe computers or other types of computers. The electronic device 310 can also be various forms of mobile terminals, such as smart phones, tablet computers, wearable devices (such as helmets, glasses, watches, etc.) and other similar mobile terminals.

[0090] Optionally, the electronic device 310 may be connected to required input / output devices, such as a keyboard, a display device, etc., through the input / output interface 314. The electronic device 310 itself may have a display device and may also externally connect other display devices through the input / output interface 314. Optionally, a storage device, such as a hard disk, etc., may also be connected through the input / output interface 314, so that the data in the electronic device 310 can be stored in the storage device, or the data in the storage device can be read, and the data in the storage device can also be stored in the memory 311. It can be understood that the input / output interface 314 may be a wired interface or a wireless interface. According to different actual application scenarios, the devices connected to the input / output interface 314 may be components of the electronic device 310 or external devices connected to the electronic device 310 when needed.

[0091] Optionally, the memory 311 may be a volatile memory and / or a non-volatile memory. The volatile memory may be a random access memory, etc., and the non-volatile memory may be a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, or a flash memory, etc.

[0092] Optionally, the computer program stored in the processor 312 may be divided into one or more modules. The one or more modules are stored in the memory 311 and executed by the processor 312 to complete the method provided by its own embodiment. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the computer program instruction segments are used to describe the execution process of the computer program in the electronic device 310.

[0093] Optionally, the processor 312 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 312 include but are not limited to a central processing unit, a graphics processing unit, a digital signal processor, various special artificial intelligence computing chips, various processors running machine learning model algorithms, and may also be any suitable controller, microcontroller, processor, etc. The processor 312 executes each method and process of this embodiment. Exemplarily, such as a multi-scale attention feature extraction method of an embodiment of the present application.

[0094] Optionally, the bus 315 may include a path for transmitting information. The bus 315 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. According to different functions, the bus 315 may be divided into an address bus, a data bus, a control bus, and the like.

[0095] Obviously, the above embodiments of the present application are merely examples for clearly illustrating the technical solutions of the present application, rather than limitations on the specific implementation manners of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the claims of the present application shall be included within the protection scope of the claims of the present application.

Claims

1. A multi-scale attention feature extraction method, characterized in that: The method comprises: Obtaining a first feature map of a target image; Inputting the first feature map into a serial convolution layer for convolution processing to obtain a second feature map; the second feature map includes several scale features; Performing a global pooling operation on each scale feature of the second feature map to obtain a channel feature vector; Convolution processing is performed on the channel feature vector to obtain a convolution feature vector, and an attention score of the convolution feature vector is calculated to obtain a channel attention score map; The second feature map is fused with the channel attention score map to obtain a channel attention multi-scale feature map.

2. A multi-scale attention feature extraction method according to claim 1, characterized in that: The serial convolution layer includes a first convolution layer and a second convolution layer connected in series; the first convolution layer is used to adjust the number of channels of the first feature map; and the second convolution layer is used to extract several scale features of the first feature map after the number of channels is adjusted.

3. A multi-scale attention feature extraction method according to claim 2, characterized in that: The convolution kernel size of the first convolution layer is 1x1, and the convolution kernel size of the second convolution layer is 3x3.

4. A multi-scale attention feature extraction method according to claim 1, characterized in that: The performing a global pooling operation on each scale feature of the second feature map to obtain a channel feature vector specifically includes: Performing a global average pooling operation on each of the scale features to obtain a first feature vector; Performing a global maximum pooling operation on each of the scale features to obtain a second feature vector; The first feature vector and the second feature vector are added to obtain a channel feature vector corresponding to each scale feature.

5. A multi-scale attention feature extraction method according to claim 1, characterized in that: The convolution processing of the channel feature vector to obtain the convolution feature vector specifically includes: The channel feature vector corresponding to each scale feature is convolved by the third convolution layer; the output of the third convolution layer is input to the fourth convolution layer for convolution operation to obtain the corresponding convolution feature vector.

6. A multi-scale attention feature extraction method according to claim 1, characterized in that: The attention score of the channel feature vector is calculated by the softmax activation function.

7. A multi-scale attention feature extraction method according to any one of claims 1 to 6, characterized in that: The step of fusing the second feature map with the channel attention score map to obtain a channel attention multi-scale feature map includes: Each scale feature of the second feature map is multiplied by an element of the corresponding channel attention score map to generate a corresponding channel attention multi-scale feature map, where the shape of the channel attention multi-scale feature map is the same as that of the first feature map.

8. A multi-scale attention feature extraction device, characterized in that: The device comprises: A first feature map acquisition module, used to acquire a first feature map of a target image; A second feature map acquisition module is used to input the first feature map into a serial convolution layer for convolution processing to obtain a second feature map; the second feature map includes several scale features; A channel feature vector acquisition module, used for performing a global pooling operation on each scale feature of the second feature map to obtain a channel feature vector; An attention score map acquisition module is used to convolve the channel feature vector to obtain a convolution feature vector, calculate the attention score of the convolution feature vector, and obtain a channel attention score map; A fusion module is used to fuse the second feature map with the channel attention score map to obtain a channel attention multi-scale feature map.

9. An electronic device, characterized in that: include: a memory for storing one or more computer programs; A processor, when the one or more computer programs are executed by the processor, implements the multi-scale attention feature extraction method as described in any one of claims 1-7.

10. A computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a processor to implement the multi-scale attention feature extraction method as described in any one of claims 1-7 when executed.