Multi-scale target detection method and device, computer equipment, readable storage medium and program product
Through the combination of graph neural network and self-attention mechanism, graph structure is constructed and local and global features are integrated, which solves the shortcomings of traditional multi-scale object detection methods in scale changes and small object detection, and achieves higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202510408364.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional multi-scale object detection methods are poorly accurate when processing images with significant scale variations, especially in small object detection and scale variation adaptability, and improper feature fusion strategies may lead to information loss or redundancy.
The graph neural network and self-attention mechanism are used to construct graph structures, combining the multi-head self-attention mechanism and feature pyramid network, local and global feature information are extracted and integrated to achieve detection of targets at different scales.
It improves the accuracy and robustness of multi-scale object detection, can effectively deal with small and large targets in complex scenarios, and enhances the model's ability to identify and adapt to different targets.
Smart Images

Figure CN120339768A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information technology, and in particular, to a multi-scale object detection method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] Traditional object detection methods usually operate at a fixed scale, which may not perform well when dealing with images with significant scale variations. Multi-scale object detection techniques introduce multi-scale or multi-resolution processing strategies to more comprehensively capture different scales of objects. However, currently, the accuracy of multi-scale object detection is poor. Summary of the Invention
[0003] Based on this, it is necessary to provide a multi-scale object detection method, apparatus, computer device, computer-readable storage medium, and computer program product for the above technical problems to improve the accuracy of multi-scale object detection.
[0004] In a first aspect, the present application provides a multi-scale object detection method, including:
[0005] Extracting feature map information of an image to be detected; the image to be detected is an image including at least two objects of different scales;
[0006] Obtaining a corresponding graph structure based on the feature map information; each node in the graph structure corresponds one-to-one with each position of the feature map information;
[0007] Inputting the graph structure data into a preset first model, and obtaining local feature information of objects of different scales in the image to be detected based on the graph neural network and self-attention mechanism of the first model;
[0008] Obtaining global feature information in the feature map information based on a second model with a multi-head self-attention mechanism;
[0009] Fusing the local feature information and the global feature information to obtain first target feature information;
[0010] Performing object detection on objects of different scales in the image to be detected based on the first target feature information.
[0011] In one embodiment, inputting the graph structure data into a preset first model, and obtaining local feature information of objects of different scales in the image to be detected based on the graph neural network and self-attention mechanism of the first model includes:
[0012] Input the feature map information into a preset first model to determine the node features of each node in the graph structure as the local feature information of targets at different scales in the to-be-detected image based on the attention weights between nodes and their neighbor nodes in the graph structure, the multi-head self-attention mechanism, and the stacking of multiple neural network layers in the first model.
[0013] In one embodiment, the second model is a Transformer network; the foregoing method further includes: flattening the two-dimensional feature map in the feature map information into a one-dimensional sequence and adding positional encoding to retain the spatial position information to obtain the one-dimensional information corresponding to the feature map information.
[0014] Based on the second model with the multi-head self-attention mechanism, obtain the global feature information in the feature map information, including: determining the global feature information in the feature map information based on the second model with the multi-head self-attention mechanism and the one-dimensional information.
[0015] In one embodiment, extracting the feature map information of the to-be-detected image includes:
[0016] Based on a preset third model, perform feature extraction on the to-be-detected image to determine the feature map information of the to-be-detected image; wherein, the third model includes multiple convolutional layers and multiple pooling layers, and each convolutional layer and each pooling layer are stacked layer by layer; the feature map information includes the low-level spatial feature information and the high-level semantic feature information corresponding to the to-be-detected image.
[0017] In one embodiment, fusing the local feature information and the global feature information to obtain the first target feature information, including:
[0018] Concatenate the local feature information and the global feature information to obtain the concatenated feature information.
[0019] Obtain the detection weights for targets at each scale.
[0020] Based on the detection weights and the concatenated feature information, determine the initial target feature information corresponding to each target at each scale; and based on the initial target feature information corresponding to each target at each scale, obtain the first target feature information.
[0021] In one embodiment, based on the first target feature information, perform target detection on targets at different scales in the to-be-detected image, including:
[0022] Based on a preset fourth model, combine the high-level semantic feature information and the low-level spatial feature information in the first target feature information to obtain the second target feature information.
[0023] Based on the second target feature information, perform target detection on targets at different scales in the to-be-detected image.
[0024] Second aspect, the present application also provides a multi-scale object detection device, including:
[0025] A basic feature module for extracting feature map information of the image to be detected; the image to be detected is an image including at least two objects of different scales; a corresponding graph structure is obtained based on the feature map information; each node in the graph structure corresponds one-to-one to each position of the feature map information;
[0026] A local feature module for inputting the graph structure data into a preset first model, and obtaining local feature information of objects of different scales in the image to be detected based on the graph neural network and self-attention mechanism of the first model;
[0027] A global feature module for obtaining global feature information in the feature map information based on a second model with a multi-head self-attention mechanism;
[0028] An object feature module for fusing the local feature information and the global feature information to obtain first object feature information;
[0029] A detection module for performing object detection on objects of different scales in the image to be detected based on the first object feature information.
[0030] Third aspect, the present application also provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in the first aspect above are implemented.
[0031] Fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in the first aspect above are implemented.
[0032] Fifth aspect, the present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect above are implemented.
[0033] The above multi-scale object detection method, device, computer device, computer-readable storage medium, and computer program product extract local feature information of a to-be-detected image through a first model and extract global feature information of the to-be-detected image through a second model to obtain first target feature information, thereby realizing the detection of objects of different scales in the to-be-detected image. Since the first target feature information is obtained by fusing local feature information and global feature information, multi-scale object detection can simultaneously focus on the local and global parts of the to-be-detected image, thereby improving the accuracy of detecting objects of different scales in the to-be-detected image. In addition, by obtaining a corresponding graph structure based on feature map information, more accurate local feature information can be obtained according to the graph structure, the graph neural network of the first model, and the self-attention mechanism; through the second model with a multi-head self-attention mechanism, more accurate global feature information can be obtained, which helps to obtain more accurate first target feature information and further improve the accuracy of detecting objects of different scales in the to-be-detected image. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0035] Figure 1 It is an application environment diagram of the multi-scale object detection method in an embodiment;
[0036] Figure 2 It is a flowchart of the multi-scale object detection method in an embodiment;
[0037] Figure 3 It is another flowchart of the multi-scale object detection method in an embodiment;
[0038] Figure 4 It is still another flowchart of the multi-scale object detection method in an embodiment;
[0039] Figure 5 It is a structural block diagram of the multi-scale object detection method device in an embodiment;
[0040] Figure 6 It is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] To make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application. As used herein, "multiple", "many", etc., without special instructions, the corresponding numerical range can be greater than or equal to 2.
[0042] The following first explains the terms involved in the technical solution of this application as follows:
[0043] Graph Neural Network (GNN): It is a class of neural network models for processing graph-structured data. Graph data consists of nodes (vertices) and edges. This model learns the feature representations of nodes or graphs through information transmission between nodes and their neighbors. Among them, the core idea of Graph Attention Network (GAT) is to assign different weights to the neighbors of each node and improve the expressive power of the model by learning the importance of each node's neighbors.
[0044] Self-Attention Mechanism: It is a computational mechanism for neural networks that enables the model to dynamically capture the correlations between different positions in an input sequence when processing the sequence. This mechanism can assign different weights according to the importance of each element in the input sequence, enabling the model to more flexibly focus on important information.
[0045] Multi-scale object detection: Aims to identify and locate objects of different scales in an image. Since the sizes of objects may vary significantly in an image, multi-scale object detection can effectively capture objects of different sizes by processing the image at multiple scales or resolutions.
[0046] Transformer, a neural network architecture based on the self-attention mechanism.
[0047] Convolutional Neural Networks (CNN): A class of neural networks that contain convolutional computations.
[0048] FPN (Feature Pyramid Networks), Feature Pyramid Network, is a neural network structure in the field of computer vision. It improves the detection of multi-scale objects by the model by using feature maps of different scales.
[0049] Bounding Box Regression: A technique used to predict the position of an object in an object detection task.
[0050] Upsampling, a technique commonly used in fields such as digital signal processing, image processing, and computer vision, is used to increase the number of data points, such as increasing the image resolution.
[0051] After the above noun explanations, the technical solutions provided by this application will be described below:
[0052] Multi-scale object detection is a key research direction in computer vision, aiming to improve the model's detection ability for various scale objects in images. Traditional object detection methods usually operate at a fixed scale, which may not perform well when dealing with images with significant scale variations. Multi-scale object detection techniques introduce multi-scale or multi-resolution processing strategies to more comprehensively capture different scales of objects, thereby improving detection performance. With the development of object detection technology and the introduction of a large number of deep learning models, the accuracy and efficiency of object detection have been improved. To further improve the model's performance in dealing with objects of different scales, more effective multi-scale detection techniques have emerged in recent years, including improving network structures, optimizing training strategies, and introducing new model architectures. The continuous progress of these techniques has made multi-scale object detection more efficient and accurate in practical applications and is widely used in fields such as autonomous driving, security monitoring, and medical image analysis. Although significant progress has been made in multi-scale object detection, there are still deficiencies, and the multi-scale object detection effect needs to be further improved:
[0053] First, the problem of small object detection: Small objects are difficult to be effectively recognized by detection algorithms due to their small size and limited features. This is particularly prominent in multi-scale object detection because small objects may only occupy a few pixels on the feature map, resulting in difficult feature extraction.
[0054] Second, adaptability to scale changes: Object detection algorithms need to be able to adapt to the scale changes of objects appearing in images, which requires the model to have good scale invariance. In practical applications, the scale of objects may vary significantly due to distance, perspective, or the size difference of the object itself.
[0055] Third, feature fusion strategy: How to effectively fuse features at different scales to obtain accurate object detection results is a challenge. Improper feature fusion may lead to information loss or redundancy.
[0056] Therefore, the related technologies of multi-scale object detection still need to be improved. The related technologies of multi-scale object detection help to improve the model's ability to recognize objects of different sizes, adapt to the diversity of object sizes in real-world scenarios, and thus enhance the accuracy, robustness, and generalization ability of the detection algorithm. Based on the above analysis, this application provides a multi-scale object detection method. This method is based on graph neural networks and self-attention mechanisms, and improves the model's detection ability for objects of different sizes by fusing local and global features while maintaining computational efficiency. The following will be described by way of examples:
[0057] Regarding the application scenarios of the technical solution of this application, it can include various scenarios that require multi-scale object detection. For example, in autonomous driving: in the field of autonomous driving, vehicles need to detect and recognize various objects in the surrounding environment in real time, including pedestrians, other vehicles, traffic signs, etc. Multi-scale object detection can ensure that objects at different distances and of different sizes can be accurately recognized, improving the safety and reliability of the autonomous driving system. Another example is intelligent monitoring: in the monitoring system, multi-scale object detection technology can be used to detect and track targets of different sizes, such as crowds, vehicles, etc., which is crucial for public safety and traffic management. Another example is during the process of taking pictures, recording videos, etc., to recognize multi-scale objects and perform intelligent tracking, focusing, etc.
[0058] In some embodiments, the multi-scale object detection method provided by the embodiments of this application can be applied to, for example, Figure 1 the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed in the cloud or on other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smartphones, tablets, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable device can be a head-mounted device, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0059] In some embodiments, the multi-scale object detection method provided by the embodiments of this application can also be applied to an integrated device for multi-scale object detection. For example, the integrated device can integrate functions such as image acquisition and multi-scale object detection, and does not need to be implemented by interacting with the server and the terminal.
[0060] In an exemplary embodiment, as Figure 2 shown, a multi-scale object detection method is provided. Taking the method applied to the Figure 1 server as an example for illustration, it includes the following steps S201 to S206. Among them:
[0061] Step S201: The server extracts the feature map information of the image to be detected; the image to be detected is an image including at least two objects of different scales.
[0062] Among them, the scale can be used to characterize the size of the object in the image to be detected. For example, the scale can characterize the area, pixel region, etc. occupied by the object in the image to be detected.
[0063] Among them, the image to be detected can be an image representing the existence of multi-scale object detection, and the image can include one or more objects to be detected. Exemplarily, the image to be detected can be a road surface image during vehicle driving, and the objects in the image can include displaceable objects such as road vehicles and pedestrians, and can also include immovable objects such as warning signs and roadblocks.
[0064] In some embodiments, the image to be detected can be a single image or multiple images. For example, three images obtained by an image acquisition device by taking pictures every 1 second and continuously for 3 seconds.
[0065] In some embodiments, the server can obtain the image to be detected from terminal devices such as surveillance cameras and dash cams.
[0066] In some embodiments, at least two objects of different scales can be understood as at least two objects and the scales of at least two objects are different, and can also be understood as at least one object and the object has at least two scales. For example, in the intelligent driving scenario, among multiple road surface images continuously captured by a dash cam, taking oncoming vehicle A as one object, the scale of oncoming vehicle A gradually becomes larger from far to near. Therefore, at least oncoming vehicle A is taken as one object and the object has at least two scales.
[0067] In some embodiments, the server can classify the objects in the image to be detected into: small objects, medium objects, and large objects according to the different scale sizes of the objects to be detected, and perform detection based on objects of different categories. For example, for small objects, specific algorithms or artificial intelligence models, etc. can be used to perform feature extraction and detection specifically, so as to avoid being ignored due to the too small scale of small objects during multi-scale object detection.
[0068] In some embodiments, the same image to be measured may be an image including targets of at least two different scales. The sizes of the targets included in two different images to be measured may be the same, partially the same, or different. For example, the image to be measured includes Image A and Image B. Image A includes Target a, Target b, and Target c, and Image B includes Target b and Target c. Exemplarily, the corresponding application scenario may be an intelligent driving scenario. Images A and B may be driving record images taken 1 second apart. Target a may be Vehicle a, Target b may be Pedestrian b, and Target c may be Pedestrian c. After 1 second, Vehicle a drives out of the field of view corresponding to the driving record image, while Pedestrians b and c are still within the field of view corresponding to the driving record image.
[0069] Step S202: The server obtains a corresponding graph structure based on the feature map information; each node in the graph structure corresponds one-to-one to each position of the feature map information.
[0070] In some embodiments, the position in the feature map may be the specific position coordinates in the feature map matrix.
[0071] In some embodiments, the server may define the adjacency relationship between the nodes in the graph structure based on the spatial information. For example, a position in a feature map may correspond to a node in a graph structure.
[0072] Exemplarily, the feature map information may be a feature matrix.
[0073] Step S203: The server inputs the graph structure data into a preset first model, and obtains the local feature information of the targets of different scales in the image to be measured based on the graph neural network and the self-attention mechanism of the first model.
[0074] In some embodiments, the first model may be an artificial intelligence model with high accuracy in small target detection and / or processing complex backgrounds.
[0075] In some embodiments, the first model may include one or more local connection layers. Different from the fully connected layer, the local connection layer (such as a convolutional layer) only applies weights to the local area of the input data, which helps to retain and emphasize the local feature information.
[0076] Step S204: The server obtains the global feature information in the feature map information based on the second model of the multi-head self-attention mechanism.
[0077] In some embodiments, the second model may be an artificial intelligence model with high accuracy in capturing the information of the entire image.
[0078] In some embodiments, the second model may adopt a multi-head self-attention mechanism, which enables the model to dynamically consider the relationship between any two positions when processing an image, regardless of their distance. This mechanism is particularly suitable for capturing long-range dependencies, thereby effectively utilizing the global information of the image.
[0079] Step S205: The server fuses the local feature information and the global feature information to obtain first target feature information.
[0080] In some embodiments, the server may implement the fusion of the local feature information and the global feature information through a preset fusion model. The fusion model may be constructed and trained based on a specific algorithm.
[0081] In some embodiments, the first model and the second model may be integrated into the same integrated model, and the integrated model may fuse the local feature information and the global feature information to obtain first target feature information.
[0082] Step S206: The server performs object detection on objects of different scales in the image to be detected based on the first target feature information.
[0083] In some embodiments, the server may perform object detection on objects of different scales in the image to be detected based on the first target feature information to obtain multiple detection results. For example, multiple object detection result boxes, and filter the multiple detection results. For example, by removing duplicates to ensure that the number of finally output object boxes is appropriate and non-repetitive.
[0084] In some embodiments, the server may, from the perspective of the low layer (high resolution but weak semantic information) and high layer (low resolution but strong semantic information) of the features, perform object detection on objects of different scales in the image to be detected based on the first target feature information.
[0085] The local feature information of the image to be detected is extracted by the first model, and the global feature information of the image to be detected is extracted by the second model to obtain first target feature information, thereby realizing the detection of objects of different scales in the image to be detected. Since the first target feature information is obtained by fusing the local feature information and the global feature information, multi-scale object detection can simultaneously focus on the local and global parts of the image to be detected, thereby improving the accuracy of detecting objects of different scales in the image to be detected. In addition, by obtaining the corresponding graph structure based on the feature map information, and according to the graph structure, the graph neural network and the self-attention mechanism of the first model, more accurate local feature information can be obtained; through the second model with the multi-head self-attention mechanism, more accurate global feature information can be obtained, which helps to obtain more accurate first target feature information and further improve the accuracy of detecting objects of different scales in the image to be detected.
[0086] In one embodiment, the foregoing "inputting the graph structure data into a preset first model, and obtaining the local feature information of the targets at different scales in the image to be detected based on the graph neural network and the self-attention mechanism of the first model" may include: the server inputs the feature map information into the preset first model, and based on the attention weights between the nodes and their neighbor nodes in the graph structure, the multi-head self-attention mechanism, and the superposition of multiple neural network layers in the first model, determines the node features of each node in the graph structure as the local feature information of the targets at different scales in the image to be detected.
[0087] Among them, the neighbor node may refer to a node directly connected to the current node in the graph structure.
[0088] In some embodiments, the first model may adopt a multi-head attention mechanism. Each attention head calculates the attention weights between each node and its neighbor nodes in the graph structure respectively, and then fuses the calculation results of the attention weights of each attention head, and based on the superposition of multiple neural network layers, obtains the local feature information.
[0089] In some embodiments, the first model may adopt a graph attention network, thereby introducing a self-attention mechanism in the process of aggregating neighbor node information to allow the model to assign different weights to different neighbors of each node and dynamically adjust the weights of each neighbor node.
[0090] By using the graph structure and the self-attention mechanism, the local nodes in the feature map are weighted, thereby realizing the dynamic update of the node features; through multi-layer superposition, the local features of the image to be detected can be effectively captured, such as the details of complex regions in the image to be detected, so that the local features are suitable for the processing of small targets in multi-scale object detection.
[0091] In one embodiment, the second model is a Transformer network; the foregoing method further includes: the server flattens the two-dimensional feature map in the feature map information into a one-dimensional sequence, and adds position encoding to retain the spatial position information, obtaining the one-dimensional information corresponding to the feature map information; based on the second model with a multi-head self-attention mechanism, obtains the global feature information in the feature map information, including: based on the second model with a multi-head self-attention mechanism and the one-dimensional information, determines the global feature information in the feature map information.
[0092] Among them, the Transformer network may be a neural network based on the Transformer model architecture.
[0093] In some embodiments, since the Transformer was originally designed for sequence data, before processing the feature map information of the image to be measured using the transformer network, the two-dimensional feature map in the feature map information can be flattened into a one-dimensional sequence, and position encoding can be added to retain the spatial position information, so as to obtain one-dimensional information, and then the one-dimensional information can be processed using the transformer network.
[0094] By means of the multi-head self-attention mechanism, the correlation between features is calculated at different positions of the sequence in the one-dimensional information, so as to capture long-distance dependencies and global context, thereby obtaining more accurate global feature information, making the global feature information better suitable for the processing of large targets in multi-scale object detection.
[0095] In one of the embodiments, the aforementioned "extracting the feature map information of the image to be measured" may include: the server performs feature extraction on the image to be measured based on a preset third model to determine the feature map information of the image to be measured; wherein, the third model includes multiple convolutional layers and multiple pooling layers, and the convolutional layers and the pooling layers are stacked layer by layer; the feature map information includes low-level spatial feature information and high-level semantic feature information corresponding to the image to be measured.
[0096] In some embodiments, the third model may be a Convolutional Neural Networks (CNN) model. The CNN can generate feature maps with different semantic levels through stacked convolutional and pooling operations layer by layer, so as to obtain the feature map information corresponding to the image to be measured that includes both low-level spatial feature information and high-level semantic feature information.
[0097] In some embodiments, the first model, the second model, and the third model may be integrated into the same model.
[0098] By gradually extracting features from the image to be measured through the third model, rich basic features, that is, feature map information, are provided for multi-scale object detection. The feature map information not only contains low-level spatial feature information, such as low-level fine-grained spatial information, but also combines high-level semantic feature information, which helps to more accurately detect objects of different scales.
[0099] In one of the embodiments, the aforementioned "fusing the local feature information and the global feature information to obtain the first target feature information" may include: the server splices the local feature information and the global feature information to obtain spliced feature information; obtains the detection weights for the targets of each scale; determines the initial target feature information corresponding to each target of each scale based on the detection weights and the spliced feature information; and obtains the first target feature information based on the initial target feature information corresponding to each target of each scale.
[0100] Among them, the detection weight can be a weight for a target at a certain scale during the multi-scale object detection of the image to be detected. Exemplarily, the detection weights corresponding to different scale targets can be the same or different.
[0101] In some embodiments, the detection weight can represent the degree of attention or importance to different scale targets during the multi-scale object detection of the image to be detected. Exemplarily, by assigning a larger detection weight to small targets, the multi-scale object detection can pay more attention to the detection of small targets.
[0102] In some embodiments, the server can obtain the multi-scale object detection requirements for the image to be detected. For example, in the intelligent driving scenario, the multi-scale object detection requirements may include detecting pedestrians, especially children. The server determines the detection weights for the targets at each scale according to the multi-scale object detection requirements. For example, the server can divide pedestrians into adults and children. Adults belong to large targets and children belong to small targets. The detection weight of small targets is greater than that of large targets. By determining the detection weights for the targets at each scale according to the multi-scale object detection requirements, more accurate and more practical multi-scale object detection can be achieved.
[0103] In some embodiments, the server can fuse the local feature information and the global feature information based on methods such as weighted average and layer-by-layer splicing.
[0104] By splicing the local feature information and the global feature information and combining the detection weights, the initial target feature information corresponding to each scale of the target is determined, which realizes the balanced processing between different scale targets, solves the detection deviation problem between small targets and large targets, and thus obtains more accurate first target information.
[0105] In one of the embodiments, the aforementioned "performing object detection on the targets of different scales in the image to be detected based on the first target feature information", as Figure 3 shown, may include steps S301 to S302:
[0106] Step S301: Based on a preset fourth model, combine the high-level semantic feature information and the low-level spatial feature information in the first target feature information to obtain second target feature information.
[0107] Among them, the high-level semantic feature information and the low-level spatial feature information can be information of different levels in the first target feature information extracted from the image to be measured. The high-level semantic feature information can be information of feature representations obtained through deeper layers of a deep neural network (such as a convolutional neural network CNN), which can contain highly abstract information about the image content, such as the category of objects (such as people, cars, animals, etc.), the type of scene (such as indoor, outdoor, etc.); the low-level spatial feature information can be feature information extracted from the shallow layer or the first few layers of the network, which pays more attention to the basic components of the image, such as local detail information such as edges, color patches, textures, etc.
[0108] In some embodiments, the first model, the second model, the third model, and the fourth model can all be integrated in the same model.
[0109] In some embodiments, the fourth model can adopt a feature pyramid network.
[0110] Step S302: Based on the second target feature information, perform object detection on objects of different scales in the image to be measured.
[0111] By combining the high-level semantic feature information with the low-level spatial feature information, the detection ability of the model for multi-scale objects is enhanced. Specifically, the high-level semantic feature information can help identify what the object is, while the low-level spatial feature information helps to accurately locate the position and its boundary of the object. This combination enables the model to not only understand the overall situation of the image but also capture the necessary details, thereby improving the overall performance.
[0112] In an exemplary embodiment, the present application proposes a multi-scale object detection method. As Figure 4 shown, the overall process of the method is given. The design idea of combining GAT and Transformer based on graph neural network and self-attention mechanism is to improve the detection accuracy by integrating local and global feature modeling. First, GAT adaptively assigns neighborhood weights on the graph structure through the self-attention mechanism, accurately capturing the local relationships and detail features of objects of different scales. Then, Transformer models the long-range dependencies and global context information in the feature map through the global self-attention mechanism, helping the model understand the complex relationships between cross-scale objects and the background. This method gives play to the advantages of GAT in local modeling and at the same time utilizes the global information capture ability of Transformer to achieve more comprehensive and accurate detection of multi-scale objects. The following will further explain the above method from the perspective of modules in combination with the Figure 4 process shown.
[0113] I. Regarding the feature extraction module:
[0114] Use a convolutional neural network (CNN) to extract the basic features of the input image and generate multi-layer feature maps. These feature maps have different scales and resolutions, providing inputs for subsequent multi-scale detection. The content of this module corresponds to the relevant parts such as "extracting the feature map information of the image to be detected" mentioned above. Specifically:
[0115] (I) Convolution operation
[0116] Convolution operation is a key step in CNN for extracting local features. In the convolutional layer, each filter (convolution kernel) slides over the input feature map to calculate the weighted sum of the local region and generate the output feature map. The formula for the convolution operation is:
[0117] y[i, j, k] = ∑ m ∑ n x[i + m, j + n, c] · w[m, n, c, k] + b k
[0118] where x[i, j, c] represents the pixel value of the c-th channel at the position (i, j) in the input feature map. w[m, n, c, k] represents the weight of the c-th input channel of the convolution kernel k at the position (m, n). y[i, j, k] is the value of the k-th channel at the position (i, j) in the output feature map. b k is the bias term of the k-th channel.
[0119] (II) Multi-level design of feature maps
[0120] CNN generates feature maps with different semantic levels by stacking convolution and pooling operations layer by layer. The generation of feature maps can be represented by the following recursive formula:
[0121] x l+1 = f(x l ; W l );
[0122] where x l represents the input feature map of the l-th layer, and x l+1 represents the input feature map of the (l + 1)-th layer; W l is the weight matrix of the l-th layer.
[0123] (III) Pooling operation
[0124] Pooling operation is used to downsample the feature map, reducing the spatial dimension while retaining important features. The formula for the max-pooling operation is:
[0125] y[i, j, k] = max m,n x[i + m, j + n, k];
[0126] x[i, j, k] is the pixel value of the k-th channel at the position (i, j) in the input feature map; y[i, j, k] is the value of the k-th channel at the position (i, j) in the pooled feature map.
[0127] (IV) Multi-scale feature output
[0128] Finally, feature maps {C1, C2, C3} with multiple resolutions are generated through a convolutional network. Among them, C1, C2, and C3 represent feature maps at different levels, representing different scale information, that is, the features corresponding to the feature maps are scale features of different scale sizes.
[0129] The feature extraction module provides rich basic features for multi-scale object detection by extracting features layer by layer. These features not only contain low-level fine-grained spatial information but also combine high-level semantic information, which helps to accurately detect objects of different scales.
[0130] II. GAT module:
[0131] The construction and feature extraction steps of GAT are to enhance the model's ability to capture local and global features. The content of this module corresponds to the relevant parts such as "obtaining the corresponding graph structure based on the feature map information" and "inputting the graph structure data into a preset first model, and obtaining the local feature information of objects of different scales in the image to be measured based on the graph neural network and self-attention mechanism of the first model".
[0132] (I) Constructing the graph structure
[0133] Each position in the feature map extracted by the convolutional neural network is regarded as a node in the graph. The construction of the graph structure usually defines the adjacency relationship between nodes based on spatial information or channel information. Each position (i, j) of the feature map is regarded as a node, and the feature map itself is a matrix of size H×W with C channels.
[0134] (II) Message passing and attention calculation
[0135] After constructing the graph structure, GAT uses the self-attention mechanism to perform message passing and feature update in the neighborhood. For each node i, GAT calculates the attention weight between it and each neighbor node j, and the calculation formula of the weight is:
[0136] e ij =LeakyRelU(a T [Wh i ||Wh j );
[0137] Where, h i and h jRespectively represent the feature vectors of node i and neighbor node j; || represents the vector concatenation operation; W: This is a learnable weight matrix; a T This is the transpose of the attention vector a, used for dot product operation with the concatenated feature vector; LeakyReLU: A non-linear activation function, used to introduce non-linearity; attention weight e ij .
[0138] The normalized attention weight is expressed as:
[0139]
[0140] Use the attention weight to perform weighted summation on the features of neighbor nodes to update the features of node i, expressed as:
[0141] h′ i = σ(∑ j∈N(i) a ij Wh j ).
[0142] (III) Multi-Head Attention Mechanism
[0143] To enhance the learning ability of the model, GAT uses a multi-head attention mechanism. Concatenate or average the results of multiple attention heads, expressed as:
[0144]
[0145] Among them, K represents the number of multi-head attention; is the attention weight of the k-th attention head. The updated feature vector of node i. σ: Activation function, exemplarily, LeakyReLU or other non-linear activation functions can be used. W k The weight matrix in the k-th attention head. h j The original feature vector of node j. N(i): The neighborhood set of node i, that is, the set of all nodes directly connected to node i.
[0146] (IV) Multi-Layer Stacking
[0147] To further expand the receptive field of features, multiple GAT layers can be stacked. The output of each GAT layer is used as the input of the next layer, gradually propagating the features to more distant neighbor nodes, thereby enhancing the ability to capture local features.
[0148] H l+1 = GA T (H l , A);
[0149] Among them, H l+1: The updated node feature matrix, where each row represents the feature vector of a node at the (l + 1)-th layer. H l : The node feature matrix of the current layer (the l-th layer), where each row represents the feature vector of a node at this layer. A: The adjacency matrix, representing the connection relationship between nodes in the graph. The adjacency matrix A is a binary matrix. If there is an edge between node i and node j, then A ij = 1; otherwise, A ij = 0.
[0150] The GAT module constructs a graph structure and uses the self-attention mechanism to weight local nodes in the feature map and dynamically update node features. Through multi-layer stacking, GAT can effectively capture local features and details of complex regions, making it suitable for processing small targets in multi-scale object detection.
[0151] III. Transformer Module:
[0152] The Transformer module is used for global feature modeling, capable of capturing long-range dependencies and global context information, and is particularly suitable for the detection of large-scale targets. The content of this module corresponds to the relevant parts such as "the second model based on the multi-head self-attention mechanism, obtaining the global feature information in the feature map information" mentioned above.
[0153] (I) Input Embedding and Position Encoding
[0154] Since Transformer was initially designed for sequence data, when processing image data, the two-dimensional feature map needs to be flattened into a one-dimensional sequence, and position encoding is added to retain spatial position information. Thus, one-dimensional information is obtained.
[0155] Flatten the feature map F ∈ R H×W×C into a sequence X ∈ R N×C , where N = H × W is the sequence length.
[0156] Generate position encoding for each position using sine and cosine functions:
[0157]
[0158] where PE(pos, 2i) represents the position embedding value of the 2i-th dimension at position pos; pos, the position index in the sequence; d model is the hidden layer dimension of the model.
[0159] (II) Multi-Head Self-Attention Mechanism
[0160] The multi-head self-attention mechanism is the core of Transformer, used to capture the correlation between different positions in the sequence.
[0161] The query Q, key K, and value V matrices are obtained through linear transformation:
[0162] Q = X input W Q , K = X input W K , V = X input W V ;
[0163] Calculate the attention weight matrix A and normalize it:
[0164] Use the attention weight matrix A to perform weighted summation on the value matrix V to obtain the output Z = AV;
[0165] Compute the self-attention of multiple heads in parallel and then concatenate the results:
[0166] Mul(Q, K, V) = Concat(h1, h2,..h n )W o .
[0167] (III) Residual connection and normalization
[0168] Add the output of the self-attention layer to the input: Y = X input + Z.
[0169] Normalize the output of the residual connection: Y = LN(Y).
[0170] (IV) Feed-forward network
[0171] Each Transformer layer also includes a feed-forward network for further processing of features.
[0172] Through two fully connected layers and the ReLU activation function:
[0173] F(Y) = ReLU(W1Y + b1)W2 + b2.
[0174] Use the residual connection and LayerNormalization again:
[0175] Y = LN(Y + F(Y)).
[0176] (V) Stacking multiple layers of Transformer
[0177] The output of each layer is used as the input of the next layer to gradually build deeper feature representations:
[0178]
[0179] Calculate the correlation between features at different positions in the sequence through the multi-head self-attention mechanism to capture long-distance dependencies and global context. Combining residual connections and normalization mechanisms, the Transformer can not only be trained stably but also effectively model global features, making it particularly suitable for processing large-scale targets.
[0180] IV. Multi-scale feature fusion:
[0181] The purpose is to combine local details and global context information to ensure that the model can evenly process targets of different scales.
[0182] (I) Feature fusion
[0183] Concatenate the features of the GAT and Transformer modules in the feature dimension:
[0184] F fused = concat(F GAT , F Transformer ),
[0185] where F fused represents the concatenated features, i.e., the concatenation feature; F GAT represents the features obtained from the aforementioned GAT module; F Tranformer represents the features obtained from the aforementioned Transformer module; concat() is a function for concatenation.
[0186] (II) Scale balance
[0187] Process features of different scales through a specific fusion module.
[0188] Use a feature weighting mechanism to adjust the weighting value according to the detection requirements of targets at different scales:
[0189] F fused = softmax(W F-scale ) * F input , where,
[0190] F input is the aforementioned concatenated feature, softmax(W F-scale ) represents the weights corresponding to targets at different scales, i.e., the corresponding detection weights, and F fused represents the features obtained after weighted adjustment based on the aforementioned concatenated features, i.e., the corresponding initial target features.
[0191] In some embodiments, targets at each scale correspond to their respective weighted features F fused , so multiple F fused can be obtained.
[0192] In some embodiments, a normalization method can be used to ensure the balance of features at different scales during fusion. Exemplarily, the features at different scales are normalized:
[0193]
[0194] where F represents the weighted features of different-scale targets; ||F|| is the norm of F, for example, the Euclidean norm (i.e., the L2 norm); F norm is the feature of different-scale targets after normalization. Thus, the first target feature information is obtained based on the initial target feature information corresponding to each scale's target.
[0195] By performing weighted averaging, layer-by-layer concatenation, or other fusion strategies on the features of the GAT and Transformer modules, local details and global context information are combined. Specific fusion modules and scale balance strategies ensure that the model can evenly process different-scale targets, solving the detection bias problem between small and large targets.
[0196] V. Multi-scale prediction:
[0197] To process targets of different scales, object detection is performed on feature maps at different levels.
[0198] (I) Multi-scale prediction
[0199] The FPN (Feature Pyramid Network) provides rich feature representations for targets of different scales by combining high-level semantic information (i.e., high-level semantic feature information) and low-level spatial information (i.e., low-level spatial feature information). The high-level semantic features are combined with the low-level spatial features through sampling and fusion:
[0200] P l = Conv(Upsample(P l+1 ) + C l ),
[0201] where C l represents the spatial features of the l-th layer (low layer), P l+1 represents the semantic features of the (l + 1)-th layer (high layer), and P l represents the features of the l-th layer obtained by combining the high-level semantic features and the low-level spatial features.
[0202] In some embodiments, both the high-level semantic features and the low-level spatial features here can be the features obtained through the aforementioned step of "IV. Multi-scale feature fusion".
[0203] Object classification is performed for each scale at each feature map position, and the class probability is output:
[0204] Cl = softmax(W l P l + b l ),
[0205] where C l is the class probability, W l is the weight matrix, and b l is the bias vector;
[0206] Regress the bounding boxes at each feature map location to predict the location information of the target:
[0207] B l = W l P l + b l ,
[0208] where B l is the location information of the target.
[0209] (II) Loss Function
[0210] The classification loss measures the prediction:
[0211]
[0212] where L cls is the classification loss; y i is the true correct target classification. Exemplarily, y i=0 can represent incorrect, and y i=1 can represent correct; The predicted probability represents the prediction probability of the model for the i-th class, i.e., the class probability.
[0213] The bounding box bbox regression loss:
[0214]
[0215] where L bbox is the bounding box bbox regression loss; is the predicted bounding box parameter; b i is the true bounding box parameter. A specific loss function is used to balance the detection effects of targets at different scales:
[0216] L = λ1L1 + λ2L2,
[0217] where L is the total loss function; L1 and L2 are two different loss functions, which can represent the losses of the model on different tasks or different aspects of the same task; λ1 and λ2 are weight coefficients used to adjust the relative importance of L1 and L2 in the total loss L. Exemplarily, the total loss function is the weighted sum of the classification loss and the bounding box regression loss:
[0218] L = λ1L cls + λ2L bbox 。
[0219] Object detection is achieved through multi-scale prediction and loss function design. In multi-scale prediction, techniques such as FPN are combined to perform object classification and bounding box regression on feature maps at different levels to handle objects of different scales. Loss function design uses a multi-task loss function, combining classification loss and bounding box regression loss, to ensure that the model can accurately classify objects and precisely locate objects.
[0220] The above technical solutions have the following effects:
[0221] First, the combination of local and global features. GAT is good at capturing local neighbor relationships in the graph structure through the message passing mechanism. It can carefully model target regions and their local associations at different scales, and effectively process target features at different scales by adaptively assigning weights to neighbor nodes. Transformer, through the self-attention mechanism, can perform global modeling to capture long-distance dependencies and global context information. It can understand the feature relationships at different positions in the image over a larger range, thus making up for the deficiency of GAT's local processing only. GAT provides accurate modeling of the local neighborhood and is suitable for processing fine-grained features of objects, while Transformer supplements the ability to capture global information, which can help the model better understand multi-scale objects and background information in the whole image. GAT strengthens the interaction between local features through the attention mechanism and captures the complex relationships between adjacent regions in the image, while Transformer uses the self-attention mechanism to capture global context information and long-distance dependencies. This combination enables the model to simultaneously focus on local details and global structures, improving the detection ability for objects of different sizes. In complex scenes, objects may be difficult to detect due to occlusion, lighting changes, or background interference. The combination of GAT and Transformer improves the adaptability and detection robustness to these complex scenes by enhancing the model's understanding of local and global features.
[0222] Second, adaptive multi-scale feature fusion. GAT can weight different regions in multi-scale feature maps through the attention mechanism, dynamically adjusting the fusion method of target information at each scale. This feature fusion enables the model to more effectively distinguish targets from the background, especially suitable for processing complex scenarios with coexisting small and large targets. Transformer can establish global relationships between features at different scales and different levels through the multi-head self-attention mechanism, which helps to achieve deep fusion across feature layers at multiple scales and improve the model's perception ability of multi-scale targets. GAT can adaptively adjust the attention weights at the local scale, enhancing the local sensitivity of feature fusion; while Transformer can perform cross-scale information fusion at the global scale, achieving an efficient combination of detailed features and context information. In addition, through the Feature Pyramid Network (FPN) technology, by combining the output features of GAT and Transformer, the model can perform object detection on feature maps at different levels. This method not only improves the detection accuracy of small targets but also enhances the recognition ability of large targets.
[0223] Third, handling long-range dependencies and complex backgrounds. GAT can capture the complex dependency relationships between targets and their local contexts through the attention mechanism in the graph structure, especially suitable for processing image data with irregular and non-Euclidean structures. Transformer can span long pixel positions in the image and effectively capture the long-range feature relationships in the image through the global attention mechanism, enabling it to better handle targets in complex backgrounds. GAT enhances the expression of target features in the local area, making it better able to distinguish targets in complex backgrounds, while Transformer provides a powerful modeling ability for long-range dependencies, helping the model process global background information and reduce false detections.
[0224] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential either, but can be executed alternately or in rotation with at least a part of other steps or steps or stages in other steps.
[0225] Based on the same inventive concept, an embodiment of the present application further provides a multi-scale object detection device for implementing the multi-scale object detection method involved above. The implementation solution provided by this device to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the multi-scale object detection device provided below can refer to the limitations on the multi-scale object detection method in the above text, and will not be elaborated here.
[0226] In an exemplary embodiment, as Figure 5 shown, a multi-scale object detection device 500 is provided, including:
[0227] A basic feature module 501, configured to extract feature map information of an image to be detected; the image to be detected is an image including at least two objects of different scales; a corresponding graph structure is obtained based on the feature map information; each node in the graph structure corresponds one-to-one with each position of the feature map information;
[0228] A local feature module 502, configured to input the graph structure data into a preset first model, and obtain local feature information of objects of different scales in the image to be detected based on the graph neural network and self-attention mechanism of the first model;
[0229] A global feature module 503, configured to obtain global feature information in the feature map information based on a second model of the multi-head self-attention mechanism;
[0230] An object feature module 504, configured to fuse the local feature information and the global feature information to obtain first object feature information;
[0231] A detection module 505, configured to perform object detection on objects of different scales in the image to be detected based on the first object feature information.
[0232] In one of the embodiments, the local feature module 502 is further configured to input the graph structure data into a preset first model, and obtain local feature information of objects of different scales in the image to be detected based on the graph neural network and self-attention mechanism of the first model, including: inputting the feature map information into the preset first model, and determining the node features of each node in the graph structure as the local feature information of objects of different scales in the image to be detected based on the attention weights between the nodes and their neighbor nodes in the graph structure, the multi-head self-attention mechanism, and the superposition of multiple neural network layers in the first model.
[0233] In one embodiment, the second model is a transformer network; the global feature module 503 is further configured to flatten the two-dimensional feature map in the feature map information into a one-dimensional sequence, and add a positional encoding to retain the spatial position information, so as to obtain the one-dimensional information corresponding to the feature map information; the second model based on the multi-head self-attention mechanism is used to obtain the global feature information in the feature map information, including: determining the global feature information in the feature map information based on the second model based on the multi-head self-attention mechanism and the one-dimensional information.
[0234] In one embodiment, the basic feature module 501 is further configured to extract the feature map information of the image to be detected, including: performing feature extraction on the image to be detected based on a preset third model to determine the feature map information of the image to be detected; wherein, the third model includes a plurality of convolutional layers and a plurality of pooling layers, and the convolutional layers and the pooling layers are stacked layer by layer; the feature map information includes the low-level spatial feature information and the high-level semantic feature information corresponding to the image to be detected.
[0235] In one embodiment, the target feature module 504 is further configured to fuse the local feature information and the global feature information to obtain the first target feature information, including: splicing the local feature information and the global feature information to obtain the spliced feature information; obtaining the detection weights for the targets of each scale; determining the initial target feature information corresponding to each target of each scale based on the detection weights and the spliced feature information; and obtaining the first target feature information based on the initial target feature information corresponding to each target of each scale.
[0236] In one embodiment, the detection module 505 is further configured to perform object detection on objects of different scales in the image to be detected based on the first target feature information, including: combining the high-level semantic feature information and the low-level spatial feature information in the first target feature information based on a preset fourth model to obtain the second target feature information; performing object detection on objects of different scales in the image to be detected based on the second target feature information.
[0237] Each module in the above multi-scale object detection device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0238] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 6As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data required for the multi-scale object detection method, such as feature map information, first object feature information, etc. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a multi-scale object detection method.
[0239] Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0240] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0241] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0242] In an embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0243] Those of ordinary skill in the art can understand that all or part of the processes in the above-described method embodiments can be implemented by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-described method embodiments. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the various embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0244] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in this application.
[0245] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A multi-scale object detection method, characterized in that, The method includes: Extracting feature map information of the image to be detected; the image to be detected is an image including at least two targets of different scales; Obtaining a corresponding graph structure based on the feature map information; each node in the graph structure corresponds one-to-one to each position of the feature map information; Inputting the graph structure data into a preset first model, and obtaining local feature information of targets of different scales in the image to be detected based on the graph neural network and self-attention mechanism of the first model; Obtaining global feature information in the feature map information based on a second model with a multi-head self-attention mechanism; Fusing the local feature information and the global feature information to obtain first target feature information; Performing target detection on targets of different scales in the image to be detected based on the first target feature information.
2. The method according to claim 1, wherein The step of inputting the graph structure data into a preset first model and obtaining local feature information of targets of different scales in the image to be detected based on the graph neural network and self-attention mechanism of the first model includes: Inputting the feature map information into a preset first model, and determining node features of each node in the graph structure as local feature information of targets of different scales in the image to be detected based on the attention weights between nodes and their neighbor nodes in the graph structure, the multi-head self-attention mechanism, and the stacking of multiple neural network layers in the first model.
3. The method according to claim 1, wherein The second model is a Transformer network; The method further includes: Flattening the two-dimensional feature map in the feature map information into a one-dimensional sequence and adding positional encoding to retain spatial position information, obtaining one-dimensional information corresponding to the feature map information; The step of obtaining global feature information in the feature map information based on a second model with a multi-head self-attention mechanism includes: Determining global feature information in the feature map information based on a second model with a multi-head self-attention mechanism and the one-dimensional information.
4. The method according to claim 1, wherein The step of extracting feature map information of the image to be detected includes: Performing feature extraction on the image to be detected based on a preset third model to determine the feature map information of the image to be detected; where The third model includes multiple convolutional layers and multiple pooling layers, and each convolutional layer and each pooling layer are stacked layer by layer; the feature map information includes low-level spatial feature information and high-level semantic feature information corresponding to the image to be detected.
5. The method according to claim 1, characterized in that, The step of fusing the local feature information and the global feature information to obtain first target feature information includes: Concatenating the local feature information and the global feature information to obtain concatenated feature information; Obtaining detection weights for each scale of the target; Determining initial target feature information corresponding to each scale of the target based on the detection weights and the concatenated feature information; And obtaining the first target feature information based on the initial target feature information corresponding to each scale of the target.
6. The method according to any one of claims 1 to 5, characterized in that The step of performing target detection on targets of different scales in the image to be detected based on the first target feature information includes: Based on a preset fourth model, combining the high-level semantic feature information and the low-level spatial feature information in the first target feature information to obtain second target feature information; Based on the second target feature information, perform target detection on targets of different scales in the image to be measured.
7. A multi-scale object detection device, characterized in that, The device includes: A basic feature module, configured to extract feature map information of the image to be measured; the image to be measured is an image including at least two targets of different scales; obtain a corresponding graph structure based on the feature map information; each node in the graph structure corresponds one-to-one to each position of the feature map information; A local feature module, configured to input the graph structure data into a preset first model, and obtain local feature information of targets of different scales in the image to be measured based on the graph neural network and self-attention mechanism of the first model; A global feature module, configured to obtain global feature information in the feature map information based on a second model of the multi-head self-attention mechanism; A target feature module, configured to fuse the local feature information and the global feature information to obtain first target feature information; A detection module, configured to perform target detection on targets of different scales in the image to be measured based on the first target feature information.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.