A neural network training method and system for few-shot object detection
By combining deep meta-learning and deep separable convolutional structures, the problem of insufficient cross-scene adaptability of target detection models under conditions of few samples is solved, and efficient and accurate airport bird detection and security monitoring are achieved.
Patent Information
- Application Number
- CN202511660695.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2045-11-13
AI Technical Summary
Existing target detection methods are prone to overfitting under conditions of few samples and are difficult to adapt to cross-scene tasks, especially in airport bird detection, where changes in lighting and weather conditions lead to insufficient model generalization ability, and multimodal data lacks an effective feature fusion mechanism.
We employ a deep meta-learning framework and a deep separable convolutional structure. We train a feature extraction network using a cross-scene sample dataset, and use a meta-learner to perform inner and outer loop optimization on the support set and query set to generate initial weight parameters. We then combine multimodal data fusion and refined post-processing to construct a lightweight object detection network.
It improves the model's generalization ability and detection accuracy in complex environments, and enables rapid adaptation and efficient detection under conditions with few samples. It is suitable for cross-scenario applications such as airport bird detection and security monitoring.
Smart Images

Figure CN121119059B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning technology, and in particular to a neural network training method and system for few-shot object detection. Background Technology
[0002] Object detection is one of the core tasks in computer vision, aiming to locate and identify objects of interest in images or videos. In recent years, deep learning-based object detection methods have made significant progress, especially with the support of large-scale labeled datasets, with many models demonstrating excellent performance in specific scenarios. However, these methods heavily rely on large amounts of high-quality labeled data. In practical applications, such as airport bird detection and security monitoring, acquiring sufficient samples and performing detailed annotations is often costly and time-consuming in cross-scenario tasks.
[0003] In few-shot learning, traditional object detection models are prone to overfitting due to insufficient training data, leading to decreased generalization ability and difficulty in adapting to detection requirements in new scenarios or categories. Although existing research has attempted to alleviate the data scarcity problem through methods such as data augmentation and transfer learning, these methods are usually limited to a single scenario or category and cannot effectively cope with distribution differences and environmental changes across scenarios. For example, in airport bird detection, the lighting conditions, weather conditions, and background environment vary significantly between different airports, making it difficult to directly transfer models trained in a single scenario. In addition, existing methods often lack effective feature fusion mechanisms when processing multimodal data (such as visible light and infrared images), further limiting the detection performance of models in complex environments. Therefore, there is an urgent need for an object detection method that can quickly adapt to cross-scenario tasks with a small number of labeled samples to improve the model's generalization ability in real and complex environments. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, the embodiments of this application provide a neural network training method for few-sample target detection to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, this application provides a neural network training method for few-shot target detection, comprising:
[0006] A cross-scene sample dataset is obtained, which includes sample images under various lighting conditions and weather conditions to simulate the environmental conditions of the target airport.
[0007] A feature extraction network based on deep meta-learning is constructed. The feature extraction network is trained on the cross-scene sample dataset using a meta-learner to generate the initial weight parameters of the target detection network. The training process includes: dynamically extracting sample subsets from the cross-scene sample dataset according to environmental attribute classification; constructing a meta-learning task simulating specific airport environmental conditions based on each sample subset; each meta-learning task includes a support set and a query set; performing an inner loop optimization process on the support set samples and an outer loop optimization process on the query set samples; and updating the network parameters using a gradient descent algorithm.
[0008] The target detection network is initialized using the initial weight parameters. The target detection network employs a depthwise separable convolutional structure, which includes performing depthwise convolution operations and pointwise convolution operations on the input feature map.
[0009] The monitoring data of the target airport is input into the initialized target detection network for forward inference calculation, and the inference result containing the target object's location information and category label is output.
[0010] To address the aforementioned problems, this application also provides a neural network training system for few-shot target detection, the system comprising:
[0011] The sample data acquisition module is used to acquire a cross-scene sample dataset, which includes sample images under various lighting conditions and weather conditions to simulate the environmental conditions of the target airport.
[0012] The meta-learning training module is used to construct a feature extraction network based on deep meta-learning. The meta-learner is used to train the feature extraction network on the cross-scene sample dataset to generate the initial weight parameters of the target detection network. The training process includes: dynamically extracting sample subsets from the cross-scene sample dataset according to environmental attribute classification; constructing a meta-learning task simulating specific airport environmental conditions based on each sample subset; each meta-learning task includes a support set and a query set; performing an inner loop optimization process on the support set samples and an outer loop optimization process on the query set samples; and updating the network parameters through the gradient descent algorithm.
[0013] A detection network initialization module is used to initialize the target detection network using the initial weight parameters. The target detection network adopts a depthwise separable convolutional structure, which includes performing depthwise convolution operations and pointwise convolution operations on the input feature map.
[0014] The inference detection module is used to input the monitoring data of the target airport into the initialized target detection network for forward inference calculation and output the inference result containing the target object's location information and category label.
[0015] This invention effectively improves the generalization ability and computational efficiency of object detection models under limited sample conditions by introducing a deep meta-learning framework and a deep separable convolutional structure. First, a meta-learner is used to construct multiple learning tasks on a cross-scene sample dataset, and an inner-outer loop optimization strategy is employed to train the feature extraction network, enabling the model to learn initial weight parameters with cross-scene adaptability from a limited set of samples. This mechanism significantly enhances the model's ability to quickly adapt to new scenes, overcoming the overfitting and insufficient generalization problems caused by the scarcity of labeled data in traditional methods.
[0016] Secondly, a lightweight target detection network is constructed using a depthwise separable convolutional structure, significantly reducing the number of parameters and computational complexity while maintaining feature extraction capabilities. This structure achieves efficient multi-scale feature representation through the separation of spatial feature extraction and channel feature fusion, making it particularly suitable for the deployment requirements of edge computing devices. Furthermore, multimodal data fusion and refined post-processing mechanisms further enhance the model's detection accuracy and robustness in complex environments, ensuring accurate output of target location and category labels. This invention provides an efficient and reliable few-shot target detection solution for cross-scenario applications such as airport bird detection and security monitoring. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a neural network training method for few-shot target detection provided in an embodiment of this application.
[0018] Figure 2 A functional block diagram of a neural network training system for few-shot target detection provided in an embodiment of this application;
[0019] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0020] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0021] This application provides a neural network training method for few-shot target detection. The execution entity of the neural network training method for few-shot target detection includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the neural network training method for few-shot target detection can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.
[0022] Reference Figure 1 The diagram shown is a flowchart illustrating a neural network training method for few-shot target detection according to an embodiment of this application. In this embodiment, the neural network training method for few-shot target detection includes:
[0023] S1. Obtain a cross-scene sample dataset, which includes sample images under various lighting conditions and weather conditions to simulate the environmental conditions of the target airport.
[0024] In some embodiments, obtaining the cross-scenario sample dataset includes:
[0025] In this embodiment of the application, the cross-scene sample dataset is a collection of bird images containing various lighting conditions (daytime, nighttime) and various weather conditions (sunny, cloudy, rainy, foggy) to simulate bird detection tasks in different airport environments and improve the generalization ability of the model.
[0026] Multimodal sensing devices are used to collect cross-scene sample data, which includes sample images under various lighting conditions and weather conditions.
[0027] In detail, the multimodal sensing device includes: a visible light camera and an infrared thermal imaging device. The visible light camera is used to collect visible light image information of the airport and surrounding monitoring areas, and can capture the texture, color and other features of bird targets. It is suitable for scenarios with good lighting conditions. The infrared thermal imaging device is used to collect infrared thermal radiation information of the airport and surrounding monitoring areas. It can capture the thermal radiation characteristics of bird targets under adverse lighting or weather conditions such as at night, rain, and fog, and is not affected by changes in lighting.
[0028] In this embodiment, visible light camera equipment and infrared thermal imaging equipment are selected and deployed. The visible light camera equipment selected has a resolution of 1920×1080 and a frame rate of 25fps, which can cover a monitoring area of ±500 meters from the center of the airport runway and ±1000 meters in length. This area is the key area where 90% of bird strikes occur. The infrared thermal imaging equipment selected has a resolution of 640×512 and a temperature measurement range of -20℃ to 150℃, which can penetrate adverse weather conditions such as rain and fog, and supplement the need for bird target collection under no-light conditions at night.
[0029] In this embodiment of the application, the image sources of the cross-scene sample dataset include two parts: one part is the on-site images collected in the area surrounding the airport. During the collection process, areas around the airport with high bird activity are selected, and different flight postures of birds are photographed from different angles, under different lighting conditions (strong light during the day, weak light at dusk, and no light at night) and different weather conditions (sunny, cloudy, rainy, and foggy) to ensure the scene authenticity of the on-site samples; the other part is bird images supplemented from the Internet to enhance the diversity of bird species and avoid model overfitting caused by a single sample.
[0030] In this embodiment of the application, the acquired images are preprocessed, including image denoising, image size unification, and image annotation (the Labellmg tool is used to annotate the bird targets in each image, and the annotation content includes the bounding box coordinates of the bird targets and the species classification label of the bird targets. A total of 15 common airport birds, such as sparrows, doves, and eagles, are annotated), forming a complete cross-scene sample dataset.
[0031] S2. Construct a feature extraction network based on deep meta-learning, and train the feature extraction network on the cross-scene sample dataset using a meta-learner to generate the initial weight parameters of the object detection network.
[0032] In this embodiment of the application, this step constructs a bird feature extraction network architecture based on deep meta-learning. Unlike the limitations of traditional convolutional neural networks (CNNs) that are "trained and adapted in a single scene", this step achieves cross-scene generalization of the initial weight parameters through a meta-learner. At the same time, it incorporates flock bird scenes into the construction of the dataset and task, thus solving the problem of insufficient feature extraction of flock birds in traditional detection.
[0033] In this embodiment, deep meta-learning is a machine learning method that enables a model to quickly adapt to new scenarios or tasks, allowing the model to maintain good performance even with a small number of samples or across different scenarios. The feature extraction network is a neural network structure used to extract bird target features from image data, which can convert the original image data into discriminative feature vectors. The meta-learner is the core component in the deep meta-learning framework, responsible for learning the initial parameters of the feature extraction network, enabling the initial parameters to quickly adapt to different scenario tasks.
[0034] In some embodiments, the training process includes: sampling from the cross-scenario sample dataset to construct multiple learning tasks, and updating network parameters based on the inner and outer loop optimization strategy of deep meta-learning.
[0035] In this embodiment, a learning task refers to a task unit constructed based on a specific scenario (such as a rainy day environment) in a cross-scenario sample dataset, used to train a model to adapt to that scenario. Each learning task includes support set samples and query set samples. The inner and outer loop optimization strategy is the core training strategy of meta-learning. The inner loop is used to enable the feature extraction network to quickly adapt to a single learning task, while the outer loop is used to optimize the meta-learner to improve the generalization ability of the feature extraction network on all learning tasks.
[0036] In some embodiments, the construction of a feature extraction network based on deep meta-learning, and the training of the feature extraction network on the cross-scene sample dataset using a meta-learner to generate initial weight parameters for the object detection network, includes:
[0037] From the cross-scenario sample dataset, a subset of samples is dynamically extracted based on environmental attributes, and a meta-learning task simulating specific airport environmental conditions is constructed based on each subset of samples.
[0038] For each meta-learning task, an internal loop optimization process is performed on the support set samples of the meta-learning task, and the parameter gradient of the feature extraction network is calculated and the network parameters of the feature extraction network are updated by using the gradient descent algorithm.
[0039] An outer loop optimization process is performed on the query set samples of all the meta-learning tasks, and the parameter gradients of the meta-learner are calculated and the network parameters of the meta-learner are updated by using the gradient descent algorithm.
[0040] The inner loop optimization process and the outer loop optimization process are repeatedly executed iteratively. When the loss function value of the meta-learner is less than the set loss threshold, the iteration stops and the initial weight parameters of the target detection network are output.
[0041] In this embodiment, the support set samples are a subset of samples used in the learning task to quickly adapt and train the feature extraction network. By training with a small number of samples, the network parameters are initially adapted to the current task. The internal loop optimization process is a process of calculating the parameter gradient of the feature extraction network and updating the network parameters using the support set samples for a single learning task, thereby realizing the rapid adaptation of the network to the current task.
[0042] In this embodiment, gradient descent is an algorithm for optimizing neural network parameters. It calculates the gradient of the loss function with respect to the parameters and adjusts the parameters along the negative gradient direction to reduce the loss value. The parameter gradient is the partial derivative of the loss function with respect to the neural network parameters, reflecting the degree of influence of parameter changes on the loss value and serving as the basis for parameter updates. The query set samples are a subset of samples used in the learning task to verify the adaptation effect of the feature extraction network and calculate the meta-loss, and are used to evaluate the network's generalization ability on the current task. The outer loop optimization process is the process of calculating the meta-loss using the query set samples of all learning tasks and updating the meta-learner parameters to optimize the initial parameters of the feature extraction network.
[0043] In some embodiments, the loss function used in the inner loop optimization process is a meta-learning loss function, the expression of which is:
[0044]
[0045] in, It is the total loss function that the meta-learning training process needs to minimize. It is the first The loss function for individual learning tasks. It is the first Individual learning tasks It is the probability distribution of all possible meta-learning tasks. It is the first The new parameters obtained after updating the support set of the individual learning task through the inner loop. The parameter is Feature extraction network, These are the coefficients of the L2 regularization term. These are the initial parameters of the feature extraction network. The initial parameters are L2 norm squared, From task distribution Mid-sampling task The resulting losses are summed.
[0046] In this embodiment, the loss function is a mathematical function used to measure the difference between the output of the neural network and the true result. The loss function in this step is the meta-learning loss function, which takes into account the loss of each learning task and parameter regularization. The loss threshold is a critical value used to determine whether meta-learning training should stop. When the loss function value of the meta-learner is less than the threshold, it indicates that the initial parameters have good generalization ability and training can stop. The initial weight parameters refer to the initial parameters of the feature extraction network obtained after meta-learning training. These parameters have the ability to quickly adapt to different scene tasks and can be used to initialize the subsequent object detection network.
[0047] In this embodiment, the dataset is divided into 6 learning tasks based on its environmental attributes, corresponding to 6 typical airport scenarios: Task 1 is a sunny daytime scenario, Task 2 is a cloudy daytime scenario, Task 3 is a rainy daytime scenario, Task 4 is a foggy daytime scenario, Task 5 is a sunny nighttime scenario, and Task 6 is a cloudy nighttime scenario. Each learning task contains 16,700 images, which are split into support set samples and query set samples in a 7:3 ratio. That is, the support set samples of each task contain 11,700 images for internal loop optimization, and the query set samples contain 5,000 images for external loop optimization.
[0048] In this embodiment, a feature extraction network based on deep meta-learning is constructed. The basic architecture of the network adopts EfficientNet-B0 as the backbone network. This network has lightweight characteristics, with approximately 4 million parameters, which can reduce computational complexity while ensuring feature extraction capabilities. The meta-learner is integrated into the training framework of the feature extraction network and is responsible for managing and optimizing the initial parameters of the feature extraction network.
[0049] In this embodiment, an internal loop optimization process is performed, and optimization is performed sequentially for each learning task: First, the initial parameters (denoted as θ) of the feature extraction network are loaded into the training framework of the current task. The initial parameters are generated by random initialization and conform to a normal distribution. Then, the network is trained using support set samples. The optimizer is SGD (stochastic gradient descent) optimizer, the learning rate is set to 0.001, the momentum is set to 0.9, and the number of training iterations is set to 5 (a small number of iterations ensures that the network can quickly adapt to the current task and avoid overfitting to a single scenario).
[0050] In this embodiment, in each iteration of the inner loop, the loss value of the support set samples is calculated, the gradient of the loss function with respect to the network parameters is calculated using the gradient descent algorithm, and the network parameters are updated along the negative gradient direction to obtain temporary parameters adapted to the current learning task (denoted as...). The update formula is "temporary parameter = initial parameter - learning rate × parameter gradient", where i represents the i-th learning task.
[0051] In this embodiment, after completing the inner loop optimization of all learning tasks, the outer loop optimization process is executed: first, the temporary parameters of each learning task are... Loaded into the network, inference is performed using query set samples for the corresponding task, and the loss value for each task on the query set is calculated (denoted as ). Then, the total meta-loss value is calculated based on the meta-learning loss function, which is expressed as "total meta-loss value = expected value of all learning task query set loss values + L2 regularization term". The expected value is the average of the loss values of all learning tasks (since the tasks are sampled according to a uniform distribution, each task has the same weight). The L2 regularization term is "regularization coefficient × squared L2 norm of the initial parameters". The regularization coefficient is set to 0.001 to prevent the initial parameters from overfitting and to suppress the parameters from being too large in absolute value.
[0052] In this embodiment, the optimizer of the outer loop is the Adam optimizer, with a learning rate of 0.0001 and a weight decay of 0.0001. The gradient of the total meta-loss value with respect to the meta-learner parameters (i.e., the initial parameters θ of the feature extraction network) is calculated using the gradient descent algorithm, and the initial parameters θ are updated along the negative gradient direction to enable the initial parameters to have better cross-scene generalization ability.
[0053] In this embodiment, the inner loop optimization process and the outer loop optimization process are repeatedly executed. After each loop, the total element loss value is calculated and it is determined whether it is less than the set loss threshold (the loss threshold is set to 0.05). When the total element loss value is less than 0.05 for three consecutive iterations, training is stopped. At this time, the initial parameters of the output feature extraction network are the initial weight parameters of the target detection network. These parameters enable the network to achieve high detection accuracy in new scenarios (such as foggy conditions in an unfamiliar airport) with only 50 samples for fine-tuning.
[0054] In this embodiment of the application, this step solves the problem of robust feature extraction for small targets in all-weather changing environments. Through the inner and outer loop optimization of meta-learning, the initial weight parameters have good cross-scene generalization ability, avoiding the defect that the model cannot adapt to environmental changes after training in a single scene.
[0055] S3. Initialize the target detection network using the initial weight parameters. The target detection network adopts a depthwise separable convolutional structure.
[0056] In this embodiment, the target detection network is a neural network used to detect the location of bird targets and identify species categories from image data, and can output the bounding box coordinates and classification labels of bird targets; the depthwise separable convolutional structure is a lightweight convolutional structure that splits the standard convolution into two steps: depthwise convolution and pointwise convolution, which can significantly reduce the number of parameters and computation while ensuring feature extraction capabilities.
[0057] In some embodiments, initializing the object detection network using the initial weight parameters, wherein the object detection network employs a depthwise separable convolutional structure, includes:
[0058] The initial weight parameters are loaded into a preset convolutional neural network skeleton to generate a depthwise separable convolutional network structure, wherein the convolutional neural network skeleton contains multiple convolutional layers constructed by depthwise separable convolution operators.
[0059] In this embodiment, the convolutional neural network skeleton is the basic architecture of the object detection network, including a feature extraction layer, a feature fusion layer, and a detection layer, providing basic structural support for the network; the depthwise separable convolution operator is the core computational unit for realizing the depthwise separable convolution structure, including two parts: depthwise convolution operation and pointwise convolution operation.
[0060] In some embodiments, the operation of the depth-separable convolution operator includes:
[0061] A depthwise convolution operation is performed on the input feature map, where each input channel is independently convolved with a two-dimensional spatial convolution kernel to generate an intermediate feature map with the same number of channels.
[0062] A pointwise convolution operation is performed on the intermediate feature map, wherein a 1×1 convolution kernel is used to linearly combine all channels of the intermediate feature map to generate the final output feature map.
[0063] In this embodiment, the input feature map refers to the feature map input into the depthwise separable convolution operator, which contains the feature information extracted by the previous layer network; the depthwise convolution operation is the first step of the depthwise separable convolution operator, which uses a two-dimensional spatial convolution kernel to perform convolution calculation on each channel of the input feature map independently, without changing the number of channels of the feature map; the two-dimensional spatial convolution kernel is a matrix used to extract the spatial features of the feature map. In this step, a 3×3 convolution kernel is used, which can effectively capture the local spatial features of the target.
[0064] In this embodiment, the intermediate feature map is the output of the depthwise convolution operation, and the number of channels is the same as the input feature map. Figure 1The first step is to extract information from the spatial features. The second step is to perform a pointwise convolution operation, which uses a 1×1 convolution kernel to linearly combine all channels of the intermediate feature map, adjusting the number of channels. The 1×1 convolution kernel is a matrix used to fuse features from different channels, enabling information exchange between channels without changing the spatial size of the feature map. Linear combination refers to using a 1×1 convolution kernel to perform a weighted summation of the pixel values of each channel in the intermediate feature map, generating new channel features and achieving channel feature fusion. The final output feature map is the result of the pointwise convolution operation, containing complete feature information after spatial feature extraction and channel feature fusion, which is used for calculations in subsequent network layers.
[0065] In this embodiment, the basic architecture of the target detection network is first determined. A lightweight network architecture of improved YOLOv5 (denoted as EY-slim) is adopted as the convolutional neural network skeleton. The skeleton includes four parts: input layer, feature extraction layer, feature fusion layer and detection layer, which can take into account both detection accuracy and real-time requirements.
[0066] In this embodiment, the structure of each layer of the convolutional neural network skeleton is optimized by replacing all standard convolutional layers with depthwise separable convolution operators. Specifically, the feature extraction layer adopts the EfficientNet-B0 architecture in step S2, but all standard convolutions are replaced with a combination of 3×3 depthwise convolutions and 1×1 pointwise convolutions. The feature fusion layer removes two redundant standard convolutional layers, retains feature fusion paths of three scales (19×19, 38×38, 76×76), and replaces the standard convolutions in the paths with depthwise separable convolutions. The detection layer sets up three detection heads (corresponding to feature maps of three scales respectively), and the convolutional layer of each detection head adopts depthwise separable convolution.
[0067] In this embodiment of the application, a depthwise separable convolution operator of a feature extraction layer is used as an example to illustrate its specific operation process: the input feature map of the operator has 64 channels and a spatial size of 304×304; firstly, a depthwise convolution operation is performed, allocating one 3×3 two-dimensional spatial convolution kernel to each channel of the input feature map, for a total of 64 3×3 convolution kernels. Each convolution kernel performs convolution calculation only on the feature map of the corresponding channel, with a convolution stride of 1 and a padding of 1, ensuring that the spatial size of the output intermediate feature map remains 304×304 and the number of channels remains unchanged at 64.
[0068] In this embodiment, a pointwise convolution operation is then performed, using 128 1×1 convolution kernels to perform convolution calculations on the intermediate feature map. Each 1×1 convolution kernel performs a linear combination of all 64 channels of the intermediate feature map (i.e., weighted summation of the 64 channel pixel values at each spatial location) to generate 128 new channels. The final output feature map has a spatial size of 304×304 and 128 channels.
[0069] In this embodiment, the number of parameters and computational cost of the depthwise separable convolution operator are calculated as follows: The number of parameters for a standard 3×3 convolution is "number of input channels × number of output channels × kernel size × kernel size" = 64 × 128 × 3 × 3 = 73728, and the computational cost is proportional to the number of parameters; the number of parameters for the depthwise separable convolution is "number of depthwise convolution parameters + number of pointwise convolution parameters" = 64 × 3 × 3 + 64 × 128 × 1 × 1 = 576 + 8192 = 8768, and the number of parameters is only 11.9% of that of the standard convolution, which significantly reduces the computational complexity of the network.
[0070] In this embodiment, a Spatial Pyramid Pooling (SPP) structure and an Intersection over Union (IoU) prediction branch are added to the convolutional neural network skeleton. The SPP structure is set between the feature extraction layer and the feature fusion layer, and four max pooling kernels of different sizes (1×1, 2×2, 4×4, 8×8) are used to pool the feature maps to enhance the network's ability to extract features from bird targets of different sizes (especially small targets). The IoU prediction branch runs in parallel with the classification and regression branches of the detection layer to predict the IoU value between the detection box and the ground truth box, thereby improving the accuracy of bounding box regression.
[0071] In this embodiment, the initial weight parameters generated in step S2 are loaded to initialize the parameters of the target detection network: the part of the initial weight parameters corresponding to EfficientNet-B0 is loaded into the depthwise separable convolution operator of the feature extraction layer according to the correspondence between channels and layers; the parameters of the depthwise separable convolution operator of the feature fusion layer and the detection layer are initialized using the Xavier initialization method to ensure uniform parameter distribution and avoid gradient vanishing in the early stage of training.
[0072] In this embodiment of the application, after initialization, the total number of parameters of the target detection network is controlled within 4.8 million, and the inference frame rate (on Raspberry Pi 4B hardware) can reach 25fps, which meets the requirements of real-time deployment of airport edge devices. At the same time, the detection accuracy of small target birds is ensured through the SPP structure and IoU prediction branch.
[0073] In this embodiment, this step solves the problem of low detection efficiency for large flocks of birds in low-altitude areas of airports. It reduces the computational complexity of the network by using depthwise separable convolution and improves the detection accuracy of small targets by combining SPP and IoU branches, thus achieving the construction of an efficient and accurate target detection network.
[0074] In this embodiment of the application, the target detection network initialized in this step is the core model for forward inference computation in step S4. Step S4 requires inputting bimodal video stream data into the network to output detection results. The two form a connection between model construction and model inference.
[0075] S4. Input the monitoring data of the target airport into the initialized target detection network for forward inference calculation, and output the inference result containing the target object location information and category label.
[0076] In this embodiment, forward inference computation refers to the process of inputting input data into an initialized target detection network and obtaining the output result through the computation of each layer of the network. It does not involve parameter updates, but only feature transfer and prediction. Target location information refers to the coordinates of a rectangular box used to describe the position and size of the bird target in the image. It is usually based on the upper left corner of the image as the origin and includes the x-coordinate of the upper left corner, the y-coordinate of the upper left corner, the x-coordinate of the lower right corner, and the y-coordinate of the lower right corner of the rectangular box. Species classification label refers to the category label used to identify the species to which the bird target belongs, corresponding to the 15 common airport birds labeled in the dataset, such as "sparrow", "dove", and "eagle".
[0077] In this embodiment of the application, the inference result refers to the final output of the forward inference calculation, which includes the bounding box coordinates, species classification labels and corresponding confidence scores (reflecting the reliability of the detection results) of all detected bird targets.
[0078] In some embodiments, the step of inputting the monitoring data of the target airport into the initialized target detection network for forward inference calculation and outputting an inference result containing the target object's location information and category label includes:
[0079] The monitoring data of the target airport is converted into a multimodal fusion tensor;
[0080] The multimodal fusion tensor is input into the target detection network, and forward inference is performed through the target detection network to generate the network's original output containing one or more detection boxes. The detection boxes contain location coordinates, confidence scores, and classification label information.
[0081] The original output of the network is post-processed to remove redundant detection boxes and select the best detection results, and the final determined target object location information and corresponding category label are output.
[0082] In this embodiment, the multimodal fusion tensor refers to the tensor data obtained after multi-channel dimensional stitching processing, which contains feature information of both visible light and infrared modes, and the dimension is "image height × image width × number of channels".
[0083] In this embodiment, post-processing refers to the process of filtering and optimizing the original network output, mainly including confidence filtering and non-maximum suppression (NMS) processing, used to remove redundant detection boxes and low-confidence detection boxes; redundant detection boxes refer to multiple overlapping detection boxes generated for the same bird target in the original network output, which need to be removed by post-processing, retaining only the best one; the optimal detection result refers to the detection result obtained after post-processing, which contains only high-confidence, non-redundant detection boxes, and can accurately reflect the location and species information of bird targets in the image.
[0084] In some embodiments, converting the monitoring data of the target airport into a multimodal fusion tensor includes:
[0085] The monitoring data is modally segmented, and the different modal data are concatenated along the channel dimension to generate a multimodal fusion tensor.
[0086] In this embodiment of the application, the real-time data stream of a certain monitoring point at the target airport includes the synchronous output of a visible light camera and an infrared camera. The two have been aligned by timestamps to ensure that they are images from the same time and the same area.
[0087] In this embodiment of the application, the original network output refers to the output result directly obtained by the forward inference calculation of the target detection network, which includes multiple detection boxes. Each detection box contains bounding box coordinates, confidence score and classification label, but there may be redundant detection boxes.
[0088] In this embodiment, confidence level refers to the degree of confidence of the target detection network in whether the target box contains a bird target and whether the classification label is correct. The value ranges from 0 to 1, and the larger the value, the more reliable the detection result.
[0089] In this embodiment, the multimodal fusion tensor is input into the target detection network initialized in step S3, and forward inference calculation is performed: First, the multimodal fusion tensor is input into the feature extraction layer, and multi-scale features are extracted through the depthwise separable convolution operator and the EfficientNet-B0 architecture, outputting feature maps of three scales: 19×19×512, 38×38×256, and 76×76×128.
[0090] In this embodiment, the feature fusion layer fuses feature maps at three scales: the 19×19×512 feature map is upsampled to 38×38×512 using an upsampling algorithm, and then concatenated with the 38×38×256 feature map in the channel dimension to obtain a 38×38×768 feature map; this feature map is then upsampled to 76×76×768 and concatenated with the 76×76×128 feature map to obtain a 76×76×896 feature map; simultaneously, the 76×76×128 feature map is downsampled to 38×38×128 and concatenated with the 38×38×256 feature map to form multi-scale feature interaction, enhancing the propagation of small target features.
[0091] In this embodiment, three detection heads perform detection calculations on feature maps of 19×19×512, 38×38×768, and 76×76×896 respectively. Each detection head outputs detection results through a depthwise separable convolution operator, including bounding box coordinate prediction, confidence prediction, and classification label prediction. The 19×19 scale detection head is used to detect larger bird targets (such as eagles), the 38×38 scale is used to detect medium-sized targets (such as doves), and the 76×76 scale is used to detect small targets (such as sparrows).
[0092] In this embodiment of the application, the original network output contains a total of 300 detection boxes generated by three detection heads. Each detection box contains 4 bounding box coordinate values (ranging from 0 to 1, which need to be multiplied by the image resolution of 1920 and 1080 to convert to actual pixel coordinates), 1 confidence value (0 to 1), and 1 species classification label (an integer from 1 to 15, corresponding to 15 bird species).
[0093] In this embodiment, the original network output is post-processed as follows: First, confidence filtering is performed to remove detection boxes with confidence values less than 0.7 and retain those with confidence values ≥ 0.7. This threshold is determined experimentally and ensures that the false detection rate of the filtered detection boxes is ≤ 3%. Then, non-maximum suppression (NMS) processing is performed to sort the filtered detection boxes from high to low confidence and calculate the intersection-over-union (IoU) ratio between adjacent detection boxes. When IoU > 0.5, the detection boxes with lower confidence are removed and the detection boxes with the highest confidence are retained to avoid the same target being detected multiple times.
[0094] In this embodiment of the application, after post-processing, the final inference result is output, which includes the actual pixel coordinates of each detected bird target (such as the upper left corner (200,300) and the lower right corner (250,350)), the corresponding species classification label (such as "sparrow"), and the confidence value (such as 0.85), providing accurate target information for the subsequent generation of bird control commands.
[0095] In the embodiments of this application, this step solves the problem of extracting fine-grained features with different bird discrimination capabilities, and achieves accurate detection and species identification of bird targets through multimodal fusion and multi-scale detection, avoiding misjudgment of the same species and missed judgment of different species.
[0096] In some embodiments, after the output includes the reasoning result containing the target object location information and category label, it further includes:
[0097] Based on the target location information and the category label, a corresponding bird control command is generated to execute the bird deterrence action.
[0098] In this embodiment, generating corresponding bird control commands based on the target location information and the category label to execute bird-repelling actions includes: acquiring target location information and corresponding image acquisition device parameters from the inference result; establishing a mapping transformation relationship between the image coordinate system and the world coordinate system based on the image acquisition device parameters; converting the target location information into actual geospatial coordinates based on the mapping transformation relationship; performing spatial relationship analysis between the actual geospatial coordinates and a preset airport security area geofence range; generating a high-risk level assessment result when the actual geospatial coordinates are within the airport security area geofence range; generating bird-repelling control commands containing target location coordinates and repelling intensity parameters based on the high-risk level assessment result; transmitting the bird-repelling control commands to the corresponding bird-repelling execution device via a wireless communication network; and controlling the bird-repelling execution device to execute corresponding automated repelling operations according to the bird-repelling control commands.
[0099] In this embodiment, the parameters of the image acquisition device are first obtained. The intrinsic parameters are those calculated during the spatial registration process in step S1 (visible light camera: 1000 pixels horizontal focal length, 1000 pixels vertical focal length, principal point coordinates (960, 540); infrared thermal imaging device: 500 pixels horizontal focal length, 500 pixels vertical focal length, principal point coordinates (320, 256)). The extrinsic parameters are obtained through a GPS positioning device and an IMU attitude sensor. It is assumed that the visible light camera is installed on a 10-meter-high support next to the airport runway, and its position coordinates in the world coordinate system are (0, 0, 10), and its attitude is horizontal. The corresponding rotation matrix is the identity matrix.
[0100] In this embodiment, a mapping transformation relationship is established based on the parameters of the image acquisition device. This relationship is derived through the principle of perspective projection, and the specific formula is as follows: X value of actual geographic spatial coordinates = (u value of image pixel coordinates - horizontal coordinate of principal point) × target distance ÷ horizontal focal length + X value of device world coordinates; Y value of actual geographic spatial coordinates = (v value of image pixel coordinates - vertical coordinate of principal point) × target distance ÷ vertical focal length + Y value of device world coordinates; Z value of actual geographic spatial coordinates = target distance + Z value of device world coordinates; where the target distance is the straight-line distance (unit: meters) between the bird target and the device obtained by the ranging function of the infrared thermal imaging device.
[0101] In this embodiment of the application, taking the detection result of a bird target output in step S4 as an example, the coordinate transformation process is explained: the image pixel coordinates of the target are (1000, 560), and the target distance measured by the infrared thermal imaging device is 50 meters; substituting the parameters into the formula for calculation: the X value is (1000-960)×50÷1000+0=2 meters; the Y value is (560-540)×50÷1000+0=1 meter; the Z value is 50+10=60 meters; the actual geographic spatial coordinates of the target are obtained as (2,1,60).
[0102] In this embodiment of the application, a geofence for the airport security zone is preset. Based on statistical data that 75% of bird strikes occur at heights below 60 meters and 90% occur in and around the airport, the fence range is set as follows: X-axis range -500 meters to 500 meters (runway transverse), Y-axis range -1000 meters to 1000 meters (runway longitudinal), and Z-axis range 0 meters to 60 meters (height). This range covers the airport runway and surrounding key areas, ensuring that all high-risk areas are included in the monitoring.
[0103] In this embodiment of the application, spatial relationship analysis is performed, and the calculated actual geospatial coordinates of the bird target are compared with the geographical fence range of the airport security area. Taking the above target coordinates (2,1,60) as an example, the X value of 2 meters is in the range of -500 to 500 meters, the Y value of 1 meter is in the range of -1000 to 1000 meters, and the Z value of 60 meters is in the range of 0 to 60 meters. Therefore, the target is determined to be located in a high-risk area, and a high-risk level assessment result is generated.
[0104] In this embodiment of the application, based on the high-risk level assessment results and the species classification label output in step S4, the driving intensity parameters are generated: when the classification label is a bird of prey (such as an eagle), the laser device power is set to 500mW and the sound wave generator frequency is set to 20000Hz. This parameter combination has a strong deterrent effect on birds of prey; when the classification label is a common bird (such as a sparrow or a dove), the laser device power is set to 200mW and the sound wave generator frequency is set to 15000Hz. This parameter combination can effectively drive away common birds without causing them harm.
[0105] In this embodiment of the application, a bird deterrence control command is generated. The command is structured data, which includes the actual geographic spatial coordinates of the target (e.g., (2,1,60)), deterrence intensity parameters (e.g., laser power 500mW, sound frequency 20000Hz), and bird deterrence execution device identifier (e.g., UAV number 01, laser device number 02). The command format adopts JSON format to ensure that the device can parse it.
[0106] In this embodiment, the bird deterrence device executes an automated deterrence operation after receiving the instruction: the drone automatically plans its flight path based on the target's actual geographic spatial coordinates (2,1,60), flies from the takeoff point to within 10 meters of the target, and ensures a position accuracy of ≤1 meter through GPS positioning during the flight; after arriving at the target area, the laser device emits a 532nm wavelength laser at a set power (500mW) (this wavelength is sensitive to birds and will not harm their eyes), and the sound wave generator plays sound waves at a set frequency (20000Hz) (this frequency band is sensitive to birds and has no effect on humans).
[0107] In this embodiment of the application, during the bird driving process, the target detection network in step S4 continuously monitors the location of the bird target. When the actual geographic spatial coordinates of the target are detected to be outside the geographical fence of the airport security area, a stop bird driving command is generated and transmitted to the bird driving execution device through the wireless communication network. The device stops laser emission and sound wave playback, completing an automated bird driving action.
[0108] like Figure 2 The diagram shown is a functional block diagram of a neural network training system for few-sample target detection provided in an embodiment of this application.
[0109] The neural network training system 100 for few-shot target detection described in this application can be installed in an electronic device. Depending on the functions implemented, the neural network training system 100 for few-shot target detection may include a sample data acquisition module 101, a meta-learning training module 102, a detection network initialization module 103, and an inference detection module 104. The module described in this application can also be referred to as a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and which are stored in the memory of the electronic device.
[0110] In this embodiment, the functions of each module / unit are as follows:
[0111] The sample data acquisition module 101 is used to acquire a cross-scene sample dataset, which includes sample images under various lighting conditions and weather conditions to simulate the environmental conditions of the target airport.
[0112] The meta-learning training module 102 is used to construct a feature extraction network based on deep meta-learning. The meta-learner is used to train the feature extraction network on the cross-scene sample dataset to generate the initial weight parameters of the target detection network. The training process includes: dynamically extracting sample subsets from the cross-scene sample dataset according to environmental attribute classification; constructing a meta-learning task simulating specific airport environmental conditions based on each sample subset; each meta-learning task includes a support set and a query set; performing an inner loop optimization process on the support set samples and an outer loop optimization process on the query set samples; and updating the network parameters through the gradient descent algorithm.
[0113] The detection network initialization module 103 is used to initialize the target detection network using the initial weight parameters. The target detection network adopts a depthwise separable convolutional structure, which includes performing depthwise convolution operations and pointwise convolution operations on the input feature map.
[0114] The inference detection module 104 is used to input the monitoring data of the target airport into the initialized target detection network for forward inference calculation and output the inference result containing the target object location information and category label.
[0115] In the embodiments provided in this application, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0116] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0117] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0118] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application.
[0119] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of this application without departing from the spirit and scope of the technical solutions of this application.
Claims
1. A neural network training method for few-shot target detection, characterized in that, The method includes: A cross-scene sample dataset is obtained, which includes sample images under various lighting conditions and weather conditions to simulate the environmental conditions of the target airport. A feature extraction network based on deep meta-learning is constructed. The network is trained on the cross-scene sample dataset using a meta-learner to generate initial weight parameters for the target detection network. The training process includes: dynamically extracting sample subsets from the cross-scene sample dataset according to environmental attributes; constructing a meta-learning task simulating specific airport environmental conditions based on each sample subset; each meta-learning task including a support set and a query set; performing an inner loop optimization process on the support set samples and an outer loop optimization process on the query set samples; and updating network parameters using a gradient descent algorithm. Specifically, the cross-scene sample dataset is divided into six meta-learning tasks according to environmental attributes, corresponding to six airport scenarios: sunny daytime, cloudy daytime, rainy daytime, foggy daytime, sunny nighttime, and cloudy nighttime. For each meta-learning task, the network is trained on the support set samples. An internal loop optimization process is performed, in which the gradient gradient of the feature extraction network is calculated and the network parameters of the feature extraction network are updated using the gradient descent algorithm. The backbone network of the feature extraction network adopts the EfficientNet-B0 structure. The internal loop optimization process uses the SGD optimizer with a learning rate of 0.001 and 5 iterations. An external loop optimization process is performed on all query set samples of the meta-learning task, in which the gradient gradient of the meta-learner is calculated and the network parameters of the meta-learner are updated using the gradient descent algorithm. The external loop optimization process uses the Adam optimizer with a learning rate of 0.0001 and a weight decay coefficient of 0.0001. The internal loop optimization process and the external loop optimization process are repeatedly iterated. Iteration stops when the loss function value of the meta-learner is less than a set loss threshold, and the initial weight parameters of the object detection network are output. The target detection network is initialized using the initial weight parameters. The target detection network adopts a depthwise separable convolutional structure, which includes performing depthwise convolution and pointwise convolution operations on the input feature map. The initial weight parameters are loaded into a preset convolutional neural network skeleton to generate a depthwise separable convolutional network structure. The convolutional neural network skeleton adopts an improved YOLOv5 lightweight network architecture and includes an input layer, a feature extraction layer, a feature fusion layer, and a detection layer. The convolutional neural network skeleton contains multiple convolutional layers constructed by depthwise separable convolution operators. The monitoring data of the target airport is input into the initialized target detection network for forward inference calculation, and the inference result containing the target object's location information and category label is output.
2. The neural network training method for few-shot target detection as described in claim 1, characterized in that, The acquisition of cross-scenario sample datasets includes: Multimodal sensing devices are used to collect cross-scene sample data, including visible light cameras and infrared thermal imaging devices. The visible light camera covers a monitoring area of ±500 meters from the center of the target airport runway and ±1000 meters in length. The infrared thermal imaging device has a temperature measurement range of -20℃ to 150℃ and is used to collect infrared thermal radiation information of bird targets under night, rainy, or foggy conditions.
3. The neural network training method for few-shot target detection as described in claim 2, characterized in that, The loss function used in the inner loop optimization process is the meta-learning loss function, whose expression is: in, It is the total loss function that the meta-learning training process needs to minimize. It is the loss function for the i-th meta-learning task. It is the first Individual learning tasks It is the probability distribution of all possible meta-learning tasks. It is the first The new parameters obtained after updating the support set of the individual learning task through the inner loop. The parameter is Feature extraction network, These are the coefficients of the L2 regularization term. These are the initial parameters of the feature extraction network. It is the square of the L2 norm of the initial parameters. From task distribution Mid-sampling task The resulting losses are summed.
4. The neural network training method for few-shot target detection as described in claim 1, characterized in that, The operation process of the depth-separable convolution operator includes: A depthwise convolution operation is performed on the input feature map, where each input channel is independently convolved with a two-dimensional spatial convolution kernel to generate an intermediate feature map with the same number of channels. A pointwise convolution operation is performed on the intermediate feature map, wherein a 1×1 convolution kernel is used to linearly combine all channels of the intermediate feature map to generate the final output feature map.
5. The neural network training method for few-shot target detection as described in claim 1, characterized in that, The step involves inputting the monitoring data of the target airport into the initialized target detection network for forward inference calculation, and outputting an inference result containing the target object's location information and category label, including: The monitoring data of the target airport is converted into a multimodal fusion tensor; The multimodal fusion tensor is input into the target detection network, and forward inference is performed through the target detection network to generate the network's original output containing one or more detection boxes. The detection boxes contain location coordinates, confidence scores, and classification label information. The original output of the network is post-processed to remove redundant detection boxes and select the best detection results. The final determined target object location information and corresponding category label are output. The post-processing includes confidence filtering and non-maximum suppression processing, wherein the confidence threshold is set to 0.7 and the IoU threshold for non-maximum suppression is set to 0.
5.
6. The neural network training method for few-shot target detection as described in claim 5, characterized in that, The step of converting the monitoring data of the target airport into a multimodal fusion tensor includes: The monitoring data is modally segmented, and the different modal data obtained are spliced together by channel dimension to generate a multimodal fusion tensor. The input size of the multimodal fusion tensor is 1920×1080×4, where the first 3 channels are visible light images and the 4th channel is an infrared image.
7. The neural network training method for few-shot target detection as described in claim 1, characterized in that, After outputting the inference results, which include the target object's location information and category label, the following is also included: Based on the target object location information and the category label, a corresponding bird deterrence control command is generated to execute the bird deterrence action. The bird deterrence control command includes the actual geographic spatial coordinates of the target, the deterrence intensity parameter, and the bird deterrence execution device identifier. When the classification label is raptor, the driving intensity parameters are set to laser power 500mW and sound frequency 20000Hz. When the classification label is "common birds", the driving intensity parameters are set to laser power 200mW and sound frequency 15000Hz.
8. A neural network training system for few-shot target detection, used to implement the neural network training method for few-shot target detection according to any one of claims 1-7, characterized in that, The system includes: The sample data acquisition module is used to acquire a cross-scene sample dataset, which includes sample images under various lighting conditions and weather conditions to simulate the environmental conditions of the target airport. The meta-learning training module is used to construct a feature extraction network based on deep meta-learning. The meta-learner is used to train the feature extraction network on the cross-scene sample dataset to generate initial weight parameters for the target detection network. The training process includes: dynamically extracting sample subsets from the cross-scene sample dataset according to environmental attributes; constructing a meta-learning task simulating specific airport environmental conditions based on each sample subset; each meta-learning task including a support set and a query set; performing an inner loop optimization process on the support set samples and an outer loop optimization process on the query set samples; and updating network parameters using a gradient descent algorithm. Specifically, this includes: dividing the cross-scene sample dataset into six meta-learning tasks according to environmental attributes, corresponding to six airport scenarios: sunny daytime, cloudy daytime, rainy daytime, foggy daytime, sunny nighttime, and cloudy nighttime. For each meta-learning task, the support set of the meta-learning task is used to perform a meta-learning task on the support set samples and a meta-learning task on the query set samples. An inner loop optimization process is performed on the sample set of the meta-learning task. The gradient descent algorithm is used to calculate the parameter gradient of the feature extraction network and update its network parameters. The backbone of the feature extraction network uses an EfficientNet-B0 structure. The inner loop optimization process uses the SGD optimizer with a learning rate of 0.001 and 5 iterations. An outer loop optimization process is performed on the query set of all meta-learning tasks. The gradient descent algorithm is used to calculate the parameter gradient of the meta-learner and update its network parameters. The outer loop optimization process uses the Adam optimizer with a learning rate of 0.0001 and a weight decay coefficient of 0.0001. The inner and outer loop optimization processes are iteratively repeated. Iteration stops when the loss function value of the meta-learner is less than a set loss threshold, and the initial weight parameters of the target detection network are output. A detection network initialization module is used to initialize the target detection network using the initial weight parameters. The target detection network adopts a depthwise separable convolutional structure, which includes performing depthwise convolution operations and pointwise convolution operations on the input feature map. The initial weight parameters are loaded into a preset convolutional neural network skeleton to generate a depthwise separable convolutional network structure. The convolutional neural network skeleton adopts an improved YOLOv5 lightweight network architecture and includes an input layer, a feature extraction layer, a feature fusion layer, and a detection layer. The convolutional neural network skeleton includes multiple convolutional layers constructed by depthwise separable convolution operators. The inference detection module is used to input the monitoring data of the target airport into the initialized target detection network for forward inference calculation and output the inference result containing the target object's location information and category label.