Intelligent airport bird repelling system based on multi-target detection optimization
By combining radar detection and visual confirmation modules in the airport intelligent bird repelling system, using the target tracking model and deep learning framework, the attention mechanism and online update strategy are introduced, and the existing system's low image recognition accuracy in complex environments is solved, achieving higher detection accuracy and system real-time performance.
Patent Information
- Application Number
- CN202510269517.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
The existing airport intelligent bird repelling system has shortcomings in real-time radar updates, visual recognition sample collection and data processing efficiency, especially in complex environments, which is low in image recognition accuracy, which affects the system's real-time response capabilities.
The airport intelligent bird-repelling system based on multi-objective detection optimization is adopted, combined with the radar detection module and the visual confirmation module, bird target detection and tracking is carried out through the target tracking model and deep learning framework, and attention mechanisms and online update strategies are introduced to improve detection accuracy and adaptability.
It significantly improves the detection accuracy and tracking stability of bird targets, overcomes the problem of low image recognition accuracy in complex environments, and improves the system's real-time response and adaptability.
Smart Images

Figure CN120198641A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and particularly relates to an intelligent bird repelling system for airports optimized based on multi-target detection. Background Art
[0002] The intelligent bird repelling system for airports generally includes: detection technology, control technology, and bird repelling technology; the detection technology part mainly collects specific information on bird intrusion into the airport through methods such as radar detection, visual detection, and sound recognition; the control technology part is mainly responsible for communication and data processing, and processes the information collected by the detection part through an algorithm program to obtain the best bird repelling plan; the bird repelling technology part executes the bird repelling plan issued by the control part through bird repelling devices such as sound, laser, and medicine.
[0003] The current system still has deficiencies in terms of radar real-time update, visual recognition sample collection, and data processing efficiency. For example, in the case of strong bird flight mobility and weak energy, the radar detection technology is difficult to achieve continuous and stable tracking. At the same time, when birds face a large number of low-altitude target interferences during the migration season, the information redundancy problem also affects the real-time performance of the system. In addition, the accuracy of image recognition technology in complex environments (such as lighting, angle, occlusion, etc.) is relatively low, further restricting the real-time response ability of the system.
[0004] The airport environment is complex, and bird activities are affected by factors such as lighting, weather, angle, and occlusion, resulting in a decrease in the accuracy of image recognition. For example, in some cases, the appearance similarity of birds is relatively high or it is difficult to accurately identify due to environmental factors (such as insufficient light). In addition, although the radar and visual image fusion technology can improve the recognition effect, its performance still has limitations under different weather conditions.
[0005] The current image recognition system still needs to be improved in terms of processing speed and stability. For example, although deep learning models such as YOLOv5 can achieve a relatively high detection speed, there may still be problems of missed detection or false detection in small target detection. In addition, algorithms based on deep learning require a large amount of computing resources and may not be able to meet the real-time requirements.
[0006] Current image recognition algorithms often are designed for specific scenarios and lack adaptability to different bird species and behavior patterns. For example, some systems rely on specific training data sets and do not fully consider the diversity and dynamic changes of bird behavior. In addition, existing algorithms also have deficiencies in generalization ability, such as being unstable under different lighting conditions.
[0007] There is still room for improvement in the existing systems in terms of data processing and model optimization. For example, some systems rely on traditional machine learning methods and fail to fully utilize the potential of deep learning techniques. At the same time, the problem of sample imbalance in the data annotation and model training processes will also affect the recognition effect. Summary of the Invention
[0008] To solve the above problems existing in the prior art, the present invention provides an intelligent bird repelling system for airports based on multi-objective detection optimization;
[0009] The object of the present invention can be achieved by the following technical solutions:
[0010] An intelligent bird repelling system for airports based on multi-objective detection optimization, comprising: a detection unit, an identification unit, a repelling unit, and an intelligent optimization unit;
[0011] The detection unit includes a radar detection module and a visual confirmation module; the radar detection module uses a two-dimensional cylindrical phased array radar and DBF technology to obtain the three-dimensional coordinate information of all-airspace targets, sets the scanning mode and tracking mode, and performs periodic scanning on the airspace around the airport to detect the activities of birds; the tracking mode is to continuously track the detected bird targets, update their position information in real time, and send the position information to the visual confirmation module; the visual confirmation module uses a high-definition camera and an image recognition algorithm to visually confirm the bird targets detected by the radar detection module;
[0012] The identification unit is used to identify according to the image information provided by the visual confirmation module through a target tracking model, and evaluate the potential threat level of birds according to a preset bird threat level database to obtain an evaluation result; the repelling unit selects a bird repelling strategy to drive away the birds according to the evaluation result of the identification unit; the intelligent optimization unit optimizes the bird repelling strategy according to historical bird repelling data and real-time environmental information;
[0013] Specifically, the target tracking model uses a target detection algorithm to detect the bird species and locate the targets, uses multi-object tracking technology to track the birds, and assigns a unique ID to each bird target; adopts a deep sorting algorithm, defines a tracking scenario in a multi-dimensional state space, and combines the counting method of species, position, and ID information for area counting.
[0014] Specifically, the method for area counting is as follows: create a matrix with the same size as the image and fill it with 1 to represent the counting area; after detecting a bird, replace the value at the corresponding position with 0 and reassign it to 1 to mark the presence of the bird; finally, count the number of birds according to the number of 1s in the matrix; for newly emerged bird IDs, record the ID and increase the number of birds of that species; for the already recorded IDs, keep the number of birds of that species unchanged.
[0015] Specifically, the tracking scenario of the target tracking model is defined on a multi-dimensional state space, and the dimensions of the multi-dimensional state space include the centroid coordinates of the detection frame, the aspect ratio of the image, the height of the detection frame, and their respective speeds in the image coordinates; observation variables are formed according to the dimensions, Kalman filtering is used to predict the observation variables, and a survival period threshold and a confirmation mechanism are set to obtain the tracking result;
[0016] Specifically, the visual confirmation module locates the target in the image through bounding box regression, and uses the intersection over union (IoU) as the loss function. When calculating the IoU, a penalty term is introduced into the loss function according to the distance between the center points of the predicted box and the ground truth and the aspect ratio. The calculation formula is:
[0017]
[0018] where α is the weight coefficient used to adjust the influence degree of the center point distance and the aspect ratio on the loss function, β is the penalty term weight of the center point distance, IoU is the IoU value representing the overlapping degree between the predicted box and the ground truth box, d c is the Euclidean distance between the center points of the predicted box and the ground truth box, w and h are the width and height of the predicted box respectively, w g and h g are the width and height of the ground truth box respectively;
[0019] By introducing the penalty term, when the center point of the predicted box is farther from the center point of the ground truth box, or the aspect ratio of the predicted box is more different from the aspect ratio of the ground truth box, the value of the loss function is larger, thereby guiding the model to more accurately locate the target during the training process.
[0020] Specifically, the target tracking model is constructed using a deep learning framework, and an attention mechanism is introduced to improve the tracking accuracy. Image features are extracted through a convolutional neural network, and a recurrent neural network is used to model the feature sequence to capture the motion pattern of the target; on this basis, an attention mechanism is introduced to weight the key regions in the feature map to enhance the attention of the model to the target region; the target tracking model also adopts an online update strategy to fine-tune the model parameters according to the real-time tracking results.
[0021] Specifically, the recurrent neural network is composed of two parts: a region proposal network (RPN) and a Fast R-CNN detector, and the two share the convolutional layer; the RPN generates region proposals through a sliding window and an anchor mechanism, while the Fast R-CNN detector is responsible for classifying and regressing these region proposals; during the training process, an alternating training strategy is adopted to train the RPN and the Fast R-CNN respectively, and through iterative optimization, they are finally fused into a unified network.
[0022] Specifically, the training method of the region proposal network is as follows: Set the center point of the sliding window as the anchor point, and assign a binary class label to each anchor. The condition for the positive label is: the anchor point with the highest intersection-over-union ratio with a ground truth box, or the anchor point with an intersection-over-union ratio greater than the preset overlap threshold with any ground truth box; the condition for the negative label is: the anchor point with an intersection-over-union ratio less than the preset overlap threshold with all ground truth boxes. The loss function of the region proposal network includes classification loss and regression loss. The classification loss uses the cross-entropy loss function to distinguish positive and negative anchor points; the regression loss uses the smooth L1 loss function to optimize the position and size of the anchor points.
[0023] The beneficial effects of the present invention are as follows:
[0024] Through the airport intelligent bird repelling system based on multi-object detection optimization provided by the present invention, the detection accuracy and tracking stability of bird targets can be significantly improved. First, by the collaborative work of the radar detection module and the visual confirmation module, the problem of low image recognition accuracy in complex environments is effectively overcome. The radar detection module can obtain the three-dimensional coordinate information of bird targets in real time, and perform periodic scanning and continuous tracking, while the visual confirmation module accurately confirms the targets detected by the radar through a high-definition camera and an image recognition algorithm, improving the recognition accuracy.
[0025] In addition, the target tracking model in the present invention adopts a deep learning framework, and introduces an attention mechanism and an online update strategy, significantly improving the tracking accuracy and adaptability. Image features are extracted through a convolutional neural network, and a recurrent neural network models the feature sequence to capture the motion pattern of the target, and combines the attention mechanism to weight the key regions, enhancing the model's attention to the target region. At the same time, the online update strategy can fine-tune the model parameters according to the real-time tracking results, ensuring the stability and accuracy of the model in different environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] For the convenience of those skilled in the art to understand, the present invention will be further described below with reference to the drawings.
[0027] Figure 1 It is a schematic flow chart of an airport intelligent bird repelling system based on multi-object detection optimization of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will describe in detail the specific embodiments, structures, features and their effects of the present invention with reference to the drawings and preferred embodiments.
[0029] Please refer to Figure 1, an airport intelligent bird repelling system based on multi - target detection optimization, comprising: a detection unit, an identification unit, a repelling unit, and an intelligent optimization unit;
[0030] The detection unit includes a radar detection module and a visual confirmation module; the radar detection module uses a two - dimensional cylindrical phased array radar and DBF technology to obtain the three - dimensional coordinate information of all - airspace targets, sets the scanning mode and tracking mode, and performs periodic scanning on the airspace around the airport to detect the activities of birds; the tracking mode is to continuously track the detected bird targets, update their position information in real - time, and send the position information to the visual confirmation module; the visual confirmation module uses a high - definition camera and an image recognition algorithm to visually confirm the bird targets detected by the radar detection module;
[0031] The identification unit is used to identify according to the image information provided by the visual confirmation module through a target tracking model, and evaluate the potential threat level of the birds according to a preset bird threat level database to obtain an evaluation result; the repelling unit selects a bird repelling strategy to drive away the birds according to the evaluation result of the identification unit; the intelligent optimization unit optimizes the bird repelling strategy according to historical bird repelling data and real - time environmental information;
[0032] Specifically, the target tracking model uses a target detection algorithm to detect the bird species and locate the targets, uses multi - target tracking technology to track the birds, and assigns a unique ID to each bird target; adopts a deep sorting algorithm, defines a tracking scenario in a multi - dimensional state space, and combines the counting method of species, position, and ID information for area counting.
[0033] In this embodiment, the dataset used by the multi - target tracking feature extraction network is derived from the target detection dataset. We extract the complete image of each bird from the target detection dataset, including multiple angles and actions of the birds. First, a batch of image data is randomly selected from the bird dataset. Then, 4 images are randomly selected, randomly scaled, randomly distributed, and spliced into a new image, and the above operations are repeated for the number of batch processes. Finally, the Mosaic data augmentation data is used to train the neural network.
[0034] Specifically, the method for area counting is as follows: create a matrix with the same size as the image and fill it with 1 to represent the counting area; after detecting a bird, replace the value at the corresponding position with 0 and re - assign it to 1 to mark the presence of the bird; finally, count the number of birds according to the number of 1s in the matrix; for a newly - emerged bird ID, record the ID and increase the number of birds of that species; for a recorded ID, keep the number of birds of that species unchanged.
[0035] Specifically, the tracking scenario of the target tracking model is defined in a multi-dimensional state space. The dimensions of the multi-dimensional state space include the centroid coordinates of the detection frame, the aspect ratio of the image, the height of the detection frame, and their respective speeds in the image coordinates. Observation variables are constructed based on these dimensions, and Kalman filtering is used to predict the observation variables. A survival time threshold and a confirmation mechanism are set to obtain the tracking result.
[0036] In this embodiment, the tracking scenario of the depth sorting algorithm is defined in an eight-dimensional state space. In the formula, (u, v) are the centroid coordinates of the detection frame. is the aspect ratio, h is the height of the detection frame and their respective speeds in the image coordinates. Then, for the observation variables homogeneous model and linear observation model Kalman filtering are used for prediction. For each trajectory, the number of matching frames is calculated starting from the moment of the last match. If a match occurs during the prediction, the count is incremented. If the trajectory is associated with a new prediction, the count is reset to zero. In addition, a survival time threshold is set. After exceeding this threshold, if no match occurs, the target is considered to have left the tracking area and is removed from the tracking list (targets that have not been matched for a long time will be considered to have left the tracking area). Since each newly detected target may become a new trajectory or be directly classified as a trajectory, misdetection occurs from time to time. The new test result is marked as "tentative". Subsequently, if it is successfully matched for three consecutive frames, it is "confirmed" as a new trajectory. Otherwise, it is marked as "deleted" and is no longer considered a valid trajectory.
[0037] Specifically, the visual confirmation module locates the target in the image through bounding box regression and uses the intersection over union as the loss function. When calculating the intersection over union, a penalty term is introduced into the loss function according to the distance between the center points of the predicted box and the ground truth and the aspect ratio. The calculation formula is:
[0038]
[0039] Among them, α is the weight coefficient used to adjust the influence degree of the center point distance and the aspect ratio on the loss function, β is the penalty term weight of the center point distance, IoU is the intersection over union value indicating the overlapping degree between the predicted box and the ground truth box, d c is the Euclidean distance between the center points of the predicted box and the true box, w and h are the width and height of the predicted box respectively, w g and h g are the width and height of the true box respectively.
[0040] By introducing a penalty term, when the center point of the predicted bounding box is farther from the center point of the ground truth bounding box, or the aspect ratio of the predicted bounding box is more different from the aspect ratio of the ground truth bounding box, the value of the loss function becomes larger, thereby guiding the model to more accurately locate the target during the training process.
[0041] Specifically, the target tracking model is constructed using a deep learning framework, and an attention mechanism is introduced to improve the tracking accuracy. Image features are extracted through a convolutional neural network, and a recurrent neural network is used to model the feature sequence to capture the motion pattern of the target. On this basis, an attention mechanism is introduced to weight the key regions in the feature map to enhance the model's attention to the target region. The target tracking model also adopts an online update strategy to fine-tune the model parameters according to the real-time tracking results.
[0042] Specifically, the recurrent neural network consists of two parts: a Region Proposal Network (RPN) and a Fast R-CNN detector, and the two share the convolutional layer. The RPN generates region proposals through a sliding window and anchor mechanism, while the Fast R-CNN detector is responsible for classifying and regressing these region proposals. During the training process, an alternating training strategy is adopted to train the RPN and Fast R-CNN separately, and through iterative optimization, they are finally fused into a unified network.
[0043] In this embodiment, the Region Proposal Network (RPN) includes: a sliding window module composed of 3×3 convolutional kernels for extracting local features at each position on the feature map; two parallel 1×1 convolutional layers that respectively output the objectness score of the anchor box and the bounding box regression parameters; based on the preset anchor sizes and aspect ratios, at least 9 anchor boxes are generated, where the sizes include {64×64, 128×128, 256×256}, and the aspect ratios include {1:1, 1:2, 2:1}.
[0044] The alternating training strategy specifically includes:
[0045] a) Initialize the parameters of the shared convolutional layer and pre-train the RPN;
[0046] b) Freeze the shared convolutional layer and independently train the Fast R-CNN detector using the candidate regions generated by the RPN;
[0047] c) Load the Fast R-CNN parameters into the shared network and fine-tune the unique convolutional layer of the RPN;
[0048] d) Fix the parameters of the shared convolutional layer and update the fully connected layer of the Fast R-CNN;
[0049] e) Repeat steps c-d until the network converges.
[0050] The shared convolutional layer adopts the ResNet-50 or VGG-16 architecture, and its output feature map matches the input dimensions of the RPN and Fast R-CNN, reducing computational redundancy through parameter sharing.
[0051] The loss functions of the RPN and Fast R-CNN are multi-task losses, including: classification loss: calculating the objectness score error of the anchor boxes using the cross-entropy loss function;
[0052] Regression loss: calculating the coordinate offset error between the anchor boxes and the ground truth bounding boxes using the smooth L1 loss function;
[0053] The stride of the sliding window module is consistent with the downsampling rate of the shared convolutional layer, ensuring that the center points of the anchor boxes cover all spatial positions of the feature map.
[0054] Specifically, the training method of the region proposal network is as follows: Set the center point of the sliding window as the anchor box, and assign a binary class label to each anchor. The condition for the positive label is: the anchor box with the highest intersection over union ratio with a ground truth box, or the anchor box with an intersection over union ratio greater than the preset overlap threshold with any ground truth box; the condition for the negative label is: the anchor box with an intersection over union ratio less than the preset overlap threshold with all ground truth boxes. The loss function of the region proposal network includes classification loss and regression loss. The classification loss uses the cross-entropy loss function to distinguish positive and negative anchor boxes, and the regression loss uses the smooth L1 loss function to optimize the position and size of the anchor boxes.
[0055] The computer storage medium of the embodiments of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0056] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0057] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing. The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0058] As described above, the above are only the preferred embodiments of the present invention and do not impose any formal limitations on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the equivalent embodiments of equivalent changes within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modification, equivalent change, and modification made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. An airport intelligent bird-repelling system based on multi-target detection optimization, characterized in that: include: Detection unit, identification unit, expulsion unit, intelligent optimization unit; The detection unit includes a radar detection module and a visual confirmation module; The radar detection module obtains the three-dimensional coordinate information of the target in the entire airspace, and sets the scanning mode and the tracking mode to periodically scan the airspace around the airport to detect the activities of birds; the tracking mode is to continuously track the bird target after detecting it, update its position information in real time, and send the position information to the visual confirmation module; The visual confirmation module uses a high-definition camera and an image recognition algorithm to visually confirm the bird targets detected by the radar detection module; The recognition unit is used to identify the bird through the target tracking model according to the image information provided by the visual confirmation module, and evaluate the potential threat level of the bird according to the preset bird threat level database to obtain an evaluation result; the driving away unit selects a bird driving strategy to drive away the bird according to the evaluation result of the recognition unit; The intelligent optimization unit optimizes the bird-repelling strategy according to historical bird-repelling data and real-time environmental information.
2. The system according to claim 1, characterized in that The target tracking model uses a target detection algorithm to detect bird species and locate targets, uses multi-target tracking technology to track birds, and assigns a unique ID to each bird target; A deep sorting algorithm is used to define tracking scenarios on a multidimensional state space, and regional counting is performed in combination with counting methods that combine species, location, and ID information.
3. The system according to claim 2, characterized in that The method for area counting is as follows: create a matrix with the same size as the image and fill it with 1 to represent the counting area; after detecting a bird, replace the value of the corresponding position with 0 and reallocate it to 1 to mark the presence of the bird; finally, count the number of birds according to the number of 1s in the matrix; for newly appeared bird IDs, record the IDs and increase the number of birds of this species; for recorded IDs, keep the number of birds of this species unchanged.
4. The system according to claim 2, characterized in that The tracking scenario of the target tracking model is defined in a multidimensional state space, the dimensions of which include the coordinates of the centroid of the detection frame, the image aspect ratio, the detection frame height and their respective speeds in the image coordinates; observation variables are constructed according to the dimensions, the observation variables are predicted using Kalman filtering, and a lifetime threshold and a confirmation mechanism are set to obtain tracking results.
5. The system according to claim 1, characterized in that The visual confirmation module locates the target in the image through bounding box regression and uses the intersection-over-union ratio as the loss function. When calculating the intersection-over-union ratio, a penalty term is introduced into the loss function according to the distance between the predicted box and the center point of the ground truth and the aspect ratio. The calculation formula is: Among them, α is the weight coefficient, which is used to adjust the influence of the center point distance and aspect ratio on the loss function, β is the penalty term weight of the center point distance, IoU is the intersection over union ratio, which indicates the overlap between the predicted box and the ground truth box, and d c is the Euclidean distance between the center point of the predicted box and the real box, w and h are the width and height of the predicted box respectively, w g and h g are the width and height of the real box respectively; By introducing the penalty term, the farther the center point of the predicted box is from the center point of the ground truth box, or the greater the difference between the aspect ratio of the predicted box and the aspect ratio of the ground truth box, the greater the value of the loss function, thereby guiding the model to locate the target more accurately during the training process.
6. The system according to claim 1, characterized in that The target tracking model is constructed using a deep learning framework, and an attention mechanism is introduced to improve tracking accuracy. Image features are extracted through a convolutional neural network, and feature sequences are modeled using a recurrent neural network to capture the target's motion pattern. On this basis, an attention mechanism is introduced to weight key areas in the feature map to enhance the model's attention to the target area. The target tracking model also adopts an online updating strategy to fine-tune the model parameters according to the real-time tracking results.
7. The system according to claim 6, characterized in that The recurrent neural network includes a shared convolutional layer module, a region proposal network and a target detector; the output end of the shared convolutional layer module is respectively connected to the input end of the region proposal network and the input end of the target detector, and the output proposal region of the region proposal network forms a data connection channel with the ROI pooling layer of the target detector; The region proposal network includes a sliding window module and an anchor generation module, wherein the sliding window module traverses the spatial position of the feature map with a preset step size, and the anchor generation module generates at least three anchor boxes of different sizes and aspect ratios at each position, and outputs a candidate region based on the classification score and the regression offset; The object detector is connected to a shared convolutional layer, including a ROI pooling layer, a fully connected layer, and a classification regression module, for performing feature alignment, object category classification, and bounding box coordinate regression on the candidate region; The recurrent neural network optimizes network parameters through an alternating training strategy, including: first independently training a region proposal network to generate candidate regions, then training a target detector based on the candidate regions, and finally iteratively updating the shared convolutional layer parameters to achieve network fusion.
8. The system according to claim 6, characterized in that The training method of the region proposal network is as follows: the center point of the sliding window is set as an anchor point, and a binary class label is assigned to each anchor, wherein the condition of the positive label is: the anchor point with the highest intersection-and-union ratio with a basic truth box, or the anchor point whose intersection-and-union ratio with any basic truth box is greater than a preset overlap threshold; the condition of the negative label is: the anchor point whose intersection-and-union ratio with all basic truth boxes is less than a preset overlap threshold; the loss function of the region proposal network includes classification loss and regression loss, wherein the classification loss adopts a cross entropy loss function to distinguish between positive and negative anchor points; the regression loss adopts a smooth L1 loss function to optimize the position and size of the anchor point.
Citation Information
Cited By
Bird repelling device and method based on AI behavior recognition and dynamic parameters
CN121003191A
Airport bird identification and tracking method and system based on visual radar fusion
CN121686522A
Data processing method and system for following target positioning and motion control
CN122192287A