A bionic large-field-of-view cross-aperture collaborative monitoring method, model, computer program product and terminal device
Through the bionic large field of view cross-aperture collaborative monitoring method, the bionic curved surface compound eye imaging system and deep learning network are used to solve the problem of missed detection and failure during movement of wide-area monitoring equipment, and efficient object detection and trajectory tracking are achieved.
Patent Information
- Application Number
- CN202411878995.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-19
AI Technical Summary
Existing wide-area monitoring equipment is prone to missed detection and equipment failure during movement, which makes it difficult to capture, detect, identify and track targets within a large field of view.
The biological large field of view cross-aperture collaborative monitoring method is adopted to obtain the original compound eye image through the bionic curved surface compound eye imaging system, image preprocessing and model training are performed to generate a large field of view cross-aperture collaborative monitoring model. The model utilizes the overlap of field of view of multiple sub-eyes to perform object detection and trajectory tracking through a deep learning network.
It realizes monitoring of a larger field of view, and at the same time, the image distortion is smaller, which improves the reliability and accuracy of object detection. It can monitor multiple targets in the wide area in real time, suitable for real-time monitoring tasks in the wide area.
Smart Images

Figure CN119323672B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a target monitoring method and a model, and in particular to a bionic large-field-of-view cross-aperture collaborative monitoring method, a model, a computer program product and a terminal device. Background Art
[0002] The visual system is an important guarantee for the survival of organisms in nature. Due to different living environments and habits, there are great differences in the visual systems of organisms. The single-aperture eyes of some vertebrates such as humans have a larger aperture and higher spatial resolution, but correspondingly, their field of view is smaller, so they need to turn their heads to observe targets in a wide range. The compound eye visual system of some insects and mantis shrimps has a larger field of view, less distortion, and is more sensitive to moving targets. Based on the above-mentioned advantages of compound eyes, the bionic curved compound eye imaging system that imitates the design of biological compound eyes can achieve accurate capture of moving targets in a wide range.
[0003] At present, single-aperture systems are widely used in optical imaging systems in the field of wide-area monitoring. The spatial resolution of their imaging is high, but the field of view is extremely limited. In order to achieve wide-area monitoring, moving parts are needed to expand the monitoring field of view. Therefore, the single-aperture system may miss the monitored target during movement. At the same time, there is a possibility of failure of its moving parts, which further reduces the reliability of wide-area monitoring equipment. In addition, fisheye lenses are often used in the aerospace field to expand the field of view, but the distortion on all sides is extremely large, and extremely complex image processing algorithms are required for distortion correction. Therefore, it is difficult for fisheye lenses to achieve real-time target recognition, and their application in wide-area monitoring is greatly limited. Summary of the invention
[0004] The purpose of the present invention is to solve the shortcomings of the prior art that wide-area monitoring equipment may miss detections and fail during movement, making it difficult to capture, detect, identify and track targets within a large field of view, and to provide a bionic large-field-of-view cross-aperture collaborative monitoring method, model, computer program product and terminal device.
[0005] To achieve the above objectives, the technical solutions provided by the present invention are as follows:
[0006] A bionic large-field-of-view cross-aperture collaborative monitoring method is characterized in that it comprises the following steps:
[0007] Step 1: Model training phase
[0008] Step 1.1, obtaining original compound eye images of multiple scenes and multiple targets;
[0009] Step 1.2, the original compound eye image is subjected to image preprocessing to obtain each sub-eye image;
[0010] Step 1.3, divide each sub-eye image into a training set, a validation set, and a test set, and perform target labeling on the sub-eye images of the training set and the validation set;
[0011] Step 1.4, load the deep learning network model for target detection, use the training set, validation set, and test set to perform model training, model validation, and model testing respectively, and then output a large field of view cross-aperture collaborative monitoring model;
[0012] Step 2: Target monitoring phase
[0013] Step 2.1, obtain the original compound eye image of the target, perform image preprocessing to obtain multiple sub-eye images, and save the contour radius and center position of each sub-eye; the same target will appear in multiple different sub-eye images;
[0014] Step 2.2, using the large field of view cross-aperture collaborative monitoring model to identify the predicted frame of the target in each sub-eye image and the corresponding confidence;
[0015] Step 2.3, reconstructing the sub-eye image into a large field of view image according to the contour radius and center position of each sub-eye;
[0016] In step 2.4, the average of multiple prediction boxes of the same target is taken and the output is the prediction box of the target in the large field of view image; the collaborative confidence is calculated based on the multiple confidences of the same target.
[0017] Furthermore, in step 2.4, the collaborative confidence is calculated based on multiple confidences of the same target, specifically by the following formula:
[0018]
[0019]
[0020] in, Predict is the collaborative confidence, m and b are the collaborative confidence Predict The adjustment factor is used to make the collaborative confidence Predict The exponential function, α i is the confidence corresponding to the i-th sub-eye image, k i is the influence weight of the confidence corresponding to the i-th sub-eye image on the collaborative confidence, θ i is the angle between the sub-eye optical axis and the central optical axis corresponding to the i-th sub-eye image, i is 1, 2, …, n, n is the number of sub-eyes with the same target, p i is the adjustment factor that affects the weight, by adjusting p iReduce the impact of sub-eyes far away from the central optical axis on the coordination confidence.
[0021] Furthermore, in step 2.2, the prediction box includes parameters of the prediction box center position and the prediction box size;
[0022] Also includes step 3:
[0023] Acquire multiple frames of original compound eye images of the target, obtain the center position of the prediction box corresponding to each frame of the original compound eye image through step 2, and obtain the tracking trajectory of the target.
[0024] Furthermore, in step 1.3, the ratio of sub-eye images in the training set, validation set, and test set is 6:2:2.
[0025] At the same time, the present invention also provides a bionic large field of view cross-aperture collaborative monitoring model, which is used to realize the above-mentioned bionic large field of view cross-aperture collaborative monitoring method, and its special features are: it includes a backbone network, a neck module, and a detection head module; the backbone network is used to extract each level of features of the input image by downsampling; the input image is a sub-eye image captured by a bionic curved compound eye imaging system; the neck module is used to perform feature fusion on some level features in each level of features, and some level features include features that are smaller than 1 / 2 of the input image pixel after downsampling and greater than a set threshold, and the features are used to predict small targets; the detection head module includes detection heads corresponding to some level features, and the detection heads are used to predict targets respectively according to the level features after feature fusion.
[0026] Furthermore, it also includes a convolution layer Conv connected to the input end of the backbone network, and the convolution layer Conv is used to adjust the feature size and dimension of the input image;
[0027] The feature size and dimension of the input image are 320×320×1, and after being processed by the convolution layer Conv, the feature size and dimension are adjusted to 640×640×3;
[0028] The neck module includes multiple groups of Concat modules and C2F modules, which are used to fuse features of each level of features, whose feature sizes and dimensions are 20×20×1024, 40×40×512, 80×80×256, and 160×160×128 respectively;
[0029] The threshold is set to 80 pixels.
[0030] The present invention also provides a computer program product, including a computer program, which is special in that when the program is executed by a processor, the steps of the above-mentioned bionic large-field-of-view cross-aperture collaborative monitoring method are implemented.
[0031] The present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the terminal device is special in that when the processor executes the computer program, the steps of the above-mentioned bionic large field of view cross-aperture collaborative monitoring method are implemented.
[0032] Beneficial effects of the present invention:
[0033] 1. A bionic large-field-of-view cross-aperture collaborative monitoring method of the present invention is based on the original compound eye image obtained by the bionic curved compound eye imaging system, which achieves a larger field of view while reducing image distortion. The curved bionic compound eye imaging system is used to enable multiple sub-eyes to image the object space. There is an overlap in the field of view between adjacent sub-eyes. The same target in the scene can be observed by multiple sub-eyes at the same time, and multiple sub-eye images of the same target can be obtained simultaneously in one imaging.
[0034] 2. In a bionic large-field-of-view cross-aperture collaborative monitoring method of the present invention, the images used for model training and model verification are collected by the bionic curved compound eye imaging system itself, which is more suitable for the prediction of targets captured by the compound eye than the public data sets used in existing target detection algorithms; and the data sets required for model training and model verification are automatically labeled with target categories, which improves the reliability and accuracy of target prediction in actual applications.
[0035] 3. In a bionic large-field-of-view cross-aperture collaborative monitoring method of the present invention, the same target may appear in multiple sub-eye images. Multiple confidences are output after target prediction. Different weights are comprehensively assigned to the multiple confidences according to the different angles of the corresponding sub-eye optical axes. After the multiple confidences are collaboratively weighted and normalized, all the confidences are unified, and the output collaborative confidence is more reliable and more accurate.
[0036] 4. In a bionic large-field-of-view cross-aperture collaborative monitoring method of the present invention, the bionic curved compound eye imaging system can continuously capture multiple frames of images of the same target, and then continuously perform large-field-of-view image reconstruction and confidence collaborative weighted normalization based on the multiple frames of images, and continuously record and output the center position of the prediction frame output by the multiple frames of images, thereby realizing fast and convenient trajectory tracking of specific targets.
[0037] 5. The image preprocessing process in the bionic large-field-of-view cross-aperture collaborative monitoring method of the present invention is simpler, the large-field-of-view cross-aperture collaborative monitoring model used is lighter and smaller, the network model parameters are fewer, and near-real-time target detection can be achieved on the host. The next step is to deploy it in actual monitoring equipment to achieve real-time monitoring of multiple targets in a wide area.
[0038] 6. The bionic large-field-of-view cross-aperture collaborative monitoring model of the present invention is based on the existing deep learning network model, and adds a small target detection module for small targets of compound eyes. Since the input image is a sub-eye image, small targets occupy more pixels, and thus the small target detection module recognizes features more accurately. As a result, the large-field-of-view cross-aperture collaborative monitoring network model has higher detection efficiency and accuracy for capturing small targets by compound eyes.
[0039] 7. The present invention also provides a computer program product and a terminal device capable of executing the above method steps, which can promote and apply the method of the present invention and realize monitoring on corresponding hardware devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flow chart of an embodiment of the bionic large field of view cross-aperture collaborative monitoring method of the present invention;
[0041] Figure 2 It is a structural schematic diagram of an embodiment of a bionic large-field-of-view cross-aperture collaborative monitoring model of the present invention;
[0042] Figure 3 It is the sub-eye image corresponding to the same target in step 2.3 of the embodiment of the bionic large field of view cross-aperture collaborative monitoring method of the present invention;
[0043] Figure 4 It is the large field of view image reconstructed in step 2.3 of the embodiment of the bionic large field of view cross-aperture collaborative monitoring method of the present invention;
[0044] Figure 5 2 is a schematic diagram of the structure of a bionic curved compound eye imaging system in an embodiment of a bionic large-field-of-view cross-aperture collaborative monitoring method of the present invention;
[0045] Figure 6 is the original compound eye image in the embodiment of the bionic large-field-of-view cross-aperture collaborative monitoring method of the present invention;
[0046] Figure 7 yes Figure 6 The enlarged schematic diagram of point A in the middle;
[0047] Figure 8 yes Figure 7 The enlarged schematic diagram of point B in the middle;
[0048] Fig. 9 yes Figure 8 Enlarged schematic diagram of point C in the middle.
[0049] Description of reference numerals:
[0050] 1-ball cover assembly, 11-double cemented lens, 12-curved spherical shell, 2-relay lens assembly, 3-focal plane camera assembly. DETAILED DESCRIPTION
[0051] like Figure 1 As shown, the bionic large-field-of-view cross-aperture collaborative monitoring method of the present invention is mainly divided into two stages: a model training stage and a target monitoring stage.
[0052] The model training phase is mainly responsible for training, verifying and testing the large field of view cross-aperture collaborative monitoring model. It includes the following steps:
[0053] 1.1. Obtain the original compound eye image as a training image. In this embodiment, the target is a human target. The original compound eye image is captured multiple times by the bionic curved compound eye imaging system for multiple scenes and multiple targets.
[0054] 1.2. Perform image preprocessing on the original compound eye image to obtain the corresponding sub-eye image; image preprocessing mainly includes four steps: circle detection, positioning, segmentation, and correction;
[0055] Circle detection is to detect the circular sub-eye contour of the original compound eye image; positioning refers to drawing the sub-eye edge for the detected circular sub-eye contour, and saving the corresponding contour radius, circle center position and other parameters. In this embodiment, the Hough circle detection algorithm is used to implement the above circle detection and positioning. In other embodiments of the present invention, Sobel operator detection or canny contour detection algorithm can also be used. Segmentation refers to segmenting the original compound eye image according to the corresponding contour radius and circle center position parameters saved, that is, cutting the corresponding sub-eye image; correction refers to distortion correction of the segmented sub-eye image.
[0056] 1.3. The sub-eye images obtained in step 1.2 are divided into a training set, a validation set, and a test set according to a ratio of 6:2:2. In other embodiments of the present invention, the sub-eye images may also be classified according to other ratios;
[0057] The sub-eye images in the training set and the validation set are annotated with target anchor boxes, which are converted into text format and saved as label files.
[0058] 1.4. Load the deep learning network model for target detection, use the training set to train the model, and generate a pre-model; use the validation set to verify the model, and generate the best model containing the optimal weight file, that is, the large field of view cross-aperture collaborative monitoring model; after using the test set to test the model, output the large field of view cross-aperture collaborative monitoring model.
[0059] like Figure 2 The figure shows the large field of view cross-aperture collaborative monitoring model of the present invention, the main body of which is the YOLOv8s model. A small target detection module is added on this basis, which mainly includes three parts: backbone network (Backbone), neck module (Neck module), and detection head module (Head module).
[0060] In order to improve the target prediction accuracy of small targets, this embodiment further sets a convolution layer Conv at the input end of the backbone network. After the convolution layer Conv processing, the feature size and dimension of the input image are adjusted from 320×320×1 to 640×640×3. The input image is a sub-eye image of a training set, a validation set, or a test set.
[0061] After the input image is dimensionally adjusted by the convolution layer Conv, it enters the backbone network. The backbone network extracts the features of each level through downsampling. The three-layer features with feature sizes and dimensions of 20×20×1024, 40×40×512, and 80×80×256 are selected from the features of each level and enter the neck module. The neck module includes multiple groups of Concat modules and C2F modules. Among them, the Concat module captures more image feature information by splicing and combining features of different levels, and realizes the fusion of features of different levels. The C2F module is used to split and enhance features of different levels. The split part of the level features is then fused with the level features extracted by the backbone network through the Concat module. The remaining part of the level features can capture more complex detail information through further convolution processing, thereby enhancing the features. The combined use of the Concat module and the C2F module can effectively improve the target detection performance and accuracy of the large field of view cross-aperture collaborative monitoring model. The above three layers of features are fused by a set of Concat modules and C2F modules respectively, and then input into three detection heads of different sizes, which are respectively recorded as Detect1, Detect2, and Detect3. At the same time, the features of each level with a feature size and dimension of 160×160×128 are fused by a set of Concat modules and C2F modules respectively, and then input into the small target detection head (Small Object Detect) which is specially used to detect small compound eye targets. The three detection heads of different sizes and the small target detection head constitute the detection head module, which performs target prediction according to the corresponding input features. Among them, the convolution layer Conv, the set of Concat modules and C2F modules corresponding to the features with a feature size and dimension of 160×160×128, and the small target detection head constitute the small target detection module.
[0062] In this embodiment, the input image is a grayscale image of 320×320px. In other embodiments of the present invention, other YOLO series models, or target detection models based on Fast-CNN, or other deep learning network models for target detection can also be used to select the adapted input image, and at the same time add a small target detection module to the corresponding model to perform target prediction based on features that are smaller than 1 / 2 of the input image pixels and larger than 80 pixels after downsampling.
[0063] During the target prediction process of the detection head module, the large field of view cross-aperture collaborative monitoring model learns image features through training sets and validation sets, generates weight parameters of each level of the large field of view cross-aperture collaborative monitoring model, and finally outputs the optimal weight file, which is loaded into the large field of view cross-aperture collaborative monitoring model. The large field of view cross-aperture collaborative monitoring model is used to predict targets for sub-eye images and output the target prediction box and the corresponding confidence. The prediction box includes parameters of the center position and size of the prediction box.
[0064] like Figure 1 As shown in the figure, after the large field of view cross-aperture collaborative monitoring model is determined, the target monitoring stage is entered to realize the detection, recognition and tracking of the target on the original compound eye image actually captured. Specifically, the following steps are included:
[0065] 2.1. Obtain the original compound eye image of the target as the monitoring image, perform image preprocessing on the original compound eye image to obtain multiple sub-eye images, and input them into the large field of view cross-aperture collaborative monitoring model. The original compound eye image is captured by the bionic curved compound eye imaging system.
[0066] 2.2. The large field of view cross-aperture collaborative monitoring model recognizes the predicted frame and corresponding confidence of the target in each sub-eye image.
[0067] 2.3. Due to the overlapping fields of view of the bionic curved compound eye imaging system, the same target will appear in multiple sub-eye images, that is, there is sub-aperture cross detection. Therefore, the sub-eye images of the same target will output multiple prediction boxes and corresponding confidences. Multiple prediction boxes correspond to multiple sub-eye images. According to the corresponding contour radius, center position and other parameters of the circular sub-eye saved in the image preprocessing, the sub-eye images are merged and reconstructed into a large field of view image, such as Figure 3 and Figure 4 As shown, the large field of view image reconstruction process is the inverse process of the large field of view imaging process.
[0068] 2.4. Output the predicted box and co-confidence of each target Predict .
[0069] The average of multiple prediction boxes corresponding to the same target is taken, and the output is the prediction box of the target in the reconstructed large field of view image.
[0070] Since there is a certain angle between the optical axes of each sub-eye of the bionic curved compound eye imaging system, and the closer to the central optical axis, the higher the spatial resolution of the sub-eye, and the higher the confidence of the corresponding sub-eye image prediction, therefore, the influence of the sub-eye optical axis on target prediction is comprehensively considered, and the multiple confidences of the same target in multiple sub-eye images are collaboratively weighted and normalized, and finally the only collaborative confidence in the large field of view image is output.
[0071] The specific configuration of confidence co-weighted normalization is: when the same target appears in n Within the field of view of each eye, the target can be captured n The confidence corresponding to each sub-eye image is recorded as α 1 , α 2 ,……, α n , the angles between the optical axis of each sub-eye and the central optical axis are recorded as θ 1 , θ 2 ,……, θ n , the influence weight of the confidence corresponding to each sub-eye image on the collaborative confidence is recorded as k 1 , k 2 ,……, k n , the angle between the sub-eye optical axis corresponding to the i-th sub-eye image and the central optical axis θ i , the influence weight of confidence on collaborative confidence k i The following relationship is satisfied, and the value of i is 1, 2, ..., n.
[0072]
[0073] in, p i is the adjustment factor that affects the weight, by adjusting p i Reduce the impact of sub-eyes far away from the central optical axis on the coordination confidence.
[0074] Then, the collaborative confidence Predict for:
[0075]
[0076] Among them, m and b are the collaborative confidence Predict The adjustment factor is to adjust m and b to make the collaborative confidence Predict The exponential function, α i is the confidence corresponding to the i-th sub-eye image.
[0077] For a single frame of original compound eye image, the final output is the prediction box of multiple targets in the reconstructed large field of view image and its collaborative confidence. For the captured multiple frames of original compound eye images, after each frame of original compound eye image is input into the large field of view cross-aperture collaborative monitoring model, the center position of the prediction box corresponding to each frame of original compound eye image is recorded through the above steps 2.1 to 2.4 to obtain the tracking trajectory of the target, thereby realizing the trajectory tracking of the target.
[0078] The present invention can use the existing bionic curved compound eye imaging system to capture the original compound eye image, such as Figure 5 As shown, the bionic curved compound eye imaging system mainly consists of three parts: a ball cover assembly 1, a relay lens assembly 2, and a focal plane camera assembly 3.
[0079] The ball cover assembly 1 includes a plurality of double-cemented lenses 11 and a curved spherical shell 12. The plurality of double-cemented lenses 11 are arranged on the curved spherical shell 12 according to a certain rule to form a compound eye lens. In this embodiment, the plurality of double-cemented lenses 11 are arranged in a hexagonal honeycomb structure as a plurality of sub-eyes. There is a certain gap between the double-cemented lenses 11 to prevent aliasing of the corresponding sub-eye images. There is a certain angle between the optical axis of the double-cemented lens 11 and the main optical axis of the compound eye lens. In this embodiment, the optical axis angle between two adjacent double-cemented lenses 11 is 7°. In other embodiments of the present invention, the plurality of double-cemented lenses 11 may also be arranged from the inside to the outside in a concentric circle structure, or arranged in other ways.
[0080] The relay lens assembly 2 is used to convert the primary curved surface image of the double cemented lens 11 into a secondary plane image for reception by the focal plane camera assembly 3. The relay lens assembly 2 is composed of a plurality of lenses and can also realize the correction of object-side aberration. In this embodiment, the relay lens assembly 2 includes 10 lenses, and specifically, the structure disclosed in the Chinese patent with publication number CN116485676A can be used. The focal plane camera assembly 3 includes a planar large-area array image sensor and an imaging circuit, and the planar large-area array image sensor is used to receive the secondary plane image formed by the relay lens assembly 2.
[0081] like Figures 6 to 9 As shown, in this embodiment, the original compound eye image captured by the bionic curved compound eye imaging system is also arranged in a hexagonal honeycomb structure, and there is a certain interval between the sub-eye images, so no aliasing occurs, but there is sub-aperture cross detection, so multiple sub-eyes can detect the same target at the same time. The size of the same target is related to the imaging distance. Among them, the number of pixels occupied by an adult male target at 100 m is about 22×10 px, which can meet the detection requirements of the large field of view cross-aperture collaborative monitoring model.
[0082] The above-mentioned bionic curved compound eye imaging system is designed and manufactured to imitate the design and manufacture of insect compound eyes, achieving large field of view imaging with minimal distortion. The same target can be captured by multiple sub-eyes at the same time. Combined with the bionic large field of view cross-aperture collaborative monitoring model, it can realize detection and recognition of multiple categories of targets through model training, and can track the trajectory of specific targets based on target prediction results, thereby meeting the search, detection, recognition and tracking tasks of specific targets in a wide area.
[0083] The bionic large field of view cross-aperture collaborative monitoring method of the present invention can also form a computer program product, which includes a computer program, and when the program is executed by a processor, the steps of the above-mentioned bionic large field of view cross-aperture collaborative monitoring method are implemented. In addition, the monitoring method of the present invention can also be applied to a terminal device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of the monitoring method of the present invention are implemented. The terminal device here can be a computer, a notebook, a PDA, and various cloud servers and other computing devices, and the processor can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit or other programmable logic device, etc.
Claims
1. A bionic large field of view cross-aperture collaborative monitoring method, characterized in that: The following steps are involved: Step 1: Model training phase Step 1.1, obtaining original compound eye images of multiple scenes and multiple targets; Step 1.2, the original compound eye image is subjected to image preprocessing to obtain each sub-eye image; Step 1.3, divide each sub-eye image into a training set, a validation set, and a test set, and perform target labeling on the sub-eye images of the training set and the validation set; Step 1.4, load the deep learning network model for target detection, use the training set, validation set, and test set to perform model training, model validation, and model testing respectively, and then output a large field of view cross-aperture collaborative monitoring model; Step 2: Target monitoring phase Step 2.1, obtain the original compound eye image of the target, perform image preprocessing to obtain multiple sub-eye images, and save the contour radius and center position of each sub-eye; the same target will appear in multiple different sub-eye images; Step 2.2, using the large field of view cross-aperture collaborative monitoring model to identify the predicted frame of the target in each sub-eye image and the corresponding confidence; Step 2.3, reconstructing the sub-eye image into a large field of view image according to the contour radius and center position of each sub-eye; In step 2.4, the average of multiple prediction boxes of the same target is taken and the output is the prediction box of the target in the large field of view image; the collaborative confidence is calculated based on multiple confidences of the same target by the following formula: k i =(1-p i ·i i ) Among them, Predict is the collaborative confidence, m and b are the adjustment factors of the collaborative confidence Predict, which are used to make the collaborative confidence Predict tend to the exponential function, α i is the confidence level corresponding to the i-th sub-eye image, k i is the influence weight of the confidence corresponding to the i-th sub-eye image on the collaborative confidence, θ i is the angle between the sub-eye optical axis corresponding to the i-th sub-eye image and the central optical axis. The value of i is 1, 2, …, n, where n is the number of sub-eyes with the same target. i is the adjustment factor that affects the weight, by adjusting p i Reduce the impact of sub-eyes far away from the central optical axis on the coordination confidence.
2. According to claim 1, a bionic large field of view cross-aperture collaborative monitoring method is characterized in that: In step 2.2, the prediction box includes parameters of the prediction box center position and the prediction box size; Also includes step 3: Acquire multiple frames of original compound eye images of the target, obtain the center position of the prediction box corresponding to each frame of the original compound eye image through step 2, and obtain the tracking trajectory of the target.
3. According to claim 2, a bionic large field of view cross-aperture collaborative monitoring method is characterized by: In step 1.3, the ratio of sub-eye images in the training set, validation set, and test set is 6:2:
2.
4. A bionic large field of view cross-aperture collaborative monitoring model, used to implement a bionic large field of view cross-aperture collaborative monitoring method according to any one of claims 1 to 3, characterized in that: Including backbone network, neck module, and detection head module; The backbone network is used to extract features of each level of an input image by downsampling; the input image is a sub-eye image captured by a bionic curved compound eye imaging system; The neck module is used to perform feature fusion on some of the hierarchical features in each hierarchical feature, wherein the some hierarchical features include features that are smaller than 1 / 2 of the input image pixel after downsampling and larger than a set threshold, and the features are used to perform target prediction on small targets; The detection head module includes detection heads corresponding to some hierarchical features respectively, and the detection heads are used to perform target prediction respectively according to the hierarchical features after feature fusion.
5. According to claim 4, a bionic large field of view cross-aperture collaborative monitoring model is characterized by: It also includes a convolution layer Conv connected to the input end of the backbone network, and the convolution layer Conv is used to adjust the feature size and dimension of the input image; The feature size and dimension of the input image are 320×320×1, and after being processed by the convolution layer Conv, the feature size and dimension are adjusted to 640×640×3; The neck module includes multiple groups of Concat modules and C2F modules, which are used to fuse features of each level of features, whose feature sizes and dimensions are 20×20×1024, 40×40×512, 80×80×256, and 160×160×128 respectively; The threshold is set to 80 pixels.
6. A computer program product, comprising a computer program, characterized in that: When the program is executed by a processor, the steps of a bionic large-field-of-view cross-aperture collaborative monitoring method described in any one of claims 1 to 3 are implemented.
7. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the bionic large-field-of-view cross-aperture collaborative monitoring method as described in any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Multi-dimensional information processing method and sensing system based on curved surface bionic compound eye
CN116485676A
Freight train foreign matter detection method and system based on YOLOv8
CN117557860A