Airborne video acquisition and graded target detection system and method

By designing an onboard video acquisition and hierarchical object detection system, using a two-level processing architecture and a multiple-target box detection mechanism, the inconvenience of use caused by different power phase sequences when used in multiple sites is solved, efficient object detection and identification are achieved, and logistics work is reduced.

CN120071191APending Publication Date: 2025-05-30CHINESE FLIGHT TEST ESTAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411951635.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When existing airborne equipment is used in multiple production sites, due to the different phase sequences of the three-phase power supply, the test bench is extremely inconvenient to use, and electrical maintenance personnel are required to conduct on-site measurements and re-wires, which has a large logistics workload.

Method used

Design an on-board video acquisition and hierarchical target detection system, adopting a two-level processing architecture of the front-end camera combined with the back-end collector, implementing target detection tasks in steps, using equipment at each level to calculate resources, improving processing efficiency, and designing a detection mechanism for multiple target boxes to improve the detection accuracy of the target position.

Benefits of technology

It realizes efficient object detection and identification under the conditions of existing airborne equipment, reduces dependence on electrical maintenance personnel, and reduces logistics workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071191A_ABST
    Figure CN120071191A_ABST
Patent Text Reader

Abstract

The invention discloses an airborne video acquisition and graded target detection system and method. The detection system comprises a video imaging and processing module, a video acquisition and processing module and an output module. The video imaging and processing module is used for shooting and imaging a monitored target in an airborne application scene, performing primary target detection processing on a generated video image to obtain an initial position of the target in the image, recording the initial position as primary target detection data, and transmitting the primary target detection data to the airborne application scene; the video imaging and processing module outputs a video image containing primary target detection data to the video acquisition and processing module; the video acquisition and processing module performs second-level target detection on a video image containing first-level target detection data by adopting a CNN detection network framework based on an anchorbox and combining a target frame of a double-frame structure to obtain second-level target detection data containing target state recognition and position information; and the video acquisition and processing module is used for acquiring secondary target detection data and generating a video image for displaying the secondary target detection data in an overlapping manner, and outputting a compressed data packet to the output module after compressing and encoding the video image for displaying the secondary target detection data in the overlapping manner. The detection precision of the target position is improved, and the application requirements of target detection and recognition under the condition of existing airborne equipment are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of airborne video acquisition and processing, and particularly relates to an airborne video acquisition and hierarchical target detection system and method. Background Art

[0002] In a control loop of a certain test bench, due to the requirement for the power supply phase sequence, the designer installed a power supply phase sequence detection device at the power input end. Only when the three-phase power supply phase sequence connected to the test bench is the same as the three-phase power supply phase sequence required by the design, the indicator light on the test bench will be on, and the test bench can perform test operations. And this test bench needs to be used at multiple production sites, and the three-phase power supply phase sequences provided at each site are different, resulting in extremely inconvenient use of this test bench at multiple sites. Before use, it is necessary to first measure the three-phase power supply phase sequence of the L1, L2, and L3 phase wires of the on-site power socket, and then adjust the wiring of the L1, L2, and L3 phase wires of the test bench power input plug according to the phase sequence requirements of the test bench before the test bench can be used. And the on-site measurement of the phase sequence and on-site rewiring require electrical maintenance personnel to use special instruments to measure and implement. Due to safety reasons, ordinary users cannot and do not have the ability to perform this work. Because the test bench is frequently used at multiple sites, changing a use site requires electrical maintenance personnel to perform on-site measurement and rewiring, resulting in a large amount of logistics work. Summary of the Invention

[0003] The object of the present invention: To propose an airborne video acquisition and hierarchical target detection system and method. Based on the characteristics of the airborne video acquisition system, a two-level processing architecture combining a front-end camera and a back-end collector is designed. The target detection task is implemented step by step, and the computing resources of each level of equipment are used as much as possible to improve the processing efficiency. And a detection mechanism for multiple target frames is designed to improve the detection accuracy of the target position, so as to meet the application requirements of target detection and recognition under the existing airborne equipment conditions.

[0004] The technical solution of the present invention: To achieve the above object of the invention, according to the first aspect of the present invention, an airborne video acquisition and hierarchical target detection system is proposed, including: a video imaging and processing module, a video acquisition and processing module, and an output module;

[0005] The video imaging and processing module is used to capture images of the monitored targets in the airborne application scenario, and perform primary target detection processing on the generated video images to obtain the preliminary positions of the targets in the images, denoted as primary target detection data. The video imaging and processing module outputs the video images containing the primary target detection data to the video acquisition and processing module; the video acquisition and processing module uses a CNN detection network framework based on an anchor box and a target box with a double-frame structure to perform secondary target detection on the video images containing the primary target detection data, obtains secondary target detection data including target status recognition and position information, and generates a video image with the secondary target detection data superimposed and displayed. After the video acquisition and processing module compresses and encodes the video image with the secondary target detection data superimposed and displayed, it outputs the compressed data packet to the output module.

[0006] In a possible embodiment, the video imaging and processing module uses a CNN detection network framework based on an anchor box to perform primary target detection.

[0007] In a possible embodiment, the video acquisition and processing module receives various video signals output by multiple video imaging and processing modules on the aircraft, uses an SoC chip or a DSP chip or an FPGA to acquire data and perform format conversion to obtain a unified digital video bitstream; the video acquisition module receives the time code on the aircraft and superimposes the time code on the video screen in real time to achieve time code synchronization.

[0008] In a possible embodiment, the secondary target detection data obtained by the video acquisition and processing module is displayed in the form of a position box and a text description.

[0009] In a possible embodiment, the output module encapsulates the compressed data packet output by the video acquisition and processing module, and can be encapsulated and output in the TS format according to application requirements, or split into data packets according to the H.264 or H.265 raw stream and output, or each frame of the compressed video data can be split and packaged into a PCM data packet and output as telemetry encoded data.

[0010] According to the second aspect of the present invention, an airborne video acquisition and hierarchical target detection method is proposed. Using the above-mentioned airborne video acquisition and hierarchical target detection system, it includes the following steps:

[0011] Step 1: Capture images of the monitored targets in the airborne application scenario to obtain the original video images;

[0012] Step 2: Use a CNN detection network based on an anchor box to perform primary target detection processing on the original video images obtained in Step 1 to obtain primary target detection data;

[0013] Step 3: According to the primary target detection data in Step 2, extract the corresponding sub-images from the video image according to the positions of the target boxes. Use a CNN detection network framework based on anchor boxes combined with a target box detection network with a double-frame structure to perform secondary target detection, obtain secondary target detection data including target status recognition and position information, and superimpose and display the video image with the secondary target detection data;

[0014] Step 4: Compress and encode the video image with the superimposed secondary target detection data, and output the compressed data packets after compression encoding.

[0015] In a possible embodiment, in Step 2, the classical CNN detection network framework based on anchor boxes includes the YOLO series of networks and SSD.

[0016] In a possible embodiment, in Step 3, when extracting the corresponding sub-images from the video image according to the primary target detection data in Step 3 according to the positions of the target boxes, include the objects near the target in the sub-images according to the application scenario.

[0017] In a possible embodiment, in Step 4, encapsulate the compressed data packets, and according to the application requirements, they can be encapsulated and output in TS format or split into data packets according to the H.264 or H.265 raw stream and output. It is also possible to split each frame of the compressed video data and package it into PCM data packets and output them as telemetry encoding data.

[0018] The advantages of the present invention are as follows: Based on the characteristics of the airborne video acquisition system, a two-stage processing architecture combining a front-end camera and a back-end collector is designed. The target detection task is implemented step by step, and the computing resources of each level of equipment are used as much as possible to improve the processing efficiency. A detection mechanism with multiple target boxes is designed to improve the detection accuracy of the target position, meeting the application requirements of target detection and recognition under the existing airborne equipment conditions. Description of the Drawings

[0019] Figure 1 is a schematic structural diagram of an airborne video acquisition and hierarchical target detection system according to a preferred embodiment of the present invention;

[0020] Figure 2 is a schematic flowchart of an airborne video acquisition and hierarchical target detection method according to a preferred embodiment of the present invention;

[0021] Figure 3 is a schematic diagram of primary target detection in Step 3 according to a preferred embodiment of the present invention;

[0022] Figure 4 is a schematic diagram of secondary target detection in Step 4 according to a preferred embodiment of the present invention. Detailed implementation manners

[0023] For the purposes, technical solutions and advantages of the embodiments of the present invention to be more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present invention.

[0024] The features and illustrative embodiments of various aspects of the present invention will be described in detail below. In the following detailed description, many specific details are set forth in order to provide a comprehensive understanding of the present invention. However, it is obvious to those skilled in the art that the present invention can be implemented without some of these specific details. The following description of the embodiments is only to provide a better understanding of the present invention by showing examples of the present invention. The present invention is in no way limited to any specific settings and methods presented below, but covers any improvements, substitutions and modifications of structures, methods, and devices without departing from the spirit of the present invention. In the drawings and the following description, well-known structures and technologies are not shown to avoid unnecessarily obscuring the present invention.

[0025] It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other, and the various embodiments may refer to and cite each other. The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0026] As Figure 1 shown, an airborne video acquisition and hierarchical target detection system includes: a video imaging and processing module, a video acquisition and processing module, and an output module;

[0027] The video imaging and processing module is used to capture images of the monitored targets in the airborne application scenario and perform primary target detection processing on the generated video images to obtain the preliminary positions of the targets in the images. The video imaging and processing module is implemented by a camera integrated with a dedicated chip for deep learning model inference calculation, such as a SoC chip with an NPU from Hisilicon.

[0028] The target detection processing software adapted to the dedicated chip for deep learning model inference calculation is installed and deployed in the video imaging and processing module, and an existing classic CNN (Convolutional Neural Network) detection network framework based on anchor box is adopted, such as the yolo series network, SSD, etc. To ensure the real-time performance of the target detection processing, the target scale range in the detection network can be reduced according to the size of the targets in the application scenario, and the regression accuracy of the detection box positions can be reduced to obtain a rough target detection position.

[0029] The video imaging and processing module outputs the video images obtained by imaging to the video acquisition and processing module in accordance with the standard video data format. Common video formats include, for example, LVDS, GigE, etc.; the results of target detection are output to the video acquisition and processing module through a dedicated channel, such as network communication, or several lines can be added to the video image, and the target detection data results are placed in the added several lines and output to the video acquisition and processing module together with the video image.

[0030] The video acquisition and processing module receives various video signals output by multiple video imaging and processing modules on the machine, and uses a mature SoC chip or DSP chip or FPGA to perform data acquisition and format conversion to obtain a unified digital video bitstream. The video acquisition module on the machine receives the time code and superimposes the time code on the video picture in real time to achieve time code synchronization.

[0031] The video acquisition and processing module performs secondary target detection based on the video images received from the video imaging and processing module and the corresponding primary target detection results, obtains more accurate detection results, and superimposes the target detection and recognition results on the video image in the form of position boxes and text descriptions. According to the primary target detection results, sub-images of the video image are selected, and a detection network is implemented using a CNN detection network framework based on anchor boxes combined with a double-frame structure target box to improve the accuracy of target detection and status recognition.

[0032] The video acquisition and processing module compresses and encodes the video image stream after target detection processing, and uses a DSP chip or SoC chip to perform hard encoding in the H.264 or H.265 compression encoding format, and sends the result to the output module.

[0033] The output module encapsulates the H.264 or H.265 compressed data packets output by the video acquisition and processing module, and can be encapsulated and output in the TS format according to application requirements or split into data packets according to the H.264 or H.265 raw stream for output, or each frame of the compressed video data can be split and packaged into PCM data packets for output as telemetry encoded data.

[0034] Based on the above system, as Figure 2 shown, the present invention also proposes an airborne video acquisition and hierarchical target detection method. The flowchart and specific steps are as follows:

[0035] Step 1: Use a camera to capture and image the monitored target in the airborne application scenario to obtain the original video image frame data in RGB format;

[0036] Step 2: Perform primary target detection on the original video image obtained in Step 1. This is achieved using existing classic CNN detection network frameworks based on anchor boxes, such as the YOLO series of networks, SSD, etc. To ensure the real-time nature of target detection processing, the target scale range in the detection network can be reduced according to the size of the targets in the application scenario, and the regression accuracy of the detection box positions can be decreased to obtain a rough target detection position.

[0037] Step 3: Use a mature SoC chip, DSP chip, or FPGA to acquire and convert the format of the original video image obtained in Step 1 to obtain a unified digital video bitstream; and according to the on-board time information received, overlay the time code on the video frame in real time to achieve time code synchronization.

[0038] Step 4: According to the primary target detection results in Step 2, extract the sub-images corresponding to the primary targets from the video images of the corresponding frames in Step 3 according to the target box positions. Use a CNN detection network framework based on anchor boxes combined with a detection network with a double-frame structure for secondary target detection to obtain more accurate detection results, and overlay the target detection and recognition results on the video images in the form of position boxes and text descriptions.

[0039] Preferably, when extracting the sub-images corresponding to the primary targets from the video images of the corresponding frames in Step 3 according to the primary target detection results in Step 2, objects near the targets can be included in the sub-images according to the application scenario to facilitate the identification of the target state during secondary detection. For example Figure 3 in the case where the aircraft landing gear is the detection target, as shown by the white box according to the primary target detection results, the extended range shown by the dashed box is the sub-image, which includes the landing gear door range to facilitate the distinction of the retraction and extension states of the landing gear during secondary detection.

[0040] Step 5: Compress and encode the video images in Step 4. Preferably, use the H.264 or H.265 compression encoding format to achieve this.

[0041] Step 6: Package the H.264 or H.265 compressed data packets output in Step 5. According to the application requirements, it can be packaged and output in TS format or split into data packets according to the H.264 or H.265 raw stream for output. It can also split each frame of the compressed video data and package it into PCM data packets for output as telemetry encoded data.

[0042] Among them, in Step 4, the schematic diagram of extracting the sub-images corresponding to the primary targets from the video images of the corresponding frames according to the primary detection target box positions is as follows Figure 4As shown. A detection network for object boxes using an anchor box-based CNN detection network framework combined with a double-box structure is employed for secondary object detection. The double box used is set as shown in Figure 2 and consists of two nested rectangular boxes as shown therein. The small box contains the object, such as the landing gear, and the large box contains the surrounding objects that can reflect the change in the state of the object, such as the landing gear door. The definition of the object box and the training loss function of the anchor box-based CNN detection network includes the following steps:

[0043] 1) Set the definition of the double box in the video image as: {(x b , y b ), (x s , y s ), (w b , h b ), (w s , h s )}, where x b , y b , x s , y s , w b , h b , w s , h s are the position parameters describing the object box respectively. That is, (x b , y b ) is the coordinate value of the upper left origin of the large rectangular box, (x s , y s ) is the coordinate value of the upper left origin of the small rectangular box, w b , h b are the width and height of the large rectangular box respectively, and w s , h s are the width and height of the small rectangular box respectively;

[0044] 2) The detection network structure adopts an existing classic detection network framework based on anchor box, such as the yolo series network, SSD, etc. Then, based on the above definition of the double object box, regression processing is performed on both boxes simultaneously. The specific embodiments are as follows:

[0045] Set the ground truth box (gt box) and define it as where, is the origin coordinate of the large box in the ground truth box, is the upper left origin coordinate of the small box in the ground truth box, are the width and height of the large box in the ground truth box respectively, are the width and height of the small box in the ground truth box respectively;

[0046] Set a prediction box (predict box, pr box) and define it as where is the origin coordinate of the large box in the prediction box, are the origin coordinates of the small boxes in the prediction box respectively, are the width and height of the large box in the prediction box respectively, are the width and height of the small boxes in the prediction box respectively;

[0047] Set an anchor box (anchor box, ac box) and define it as where, (a x , a y ) is the origin coordinate of the large box in the anchor box, are the origin coordinates of the small boxes in the anchor box respectively, is the width and height of the large box in the anchor box, is the width and height of the small box in the anchor box;

[0048] 3) The relationships between the prediction box and the anchor box, and between the ground truth box and the anchor box are determined by the following calculation formulas:

[0049]

[0050] where, p l is the prediction value of the prediction box;

[0051]

[0052] g l is the prediction value of the ground truth box;

[0053] 4) The loss function is obtained through the following calculation formula:

[0054]

[0055] where, L cls is the classification loss function, which is defined as four categories: landing gear down state, landing gear up state, landing gear extending / retracting state, and negative samples. The negative samples can be selected as the part without landing gear, and a common classification loss function can be used, such as the softmax loss function;

[0056] L reg is the regression loss function for the target box position;

[0057] gc ∈ {0, 1, 2, 3} is the category of the ground truth box. When gc = 0, it means that the ground truth box is not a landing gear; when gc = 1, 2, 3, it means that the ground truth box is in the landing gear down state, up state, and extending / retracting state respectively;

[0058] pc is the probability value predicted by the prediction box for the landing gear and its status; N is the number of landing gear target boxes; α is the weight coefficient; M is the number of ground truth boxes; θ k represents the target box position parameter x b , y b , x s , y s , w b , h b , w s , h s The weight coefficient of, can appropriately reduce w according to the actual effect b , h b weight, to make the positioning accuracy of the landing gear higher; smooth L1 is the smoothing function;

[0059] 5) When the detection network performs inference calculation, it is inversely calculated from the predicted value pl of the prediction box according to formula (3). That is to say, the trained network will calculate the predicted value pl. According to the 8 sub-formulas in (3), the 8 position parameters of the target box can be calculated as x b , y b , x s , y s , w b , h b , w s , h s , that is, the position and size of the predicted target box in the image are obtained, so as to realize the detection of the accurate position of the landing gear and the recognition of its status.

Claims

1. An airborne video acquisition and hierarchical target detection system, characterized in that: include: Video imaging and processing module, video acquisition and processing module, output module; The video imaging and processing module is used to shoot and image the monitored target in the airborne application scenario, and perform primary target detection processing on the generated video image to obtain the preliminary position of the target in the image, which is recorded as primary target detection data. The video imaging and processing module outputs the video image containing the primary target detection data to the video acquisition and processing module; the video acquisition and processing module uses a CNN detection network framework based on anchorbox combined with a target frame with a double frame structure to perform secondary target detection on the video image containing the primary target detection data, obtains secondary target detection data containing target state recognition and position information, and generates a video image with superimposed display of the secondary target detection data. After the video acquisition and processing module compresses and encodes the video image with superimposed display of the secondary target detection data, the compressed data packet is output to the output module.

2. The airborne video acquisition and hierarchical target detection system according to claim 1, characterized in that: The video imaging and processing module uses a CNN detection network framework based on anchorbox to perform primary target detection.

3. The airborne video acquisition and hierarchical target detection system according to claim 1, characterized in that: The video acquisition and processing module receiver receives various video signals output by the multiple video imaging and processing modules, and uses a SoC chip, a DSP chip, or an FPGA to acquire data and convert formats to obtain a unified digital video code stream; the video acquisition module receiver receives the time code, and superimposes the time code on the video screen in real time to achieve time code synchronization.

4. The airborne video acquisition and hierarchical target detection system according to claim 1, characterized in that: The secondary target detection data obtained by the video acquisition and processing module is displayed in the form of a position frame and text description.

5. The airborne video acquisition and hierarchical target detection system according to claim 1, characterized in that: The output module encapsulates the compressed data packets output by the video acquisition and processing module. According to application requirements, it can be encapsulated into TS format for output or split into data packets according to H.264 or H.265 raw stream for output. It can also split each frame of the compressed video data and package them into PCM data packets for output as telemetry coded data.

6. An airborne video acquisition and hierarchical target detection method, using an airborne video acquisition and hierarchical target detection system according to any one of claims 1 to 5, characterized in that: The steps include: Step 1: photograph and image the monitored target in the airborne application scenario to obtain an original video image; Step 2: Use the anchorbox-based CNN detection network to perform primary target detection processing on the original video image obtained in step 1 to obtain primary target detection data; Step 3: Based on the primary target detection data in step 2, the corresponding sub-image is extracted from the video image according to the position of the target frame, and the secondary target detection is performed by using the CNN detection network framework based on the anchor box combined with the target frame detection network with a double frame structure, so as to obtain the secondary target detection data containing the target state recognition and position information, and the video image of the secondary target detection data is superimposed and displayed; Step 4: compress and encode the video image superimposed with the secondary target detection data, and output the compressed data packet after compression and encoding.

7. The method for airborne video acquisition and hierarchical target detection according to claim 6, characterized in that: In the step 2, the classic CNN detection network framework based on the anchor box includes the YOLO series network and SSD.

8. The method for airborne video acquisition and hierarchical target detection according to claim 6, characterized in that: In the step three, according to the primary target detection data in step three, when extracting the corresponding sub-image from the video image according to the position of the target frame, objects near the target are also included in the sub-image according to the application scenario.

9. The method for airborne video acquisition and hierarchical target detection according to claim 6, characterized in that: In step 4, the compressed data packet is encapsulated and output in TS format according to application requirements, or split into data packets according to H.264 or H.265 raw stream for output. Alternatively, each frame of the compressed video data can be split and packaged into PCM data packets and output as telemetry coded data.