A target tracking and detection method, detection device, program product, and storage medium for an optoelectronic pod

By detecting and tracking video images in the photoelectric pod of the drone, obtaining positive and negative samples, training the target detection update model, learning the changing characteristics of the target in real time, and determining the correct target through confidence scores, the problem of difficulty in accurately tracking and easily losing targets due to target changes is solved, and the stability and reliability of target tracking detection are improved.

CN119672324BActive Publication Date: 2025-07-01TAIZHOU HONGYUE OPTOELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510191717.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-01
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

In the photoelectric pod of drones, the target changes greatly, making it difficult to accurately track and detect the target as soon as possible, and it is easy to lose the target.

Method used

By detecting and tracking objects in video images, obtaining object detection results and tracking results, using these results to obtain positive and negative samples, training the object detection update model, learning the changing characteristics of the target in real time, and determining the correct target through confidence scores.

Benefits of technology

It improves the accuracy of tracking and detection of targets by autonomous or semi-autonomous aircraft in complex and changing environments, solves the problem of difficulty in accurately tracking and easily losing targets due to target changes, and enhances the stability and reliability of target tracking and detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672324B_ABST
    Figure CN119672324B_ABST
Patent Text Reader

Abstract

The present application provides a target tracking and detection method, a detection device, a program product, and a storage medium for an optoelectronic pod. Tracking is performed while detecting the target, shortening the delay time between detection and tracking. After obtaining the target detection result and the target tracking result, positive and negative samples are obtained based on them and the video image. This method can accurately collect target and non-target data, making the training samples more targeted and diverse. Then, the positive and negative samples are used to train the target detection update model, enabling the model to learn the changing characteristics of the target in real time and output more accurate feature data and target models for the detection and tracking of new video images. Finally, the correct target is determined through confidence scoring. Through the interaction of these technical features, the tracking and detection accuracy of the autonomous or semi-autonomous aircraft in a complex and changing environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of general image data processing or generation, and in particular to a target tracking and detection method, a detection device, a program product, and a storage medium for an optoelectronic pod. Background Art

[0002] In the civilian field, unmanned aerial vehicles (UAVs) and tethered UAVs play a crucial role. The aerospace detection devices such as optoelectronic seeker heads, optoelectronic pods, and triple-light hybrid pods carried by them are crucial for the precise detection of targets.

[0003] Currently, common target tracking and detection methods follow the traditional detection - learning - tracking process. In the initial stage, a large number of sample data are collected in a specific scenario to construct a target detection model. Subsequently, the model is applied to the actual scenario to perform target detection operations on the acquired video images. Once a target is detected, a tracking algorithm is started to continuously track the target, and in the subsequent process, the model is learned and updated using newly collected data containing the target to gradually improve the performance and accuracy of the model. This method can play a certain role in a relatively stable and less changing environment. When the target features and environmental conditions do not change much, it can achieve basic tracking and detection of the target.

[0004] However, for UAVs, the environment is large and the target changes greatly. When about to enter the tracking process, the target changes during this delay time, resulting in difficulty in precisely tracking and detecting the target in the first place and easy loss of the target. Summary of the Invention

[0005] This application provides a target tracking and detection method, a detection device, a program product, and a storage medium for an optoelectronic pod, which are used to solve the problems of difficult precise tracking and easy loss of the target caused by target changes, and enhance the stability and reliability of target tracking and detection.

[0006] In a first aspect, the present application provides a method for target tracking and detection of an optoelectronic pod, including: performing target detection on a video image to obtain a target detection result; during target detection, tracking the target in the video image to obtain a target tracking result; obtaining positive samples and negative samples according to the target detection result, the target tracking result, and the video image, where the positive samples are data containing the target in the target detection result and the target tracking result, and the negative samples are data not containing the target in the video image; training a target detection update model according to the positive samples and the negative samples; the target detection update model outputs feature data and a target model, the feature data is used to perform target detection on a new video image when receiving the new video image, and the target model is used to track the target in the new video image when receiving the new video image; performing a confidence score according to the target detection result and the target tracking result; and determining the target with the highest confidence score as the correct target.

[0007] By adopting the above technical solution, tracking is performed while detecting the target, shortening the delay time between detection and tracking. After obtaining the target detection result and the target tracking result, positive samples and negative samples are obtained based on them and the video image. This method can accurately collect target and non-target data, making the training samples more targeted and diverse. Then, the positive and negative samples are used to train the target detection update model, enabling the model to learn the changing characteristics of the target in real time and output more accurate feature data and a target model for the detection and tracking of new video images. Finally, the correct target is determined through a confidence score. With the interaction of these technical features, the tracking and detection accuracy of the target by autonomous or semi-autonomous aircraft in a complex and changing environment is improved, effectively solving the problems of difficult accurate tracking and easy loss of the target due to target changes, and enhancing the stability and reliability of target tracking and detection.

[0008] Combined with some embodiments of the first aspect, in some embodiments, the step of tracking the target in the video image to obtain a target tracking result specifically includes: dividing the video image into several regions; comparing the statistical characteristics of each region with the target feature range; if the statistical characteristics of the region do not belong to the target feature range, then performing a constant false alarm detection technique to adjust the target feature range; if the statistical characteristics of the region do not belong to the adjusted target feature range; then determining the region as a non-target region and performing suppression processing.

[0009] By adopting the above technical solution, the video image is divided into several regions and compared with the target feature range. For the regions that do not meet the requirements, the constant false alarm detection technology is used to adjust the target feature range. This way of regional processing and dynamic adjustment reduces unnecessary computational complexity. Since it is not necessary to perform comprehensive and complex calculations on the entire image, but focuses on the regions where the target may exist and the precise definition of the target feature range, it avoids the high computational complexity of processing the entire image without discrimination like the TLD algorithm. As a result, target tracking can be achieved more efficiently, reducing the demand for computing resources, improving the real-time performance of the algorithm, ensuring that the target can be tracked quickly and accurately during the operation of the aircraft, and adapting to complex and changing actual scenarios.

[0010] In combination with some embodiments of the first aspect, in some embodiments, after the step of determining the region as a non-target region and performing suppression processing, the method further includes: determining the region other than the non-target region as a potential target region; scoring the potential target region using a random fern classifier; selecting the region with the highest score among the preset data as the tracking region; and performing target tracking on the tracking region.

[0011] By adopting the above technical solution, after determining the non-target region and performing suppression processing, the remaining region is used as the potential target region, scored using a random fern classifier, and the region with the highest score among the preset data is selected as the tracking region for target tracking. The random fern classifier is characterized by high efficiency and speed, and can effectively screen and evaluate the potential target region in a short time. Compared with the complex calculation process in the TLD algorithm, it greatly reduces the computational complexity. By accurately positioning the tracking region in this way, it avoids the ineffective calculation of a large number of irrelevant regions, enables real-time target tracking, improves the operation efficiency of the entire target tracking system, and ensures that the aircraft can quickly lock the target and continuously track it in the face of complex environments and target changes, guaranteeing the smooth execution of the mission.

[0012] In combination with some embodiments of the first aspect, in some embodiments, after the step of performing confidence scoring based on the target detection result and the target tracking result, the method further includes: in the case where the target with the highest confidence score is lower than the scoring threshold or there is no target; extracting the feature data of the target in the most recent frame of the video image; determining the search region based on the target speed and the position information of the target; narrowing down the search region according to the movement direction of the target to obtain the region of interest; extracting the spare target in the subsequent video image according to the coordinates of the region of interest; performing confidence scoring on the spare target and the target; and determining the spare target with a confidence score higher than the scoring threshold as the lost target.

[0013] By adopting the above technical solution, when the target with the highest confidence score is lower than the score threshold or there is no target, the feature data of the target in the most recent frame is extracted, the search area is determined based on the target speed and position information, and the region of interest is obtained by narrowing down according to the movement direction. Spare targets are extracted in this region and their confidence scores are calculated to determine the lost target. This method avoids the aimless search and calculation in the entire image space like the TLD algorithm, but instead narrows down the search range targeted based on the motion characteristics of the target, reducing the computational amount. It ensures that the aircraft can still efficiently re-locate the target in case of a short-term loss or difficult detection of the target.

[0014] Combined with some embodiments of the first aspect, in some embodiments, after the step of determining the spare target with a confidence score higher than the score threshold as the lost target, the method further includes: calculating the average peak correlation energy of the spare target with a confidence score higher than the score threshold; if the average peak energy is less than the preset energy threshold, canceling the determination of the spare target with a confidence score higher than the score threshold as the lost target; the calculation function of the average peak correlation energy is:

[0015]

[0016] where APCE is the average peak correlation energy, is the maximum response value, is the minimum response value, is the response value of pixel point w and pixel point h.

[0017] By adopting the above technical solution, after determining the lost target, its average peak correlation energy is calculated and compared with the preset energy threshold to further confirm the authenticity of the lost target. It avoids learning a large amount of interfering information. When the target is blocked by the background or interfered by similar targets, it can more accurately determine whether the target is truly lost, improving the accuracy and stability of tracking, ensuring the continuous and effective tracking of the target by the aircraft in a complex environment, and reducing the tracking failure caused by misjudgment.

[0018] Combined with some embodiments of the first aspect, in some embodiments, the step of performing target detection on the video image to obtain the target detection result specifically includes: screening out the time-sensitive regions from the video image; extracting the feature information of the time-sensitive regions; matching and comparing the feature information with the pre-stored target feature template; and determining the time-sensitive regions with a matching result greater than the threshold as the target.

[0019] By adopting the above technical solution, time-sensitive regions are screened out from the video image, and their feature information is extracted and matched with the pre-stored target feature template. The regions with a matching result greater than the threshold are determined as the targets. The screening of time-sensitive regions can quickly focus on the key regions where target changes may occur or new targets may appear, avoiding comprehensive and time-consuming detection of the entire video image. Through the matching and comparison with the pre-stored template, using the existing target feature knowledge, the target is efficiently identified, reducing the computational amount and time consumption of ineffective detection.

[0020] Combined with some embodiments of the first aspect, in some embodiments, when detecting a target, the steps of tracking the target in the video image to obtain the target tracking result specifically include: extracting features from adjacent frame images to obtain the position information of the target in adjacent frames; obtaining the position change information according to the position information; determining the movement speed and movement direction of the target according to the position change information; predicting the position range where the target will appear in the next frame according to the movement speed and movement direction; and performing target matching and tracking according to the texture intensity and phase features of the target within the position range.

[0021] By adopting the above technical solution, features are extracted from adjacent frame images to obtain the target position information, and then the position change information is obtained to determine the movement speed and direction of the target. Next, the position range where the target will appear in the next frame is predicted, and target matching and tracking are performed in combination with the texture intensity and phase features of the target. Through the feature analysis of adjacent frames and the prediction of the target movement state, the regions where the target may appear can be planned in advance, reducing the blind search in the entire image space and improving the tracking efficiency. At the same time, using the texture intensity and phase features for matching and tracking enhances the accuracy and stability of target recognition. Especially when the appearance and posture of the target change, the target can be more accurately located.

[0022] In a second aspect, the present application provides an aerospace detection device, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions. The one or more processors call the computer instructions to enable the aerospace detection device to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0023] In a third aspect, the present application provides a computer program product containing instructions, which, when running on the aerospace detection device, enables the aerospace detection device to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0024] Fourthly, the present application provides a computer-readable storage medium, including instructions which, when running on an aerospace detection device, cause the aerospace detection device to execute the method described in the first aspect and any possible implementation manner of the first aspect.

[0025] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0026] 1. Tracking is performed while detecting the target, shortening the delay time between detection and tracking. After obtaining the target detection result and the target tracking result, positive and negative samples are obtained based on them and the video image. This method can accurately collect target and non-target data, making the training samples more targeted and diverse. Then, the positive and negative samples are used to train the target detection update model, enabling the model to learn the changing characteristics of the target in real time and output more accurate feature data and target models for the detection and tracking of new video images. Finally, the correct target is determined through confidence scoring. Through the interaction of these technical features, the tracking and detection accuracy of the autonomous or semi-autonomous aircraft for the target in a complex and changeable environment is improved, effectively solving the problems of difficult accurate tracking and easy loss of the target caused by target changes, and enhancing the stability and reliability of target tracking and detection.

[0027] 2. The video image is divided into several regions and compared with the target feature range. For regions that do not conform, the constant false alarm detection technology is used to adjust the target feature range. This method of regional processing and dynamic adjustment reduces unnecessary computational complexity. Since there is no need to perform comprehensive and complex calculations on the entire image, but rather focus on the regions where the target may exist and the accurate definition of the target feature range, it avoids the high computational complexity of processing the entire image without discrimination like the TLD algorithm. As a result, target tracking can be achieved more efficiently, reducing the demand for computing resources, improving the real-time performance of the algorithm, and ensuring that the target can be tracked quickly and accurately during the operation of the aircraft to adapt to complex and changeable actual scenarios.

[0028] 3. After determining the lost target, calculate its average peak correlation energy and compare it with a preset energy threshold to further confirm the authenticity of the lost target. This avoids learning a large amount of interfering information. When the target is blocked by the background or interfered by similar targets, it can more accurately determine whether the target is truly lost, improving the accuracy and stability of tracking, ensuring the continuous and effective tracking of the target by the aircraft in a complex environment, and reducing the tracking failure cases caused by misjudgment. Description of the Drawings

[0029] Figure 1 is a schematic flowchart of a control method for an optoelectronic pod of an unmanned aerial vehicle in an embodiment of the present application;

[0030] Figure 2 It is a schematic flowchart of a target recognition, tracking and guidance method for an optoelectronic pod in an embodiment of the present application;

[0031] Figure 3 It is a schematic flowchart of a target tracking and detection method for an optoelectronic pod in an embodiment of the present application;

[0032] Figure 4 It is a schematic framework diagram of a target tracking and detection method for an optoelectronic pod in an embodiment of the present application;

[0033] Figure 5 It is a schematic diagram of an exemplary hardware structure of an aerospace detection device in an embodiment of the present application. Detailed implementation manners

[0034] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above", "said", "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to and includes any or all possible combinations of one or more of the listed items.

[0035] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.

[0036] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of a control method for an optoelectronic pod of an unmanned aerial vehicle in an embodiment of the present application;

[0037] A control method for an optoelectronic pod of an unmanned aerial vehicle, comprising:

[0038] S101. When it is determined that the current position is within the mission area, guide the autonomous or semi-autonomous aircraft to fly according to the planned route;

[0039] Among them, the mission area refers to a spatial area that is preset according to specific aerospace detection tasks and has clear boundaries and ranges, and its boundaries are usually determined by parameters such as geographical coordinates and altitude ranges; the planned route refers to a flight path designed after comprehensive consideration of mission requirements, aircraft performance, environmental conditions (such as terrain, meteorology, etc.) and safety factors, and it includes a series of continuous waypoints and flight trajectories connecting these waypoints.

[0040] In some embodiments, after the aircraft starts and takes off, the real-time position information of the aircraft is continuously obtained. At the same time, the pre-stored mission area information is read, and the real-time position is compared with the boundary conditions of the mission area. Once it is determined that the current position is within the mission area, the aircraft retrieves the corresponding planned route data. Then, according to the starting point and the initial flight direction in the planned route, the power system and the flight attitude adjustment device of the aircraft are controlled to make the aircraft start flying along the planned route, ensuring that it can accurately enter the mission area and perform the detection task according to the predetermined plan.

[0041] S102. Scan the mission area on the planned route;

[0042] In some embodiments, when the aircraft is flying along the planned route, the detection device is activated according to the preset scanning scheme. During the scanning process, the scanning control system will monitor the position, speed, and attitude information of the aircraft in real time, and adjust the working parameters of the detection device according to this information to ensure the stability and accuracy of the scanning. For example, the shooting frame rate of the camera is adjusted according to the flight speed of the aircraft to ensure the resolution and clarity of the image.

[0043] It should be noted that for the optical detection components of aerospace detection equipment, they are generally divided into visible light television, infrared thermal imaging, laser rangefinder, and camera.

[0044] In some embodiments, the seeker includes: an optical detection component, a seeker frame, and a servo platform. The optical detection component includes: a short-focus visible light television, a long-focus visible light television, infrared thermal imaging, a laser rangefinder. The servo platform includes: a three-axis gyro, a driver, a motor, an angle measuring sensor, and a servo computer board; a power supply board, an electrical connector, and a video tracking and recording board.

[0045] In some other embodiments, the pod includes: an optical detection component, a seeker frame, and a servo platform. The optical detection component includes: a visible light television and an infrared thermal imager. The seeker frame includes: a gyro and a driver. The servo platform includes: a motor, an angle measuring sensor, a servo computer board, a power supply board, a video tracking and recording board, and an electrical connector.

[0046] In some other embodiments, the optoelectronic pod includes: an optical detection component, an optoelectronic pod frame, and a servo platform. The optical detection component includes: a zoom camera, a wide-angle camera, an infrared thermal imager, and a laser rangefinder. The optoelectronic pod frame includes: a three-axis gyro and a driver. The servo platform includes: a motor, an angle measuring sensor, a servo computer board, a power supply board, a video tracking and recording board, and an electrical connector.

[0047] The equipment in the optical detection component will be described below in sequence:

[0048] In some embodiments, the task area is scanned using visible light;

[0049] Among them, the action distance function is:

[0050]

[0051] In the formula, is the distance to the target, H is the critical dimension of the target, is the focal length of the objective lens, is the limit resolution of the CMOS, and n is the spatial frequency;

[0052] In some embodiments, the main function of the visible light television is to provide a stable daytime visible light image for the seeker. When the seeker is working, the visible light television provides real-time images. For changes in the external light intensity, the visible light television adaptively adjusts the exposure time of the CMOS device to ensure real-time output of images suitable for observation and tracking.

[0053] The working principle of the visible light television is as follows: The visible light reflected by the ground scenery is transmitted through the atmosphere, gathered on the focal plane of the detector by the optical system, converted into image electrical signals by the detector, and the image electrical signals are stored, dehazed and enhanced, target recognized, extracted, positioned, and information calculated and marked through the image processing system, and then image products are formed.

[0054] It should be noted that to complete the detection and recognition of the target, while meeting the requirements of the target imaging size (spatial resolution), the contrast requirements of the target imaging also need to be met. The combination of the two can meet the detection and recognition of the target.

[0055] The contrast function is:

[0056]

[0057] In the formula, is the contrast, is the contrast of the output signal of the television sighting device, is the signal processing modulation transfer function, is the modulation transfer function of the display device, is the inherent contrast, is the atmospheric contrast transfer coefficient, is the modulation transfer function of the optical system, is the modulation transfer function.

[0058] In some other embodiments, the visible light TV is equipped with a focusing motor (referred to as a zoom camera at this time). The rotation of the motor gear drives the rotation of the motor driving wheel, which in turn drives the focusing seat. The rotation of the focusing seat drives the forward and backward movement of the front lens group. By moving the front lens group, a slight change in the focal length of the optical system is achieved to stabilize the image plane, thereby achieving the purpose of focusing. When focusing reaches the two limits, there are automatic slipping mechanisms and microcomputer switch in-place protection mechanisms on the motor to prevent damage caused by motor jamming. The circuit control part of the entire lens has a protection circuit to prevent damage to the whole machine caused by high voltage and incorrect connection, improving the reliability of the whole machine. The motor and the lens are matched in a backpack style and distributed along the arc of the lens, with the smallest volume.

[0059] In some other embodiments, the purpose of automatic light adjustment is to adapt to the current ambient light by changing the aperture size, exposure amount, and gain value of the optical imaging system, and then changing the luminous flux of the optical imaging system to obtain good images. The light adjustment mainly includes the following three parts:

[0060] Adjustment of the aperture. The aperture controls the luminous flux per unit time. As the light intensity increases, the aperture diameter is increased by motor drive to increase the luminous flux, and vice versa. However, improper aperture control will cause noise interference and affect the image quality. The adjustment of the aperture does not consider continuous adjustment, but considers scene-based adjustment. Therefore, the design of this system selects a fixed aperture size, and the light adjustment control is completed by adjusting two variables, the gain value and the exposure time.

[0061] The exposure time controls the light input time and thus the light input amount. The light input amount is the product of the light input time and the light intensity.

[0062] Gain control adjusts the gray value of each pixel in the image through the internal amplifier circuit of the optical imaging system. While amplifying the effective signal, it will also amplify the interference noise. Generally, exposure time control is considered first, and then gain control.

[0063] It can be seen that when using visible light to scan the task area, the various parameters in the action distance function and the contrast function, such as the objective lens focal length, the ultimate resolution of the CMOS, the spatial frequency, the contrast, and the various transfer functions, cooperate with each other. The objective lens focal length and the ultimate resolution of the CMOS determine the ability to capture target details, comprehensively affecting the clarity and contrast of the image, enabling the target and the background to be more clearly distinguished during the scanning process, and more accurately determining the position and characteristics of the target, thereby improving the accuracy and efficiency of target detection.

[0064] In some embodiments, an infrared thermal imager solution is used to scan the task area;

[0065] Among them, the function of the minimum resolvable temperature difference is:

[0066]

[0067] Wherein, MRTD is the minimum resolvable temperature difference at the target spatial frequency f, is the threshold signal-to-noise ratio, MTF(f) is the transfer function of the thermal imaging system, NETD is the noise equivalent temperature difference, is the instantaneous field of view of a single pixel of the detector, is the effective integration time of the human eye, is the frame rate of the thermal imaging system, is the equivalent noise bandwidth of the circuit, is the dwell time;

[0068] In some embodiments, the infrared thermal imager can conduct reconnaissance and surveillance on targets day and night, and is mainly used in the airborne field with high reliability requirements, featuring small size, light weight, and strong environmental adaptability. It mainly consists of two parts: optical components and mechanical components.

[0069] The working principle of the infrared thermal imager is as follows: The infrared radiation of the target background is transmitted through the atmosphere to the infrared imager. First, the infrared optical system receives the infrared radiation and reflection of the target and converges the energy onto the photosensitive surface of the infrared detector. The detector converts the target radiation intensity into an analog electrical signal, and the image preprocessing circuit module buffers, amplifies, and performs A / D conversion on the analog electrical signal to a digital signal, and realizes functions such as non-uniformity correction, defective pixel detection and replacement, gain and bias control of the image signal, and outputs analog and digital video signals.

[0070] The function of the operating range is:

[0071]

[0072] Wherein, T is the temperature difference between the target and the background after atmospheric attenuation, is the actual equivalent temperature difference between the target and the background, is the average atmospheric transmittance at the operating range R, MRTD is the minimum resolvable temperature difference at the target spatial frequency f, H is the characteristic size of the target, is the equivalent number of pixels occupied by the target on the detector.

[0073] It can be seen that when using the infrared thermal imager solution for scanning, the parameters in the minimum resolvable temperature difference function and the operating range function work together. Factors such as the transfer function of the thermal imaging system, the noise equivalent temperature difference, the instantaneous field of view of a single pixel of the detector, the effective integration time of the human eye, the frame rate of the thermal imaging system, the equivalent noise bandwidth of the circuit, the dwell time, and the atmospheric transmittance enable the detection of the target and the determination of its operating range under different temperature environments and target thermal radiation characteristics. It compensates for the deficiencies of visible light scanning, further improves the comprehensiveness and accuracy of target detection, and reduces the situation of target omission caused by environmental factors.

[0074] S103. During the scanning process, run the target detection algorithm to detect the targets in the task area;

[0075] In some embodiments, preprocess the data, such as operations like grayscale conversion, noise reduction, enhancement, normalization of images, and processing such as filtering, amplification, and demodulation of signals, to improve the quality of the data and the distinguishability of features. Then, according to the type of the target detection algorithm (such as traditional algorithms based on feature extraction or deep learning algorithms based on neural networks), perform feature extraction and pattern recognition. For traditional algorithms, features such as the shape, color, texture, and edges of the target may be extracted and compared and matched with a pre-established target feature library; for deep learning algorithms, the preprocessed data is input into a trained neural network model, and the model will automatically learn and extract the target features in the data and output information such as the category, location, and confidence of the target. Once determined as a target, the algorithm will record the detailed information of the target and mark it in the original data for subsequent tracking and analysis.

[0076] It should be noted that in some embodiments, the following "target recognition, tracking, and guidance method for an optoelectronic pod" can be used for detection. Specifically, it is steps S201 to S202, or S201 to S205, or S201 to S207;

[0077] In some other embodiments, the following "target tracking and detection method for an optoelectronic pod" can be used for detection, specifically step S301. In this case, steps S103 and S104 are synchronized in time.

[0078] S104. Track the detected targets;

[0079] In some embodiments, after the target detection algorithm determines a target, the aerospace detection device first adjusts the orientation of the autonomous or semi-autonomous aircraft according to the initial position information of the target to align it with the target. Then, as the aircraft continues to fly and the target moves, the aerospace detection device obtains the motion information of the target through various methods. For example, for optical image tracking, by comparing consecutive image frames and using image feature matching algorithms (such as optical flow method, SIFT feature matching, etc.) to calculate the displacement of the target in the image, and combining the flight parameters of the aircraft (such as speed, altitude, attitude angle, etc.), the actual motion speed and direction of the target are deduced; for radar tracking, by analyzing parameters such as the frequency change and time delay change of the radar echo signal, information such as the distance, speed, and angle change of the target is calculated in real time. According to this target motion information, the aerospace detection device controls the aircraft to correspondingly adjust the flight attitude and speed of the aircraft, so that the aircraft can maintain the relative position and angle with the target, ensuring that the target is always within the effective observation range of the detection device and realizing stable tracking of the target.

[0080] It should be noted that in some embodiments, the following "target recognition, tracking and guidance method of optoelectronic pod" can be used for tracking. Specifically, it is steps S203 to S209, or S206 to S209, or S208 to S209;

[0081] In some other embodiments, the following "target tracking and detection method of optoelectronic pod" can be used for tracking, specifically step S302. In this case, steps S103 and S104 are synchronized in time.

[0082] In some embodiments, step S104 specifically includes:

[0083] S1041. When detecting a target, track the target;

[0084] S1042. When detecting multiple targets, perform a threat degree ranking on the multiple targets;

[0085] Among them, the threat degree ranking refers to evaluating and ranking the potential threat degrees of multiple targets according to a preset evaluation criterion.

[0086] In some embodiments, perform type recognition on the detected multiple targets, determine the target type by comparing with a known target database, assign different basic threat values to different types of targets, then perform weighted calculation on the basic threat values according to the motion parameters of the targets, and finally obtain the threat degree of each target and perform ranking, without calculation here.

[0087] S1043. Determine the target with the highest threat degree and track the target with the highest threat degree;

[0088] S1044. In the case where no target is detected, return to step S103.

[0089] It can be seen that when multiple targets are detected, threat level ranking is performed and the target with the greatest threat is determined for tracking. This ensures that the aircraft concentrates its limited resources on the most critical targets, avoiding resource waste and omission of key information caused by non-discriminatory tracking of multiple targets. This enables quick and accurate locking of the most threatening target in the face of complex mission scenarios.

[0090] In some embodiments, step S104 specifically includes:

[0091] S1045. In the manual tracking mode, receive the instruction input by the user;

[0092] S1046. Perform azimuth and / or pitch movement according to the instruction.

[0093] S1047. In the automatic tracking mode, according to the position information of the target in the image, determine the azimuth parameter and / or pitch parameter based on the position information and the position information of the image center, and perform azimuth and / or pitch movement according to the azimuth parameter and / or pitch parameter.

[0094] It can be seen that receiving the instruction input by the user and performing azimuth and / or pitch movement according to the instruction gives the operator direct control over target tracking in special situations. When there may be deviations in the automatic tracking system or more manual judgment is required, the operator can flexibly adjust the tracking direction based on their own experience and understanding of the mission to ensure that the target is always within the monitoring range.

[0095] It can be seen that in the automatic tracking mode, the azimuth and pitch parameters are determined based on the position information of the target in the image and the position information of the image center, and movement is performed accordingly. Through this automatic adjustment mechanism based on image information, the system can quickly and accurately respond to the movement of the target and maintain stable tracking of the target. This enhances the autonomy and stability of autonomous or semi-autonomous aircraft during the execution of detection tasks, reduces the dependence on continuous manual operation, thereby reducing labor costs and the risk of operation errors.

[0096] S105. After tracking, initiate the laser ranging strategy to obtain the azimuth, pitch angle information, and distance information of the target;

[0097] In some embodiments, after the aircraft stably tracks the target, the laser rangefinder will start to work. First, the laser rangefinder will emit a laser beam with a specific wavelength, power, and pulse width towards the target. After the laser beam encounters the target, it will be reflected back and received by the laser rangefinder on the aircraft. By measuring the time interval between the laser emission and reception with a high-precision time measurement device, according to the formula of the speed of light, the distance between the target and the aircraft can be calculated. At the same time, by using the angle measurement sensors carried on the aircraft, the current attitude angles of the aircraft (including roll angle, pitch angle, and yaw angle) and the pointing angles of the detection device (azimuth angle and pitch angle) are obtained in real time. Through coordinate transformation and geometric calculation, the azimuth angle and pitch angle of the target relative to the aircraft coordinate system are determined.

[0098] In some embodiments, the laser ranging strategy includes:

[0099] The operating range equation is:

[0100]

[0101] T = exp(-αδR)

[0102] Wherein, is the minimum detectable power at a range of R; is the transmitted power; is the receiving area; is the effective target reflection area; is the effective target reflection area; is the optical system transmittance; is the two-way atmospheric transmittance; is the laser emission beam angle; R is the range, α is the atmospheric attenuation coefficient, and δ is the elevation correction factor.

[0103] It can be seen that laser ranging is carried out based on the operating range equation. The parameters such as transmitted power, receiving area, effective target reflection area, optical system transmittance, two-way atmospheric transmittance, laser emission beam angle, atmospheric attenuation coefficient, and elevation correction factor act together, and can calculate the range and minimum detectable power according to different environmental conditions and target characteristics, so as to achieve high-precision measurement of the target distance in the complex and changeable aerospace environment, ensure that the obtained target azimuth and pitch angle information is highly matched with the distance information, provide a reliable data basis for subsequent flight path adjustment, target analysis, etc., and improve the accuracy of the detection device for target positioning.

[0104] S106. Report the azimuth, pitch angle information, and distance information.

[0105] It can be seen that when it is determined that the aircraft is in the mission area, flying according to the planned route ensures the orderliness and pertinence of the detection. During the flight, the mission area is scanned and the target detection algorithm is run to search for targets. After detecting the target, tracking is carried out to maintain continuous attention on the target and obtain more detailed and accurate target dynamic information. Then, the laser ranging strategy is activated to accurately determine the azimuth, pitch angle and distance information of the target. These information complement each other and work together to provide comprehensive and accurate data support for subsequent decision-making, avoiding decision-making difficulties caused by missing or inaccurate data, reducing the dependence on manual repeated confirmation and supplementary data, reducing labor costs, and improving the execution efficiency of the entire detection mission.

[0106] Please refer to Figure 2 , Figure 2 which is a schematic flow chart of the target recognition, tracking and guidance method of the optoelectronic pod in the embodiment of the present application;

[0107] A target recognition, tracking and guidance method for an optoelectronic pod, comprising:

[0108] S201. Perform superpixel segmentation on the acquired scanned image information to remove the background and obtain superpixel regions;

[0109] Among them, the scanned image information refers to the image data obtained by scanning the mission area through aerospace detection equipment, and these data contain the visual information of the target object and the surrounding environment. Superpixel segmentation is an image processing technology that divides an image into many small regions with similar characteristics, and these small regions are called superpixels. Compared with traditional pixel-level segmentation, superpixel segmentation can better retain the edge and region information of the image.

[0110] In some embodiments, when the aircraft reaches the predetermined mission area and starts to perform the target detection mission, once sufficient image data is obtained, the image is divided into multiple initial superpixel blocks. Then, these superpixel blocks are merged and adjusted through an optimization algorithm to ensure that the pixels within each superpixel region have higher similarity and consistency. During this process, the boundaries between the superpixel regions and the surrounding regions are continuously evaluated, and those superpixel regions that are obviously part of the background are removed, and finally the superpixel regions after removing the background are obtained, and these regions will be used as the basis for subsequent target recognition.

[0111] In some embodiments, in some embodiments, by monitoring and analyzing the pixel change situation and movement state of the target area, the target and the background are distinguished. When it is determined that the target in a certain area does not have the situation that the pixel change is greater than the established threshold and no position movement occurs, that area is determined as the background area.

[0112] S202. Perform object recognition on the superpixel region to identify the object;

[0113] In some embodiments, the following "Object Tracking and Detection Method for Optoelectronic Pod" can be used for recognition, specifically as step S301. In this case, steps S202 and S209 are synchronized in time.

[0114] In other embodiments, before object recognition, step S102 of the above "Control Method for Optoelectronic Pod of Unmanned Aerial Vehicle" can be used to obtain a scanned image.

[0115] An object refers to an object that has been determined to have specific significance and value through the object recognition process and requires subsequent tracking and processing.

[0116] In some embodiments, various feature information is extracted from the superpixel region, including color features (such as color histograms, color moments, etc.), texture features (such as gray-level co-occurrence matrices, local binary patterns, etc.), shape features (such as contour shapes, areas, perimeters, etc.), and deep learning-based features (such as high-level features extracted through convolutional neural networks), etc. Then, these extracted features are compared and matched with the object feature templates pre-stored in the aircraft database or the trained object recognition model. If the feature similarity exceeds a preset threshold, the object in the superpixel region is determined as an object, and its position, size, category, and other information are recorded to provide basic data for subsequent object tracking and processing.

[0117] In other embodiments, the object judgment is a time-sensitive object, that is, step S202 includes:

[0118] S2021. Calculate the gradient magnitude and direction of each pixel point in the superpixel region. The gradient magnitude is the intensity of the gray-level change of the pixel point, and the direction is the direction of the fastest gray-level change;

[0119] In some embodiments, for each pixel point in the superpixel region, a specific gradient operator (such as Sobel operator, Prewitt operator, etc.) is used to calculate its gradient magnitude and direction. Taking the Sobel operator as an example, it performs differential operations on the gray-level values of the pixel point in the horizontal and vertical directions respectively, and then calculates the gradient magnitude through a certain mathematical formula (such as the Pythagorean theorem), and at the same time determines the direction of the fastest gray-level change as the gradient direction according to the differential result. In this way, the gradient magnitude and direction information of each pixel point in the superpixel region can be obtained.

[0120] S2022. Perform non-maximum suppression on the gradient magnitude, and only retain the point with the largest magnitude in the direction to obtain the edge;

[0121] In some embodiments, for each pixel, its gradient direction is first determined, and then the gradient magnitudes of its adjacent pixels are checked in that direction. If the gradient magnitude of the pixel is not the maximum among the adjacent pixels in its gradient direction, the gradient magnitude of the pixel is set to zero, i.e., the point is suppressed; only the pixel with the maximum magnitude in the gradient direction is retained. By processing all pixels in the superpixel region in this way, an edge image after non-maximum suppression can be obtained, and these edge points can more accurately reflect the true contour of the target object, reducing edge blurring and misjudgment.

[0122] S2023. Determine a first threshold and a second threshold based on the statistical distribution of the gradient magnitudes, where the first threshold is greater than the second threshold.

[0123] In some embodiments, the gradient magnitudes of all pixels in the superpixel region are statistically analyzed, and statistics such as the maximum value, minimum value, mean value, and variance are calculated. Then, based on these statistics and preset rules, the first threshold and the second threshold are determined. A common method is to adopt the idea of Otsu's method, considering the histogram of the gradient magnitudes as being composed of two or more pixel points with different distributions (corresponding to the target and the background), and determining the first threshold and the second threshold by finding the threshold that maximizes the between-class variance. The thresholds determined in this way can better adapt to the characteristics of different images, reasonably divide the edge points into strong edges and weak edges, and provide more targeted edge information for subsequent target extraction and recognition.

[0124] In some specific embodiments, the histogram of the gradient magnitudes is statistically analyzed and its distribution is observed. If the histogram exhibits an obvious bimodal feature (i.e., there are two peaks, corresponding to the regions where the gradient magnitudes of the target edge and the background are concentrated), the first threshold and the second threshold are determined at the bottom of the valley between the two peaks. For example, the minimum point between the two peaks in the histogram can be found as the first threshold, and then the second threshold can be determined according to a certain proportional relationship (such as half of the first threshold), which is not limited here.

[0125] S2024. Determine the edges with gradient magnitudes higher than the first threshold as strong edges.

[0126] In some embodiments, if the gradient magnitude of a pixel is higher than the first threshold, the pixel is marked as a strong edge point, and these strong edge points will be retained for subsequent target construction and recognition steps.

[0127] S2025. Determine the edges with gradient magnitudes lower than the second threshold as weak edges.

[0128] In some embodiments, for those pixel points whose gradient magnitude is lower than the second threshold, they are marked as weak edge points. Different from the processing of strong edge points, weak edge points may not be directly used to construct the main contour of the target like strong edge points, but they will serve as auxiliary information and be associated and integrated with strong edge points in subsequent steps to improve the edge description and feature extraction of the target object.

[0129] S2026. Connect the strong edge with the corresponding weak edge to obtain the time-sensitive target;

[0130] In some embodiments, analyze the strong edge to determine its continuity and direction features, and search for possible connected weak edges within a certain range around it according to these features. For example, along the tangent direction or normal direction of the strong edge, search for weak edge points within a set neighborhood whose gradient magnitude is lower than the first threshold but higher than the second threshold and whose direction has a certain coherence with the strong edge. Once appropriate weak edge points are found, connect them to the strong edge. In this way, gradually improve the edge contour of the target to make the description of the time-sensitive target more accurate and complete.

[0131] S2027. Perform target detection of local image features on the time-sensitive target to obtain the target.

[0132] In some embodiments, select multiple representative local regions from the image region of the time-sensitive target. These regions can be determined according to the shape of the target, edge distribution, and prior knowledge. Then, for each local region, extract its corresponding local image features, such as the texture, shape, and color features described above. Next, compare and match these extracted local features with the local feature templates of various targets pre-stored in the aircraft database or the trained target recognition model. By calculating the similarity between the features (such as Euclidean distance, cosine similarity, etc.), if the similarity exceeds the preset threshold, then determine the time-sensitive target as a specific target category, and further accurately locate its position in the image, and record information such as its center coordinates, size, and direction.

[0133] It can be seen that calculating the gradient magnitude and direction of the pixels in the superpixel region can effectively extract the edge information of the image, because the edge is usually the boundary between the target and the background. By non-maximum suppression, the points with the maximum magnitude in the direction are retained to obtain the edge, which further refines the edge information and highlights the contour features of the target. According to the statistical distribution of the gradient magnitude, a reasonable threshold is determined to divide the edge into strong edges and weak edges, and the strong edges are connected to the corresponding weak edges to obtain the time-sensitive target. This target extraction method based on edge features can quickly locate the potential target region in a complex background and reduce the interference of irrelevant regions. For the time-sensitive target, local image feature-based target detection is performed, and the local features of the target are used for recognition, which improves the accuracy and speed of target recognition, especially applicable to situations where the appearance of the target varies and there are partial occlusions in a complex environment.

[0134] S203. Obtain the original position and attitude information of the sensor;

[0135] In some embodiments, the aerospace detection device directly reads the latest measurement values from the registers or data buffer areas of each sensor. For example, for the GPS sensor, the current position coordinate information is obtained from the GPS receiver through a specific communication protocol and data parsing algorithm; for the IMU, the measured acceleration and angular velocity data are read, and the attitude angle of the aircraft is preliminarily calculated through integral operation. These raw data will be stored in the memory of the aircraft after being obtained for use in subsequent steps, which will not be elaborated here.

[0136] S204. Based on the position information of the target in the scanned image information, use the distance information from the target to the sensor, the original position and attitude information to calculate the three-dimensional coordinates of the target in three-dimensional space;

[0137] In some embodiments, according to the pixel coordinates of the target in the image, combined with the internal parameters of the camera (such as focal length, optical center position, etc.) and external parameters (such as the rotation and translation matrices of the camera, which can be calculated from the attitude information and position information of the aircraft), through the perspective projection transformation formula, the pixel coordinates of the target are converted into two-dimensional coordinates in the camera coordinate system. Then, using the distance information from the target to the sensor, the two-dimensional coordinates in the camera coordinate system are extended to three-dimensional coordinates. Finally, according to the original position information of the aircraft, the three-dimensional coordinates of the target in the camera coordinate system are converted into three-dimensional coordinates referenced to the earth coordinate system or other global coordinate systems, so as to obtain the position information of the target in three-dimensional space.

[0138] In some embodiments, step S204 specifically includes:

[0139] S2041. Fuse the thermal imaging image and the scanned image information to obtain a fused image;

[0140] Among them, the fused image is a new image obtained by integrating the thermal imaging image and the scanning image information through specific algorithms and technical means. It combines the advantages of both, including both the thermal characteristic information of the object in the thermal imaging image and the details and positions of the target in the scanning image, etc.

[0141] In some embodiments, necessary preprocessing is performed on the thermal imaging image and the scanning image information in terms of spatial coordinates, resolution, data format, etc., to make them meet the basic requirements for fusion. For example, the resolution of the thermal imaging image is adjusted to match that of the scanning image; the coordinate systems of the two are calibrated so that they can correspond spatially; a suitable fusion algorithm is selected. Common ones include pixel-level fusion algorithms (such as weighted average method, wavelet transform method, etc.) and feature-level fusion algorithms (such as principal component analysis method, feature fusion matching method, etc.). Taking the weighted average method as an example, for the pixel points at the corresponding positions in the thermal imaging image and the scanning image, weighted summation is performed according to a certain weight (the weight can be preset according to the image characteristics and actual needs. For example, when the thermal imaging image focuses more on temperature information, its weight can be appropriately increased), and the obtained new pixel value is assigned to the corresponding pixel position in the fused image. By performing such operations on all pixel points in turn, the fused image can be obtained.

[0142] S2042. Based on the position information of the target in the fused image, use the distance information from the target to the sensor, the original position, and the attitude information to calculate the three-dimensional coordinates of the target in three-dimensional space.

[0143] It should be noted that the principle and process of this step are similar to those of step S204. The relevant principles and processes can refer to step S204 and are not limited here.

[0144] It can be seen that fusing the thermal imaging image and the scanning image information to obtain a fused image gives full play to the advantages of the thermal imaging image in detecting heat-emitting targets, penetrating smoke, etc. and the strengths of the scanning image in obtaining target details and shape information.

[0145] S205. Obtain the feature information of each target;

[0146] Among them, the features of the target refer to various information that can describe the unique attributes and characteristics of the target object, including but not limited to the visual features of the target (such as color, shape, texture, edges, etc.), physical features (such as size, mass, material, etc.), and motion features (such as speed, acceleration, motion direction, etc.). These feature information can be used to distinguish different target objects.

[0147] In some embodiments, after determining the position of the target and calculating its three-dimensional coordinates, the feature extraction module of the aircraft starts to work. For each target, its visual features are first extracted from the scanned image, such as calculating the color histogram of the target, shape descriptors (such as Hu moments, Fourier descriptors, etc.), and texture features (such as gray-level co-occurrence matrix features) through image segmentation and feature extraction algorithms. At the same time, combining the motion information of the target, such as calculating its velocity and acceleration vectors by the position changes of the target in consecutive image frames, and measuring the electromagnetic or thermal radiation characteristics of the target using other sensors (such as radar, infrared sensors, etc.). These different types of feature information are integrated and quantified to form a feature vector for each target, which is stored in the memory of the aircraft for subsequent tasks such as target trajectory detection and identity recognition.

[0148] S206. Perform target trajectory detection based on multiple hypothesis tracking according to the feature information and three-dimensional coordinates to obtain trajectory data;

[0149] Among them, target trajectory detection based on multiple hypothesis tracking is a method of tracking by simultaneously considering multiple possible motion states and trajectories of the target. It will infer multiple possible motion trajectory hypotheses of the target in the future based on the historical position (three-dimensional coordinates) and feature information of the target, and then as new observation data (such as three-dimensional coordinates and feature information at subsequent moments) are obtained, these hypotheses are verified and screened to gradually determine the trajectory that most conforms to the actual motion of the target. Multiple hypothesis tracking will consider these different motion possibilities, generate multiple potential driving trajectory hypotheses, and then determine the most likely actual driving trajectory of the vehicle according to the new position information.

[0150] In some embodiments, step S206 specifically includes:

[0151] S2061. Construct multiple different motion trajectories for each target according to the dynamic change data generated by the feature information and three-dimensional coordinates over time;

[0152] In some embodiments, when the aircraft continuously monitors a target, as time progresses, it continuously collects the characteristic information and three-dimensional coordinate data of the target and records its dynamic changes. For each target, first analyze its historical characteristic information and the changing trend of three-dimensional coordinates. For example, if the speed of the target shows a gradually increasing trend over a certain period of time and the color characteristics remain stable, based on this information, combined with the possible motion patterns of the target (such as uniformly accelerated linear motion, uniform circular motion, etc.) and environmental constraints (such as terrain, obstacle distribution), multiple different motion trajectory hypotheses are generated through mathematical models and algorithms. These hypotheses not only consider the range of changes in motion parameters such as the speed and acceleration of the target, thereby constructing a series of trajectory models covering various possible motion paths of the target, providing multiple candidate solutions for subsequent tracking and verification to cope with the complex and changeable motion states of the target and improving the reliability and accuracy of target tracking.

[0153] In some embodiments, step S2061 specifically includes:

[0154] S20611. Calculate the velocity components of the target in the directions of each coordinate axis according to the changes in the three-dimensional coordinates of the target at consecutive time points;

[0155] In some embodiments, when the aircraft continuously obtains the three-dimensional coordinate information of the target and accumulates the coordinate data of a certain number of consecutive time points, the calculation of the velocity components begins. First, select the three-dimensional coordinates of the target corresponding to two adjacent time points, and then obtain the velocity components in the directions of each coordinate axis by calculating the ratio of the coordinate difference to the time interval.

[0156] S20612. Calculate the acceleration components of the target in the directions of each coordinate axis according to the changes in the velocity components at different times;

[0157] In some embodiments, after obtaining the data of the velocity components of the target in the directions of each coordinate axis changing with time, the calculation of the acceleration components begins. Select the velocity component values corresponding to at least three consecutive time points. Obtain the acceleration components by calculating the ratio of the difference between adjacent velocity components to the time interval;

[0158] In some embodiments, in order to improve the accuracy, usually, methods such as calculating the average value multiple times can be used for optimization.

[0159] S20613. Construct a motion hypothesis space by changing the acceleration, velocity, and direction angle;

[0160] In some embodiments, after obtaining the acceleration components, velocity components of the target in each coordinate axis direction and the relevant descriptions of the determined direction angle, the construction of the motion hypothesis space begins. The acceleration, velocity, and direction angle are combined within their respective value ranges to form a multi-dimensional space, which is the motion hypothesis space. This space contains various possible combinations of the target's future motion states, providing rich possibilities for subsequent extraction of specific motion hypotheses.

[0161] S20614. Extract the next motion state parameters of the target from the motion hypothesis space to generate multiple motion hypotheses that conform to the probability distribution. The motion state parameters include acceleration, velocity, and direction angle. When extracting the velocity, the current velocity of the target is used as the midpoint, and the velocity change amount generated by the current acceleration within the preset time step is used as the radius to determine the extraction range. When extracting the acceleration, the current acceleration of the target is used as the midpoint, and the acceleration change amplitude is used as the radius to determine the extraction range. When extracting the direction angle, the current direction angle of the target is used as the midpoint, and the maximum steering angle is used as the radius to determine the extraction range.

[0162] In some embodiments, after constructing the motion hypothesis space, it is necessary to extract the next motion state parameters of the target from this space to generate multiple motion hypotheses that conform to the probability distribution. First, according to the pre-determined probability distribution model (this model can be constructed based on factors such as the type of the target, historical motion data, and the environment. For example, for some common target motion patterns, a higher probability is assigned; for some extreme cases that do not conform to physical laws or the target's normal operations, a lower probability is assigned), the three motion state parameters of acceleration, velocity, and direction angle are respectively extracted in the motion hypothesis space. When extracting the velocity, the current velocity of the target is used as the midpoint, and the velocity change amount generated by the current acceleration within the preset time step is used as the radius to determine the extraction range. For extracting the acceleration, the current acceleration of the target is used as the midpoint, and the acceleration change amplitude is used as the radius to determine the extraction range. When extracting the direction angle, the current direction angle of the target is used as the midpoint, and the maximum steering angle is used as the radius to determine the extraction range. In this way, each time a set of motion state parameters of acceleration, velocity, and direction angle is extracted, a motion hypothesis is generated. By repeating the extraction process multiple times, multiple motion hypotheses that conform to the probability distribution can be obtained. These motion hypotheses cover various possible motion situations of the target in the future and consider the probability of their occurrence, which is more conducive to subsequent tracking and judgment of the target's actual motion trajectory.

[0163] It can be seen that by calculating the velocity and acceleration components of the target in each coordinate axis direction, the motion state and trend of the target can be deeply understood. Based on the current motion parameters of the target, a reasonable extraction range is determined to construct a motion hypothesis space. This adaptive construction method based on the current motion state of the target makes the generated multiple motion hypotheses more conform to the actual possible motion of the target. When extracting velocity, acceleration, and direction angle, factors such as the motion characteristics and physical limitations of the target, such as the maximum steering angle and the amplitude of acceleration change, are fully considered to avoid generating unreasonable hypotheses. Thus, in the face of complex and diverse target motion scenarios, it is possible to more comprehensively and accurately cover the possible motion paths of the target, improve the accuracy and efficiency of multi-hypothesis tracking, enhance the adaptability and response ability to changes in the target motion state, ensure the accuracy and stability of target trajectory detection, and provide strong support for the aircraft to effectively track the target in a complex environment.

[0164] S2062. Use the newly received subsequent three-dimensional coordinates to verify the correlation of the motion hypothesis;

[0165] In some embodiments, an appropriate distance metric method (such as Euclidean distance) is used to calculate the error value between the predicted coordinates and the newly received actual coordinates. The smaller this error value, the higher the correlation between the motion hypothesis and the actual motion trajectory of the target.

[0166] S2063. Select the motion trajectory with the highest correlation as the trajectory data.

[0167] It should be noted that although a certain motion trajectory shows the highest correlation among the multiple constructed trajectories, if its correlation value is lower than the pre-set threshold, then at this time, the motion trajectory still needs to be corrected, or new motion trajectories need to be directly regenerated.

[0168] It can be seen that multiple motion trajectory hypotheses are constructed based on the characteristic information and the dynamic change data of the three-dimensional coordinates over time, considering various possibilities of target motion. Using the subsequent new three-dimensional coordinates for correlation verification can timely eliminate unreasonable hypotheses and ensure the accuracy of tracking. Selecting the trajectory with the highest correlation as the final trajectory data, this data-driven and probability-statistics-based method effectively avoids tracking loss caused by sudden acceleration, deceleration, turning, etc. of the target when the target motion state is complex and changeable, improves the stability and reliability of target tracking, and provides a trajectory basis for subsequent target recognition, decision-making, and guidance and strike operations.

[0169] S2064. In the case that the three-dimensional coordinates of the current target have not been updated within the preset time, perform feature comparison on the feature information of the subsequent target in the subsequent scanned image information with the feature information of the current target; the current target is any target, and the subsequent scanned image information is the subsequently obtained scanned image information of the scanned image information of the current target;

[0170] In some embodiments, when the aircraft finds that the three-dimensional coordinates of the current target have not been updated within a preset time, in order to avoid losing the target, it starts to use the feature information of the target for subsequent processing. First, extract the feature information of the current target from the previously stored target information, including color, shape, texture, and motion features, etc. Then, after obtaining the subsequent scan image information, extract the features of each potential target area in the image, and also obtain its feature information such as color, shape, and texture. Next, use a specific feature comparison algorithm to compare the feature information of the current target with the feature information of each potential target in the subsequent image one by one, and calculate the similarity between them. For example, for the color feature, the histogram intersection method can be used to calculate the similarity; for the shape feature, the shape context algorithm can be used to calculate the matching degree. Through these methods, find the subsequent targets with a higher feature similarity to the current target, providing a basis for further determining whether the target is lost and repositioning the target, ensuring that there is still a chance to retrieve the target through the feature information in the case of a short-term loss of the target, and maintaining the continuity and stability of target tracking.

[0171] S2065. Determine the lost target of the current target as the subsequent target with a feature comparison result greater than the similarity threshold and the highest similarity.

[0172] S2066. Stitch the trajectory data of the current target and the trajectory data of the lost target, where the missing trajectory data is supplemented using differences.

[0173] In some embodiments, after determining the lost target of the current target, it is necessary to stitch the trajectory data of the two to restore the complete motion trajectory of the target. First, analyze the last known position and state information of the current target before it is lost, and the first position and state information of the lost target after it reappears, including three-dimensional coordinates, speed, direction, etc. Then, according to the known motion characteristics of the target and the environmental information, estimate the possible motion path and state changes of the target during the loss period, and calculate the difference information of the missing trajectory part. For example, use the historical speed data of the target and the loss time to calculate the possible distance and direction change of the target during the loss period through kinematic formulas, so as to obtain a series of intermediate position points and state parameters for supplementing the missing trajectory. Next, supplement these differences to the trajectory data of the lost target, so that it can be reasonably connected with the trajectory data of the current target in time and space, forming a continuous and complete target motion trajectory.

[0174] It can be seen that when the target three-dimensional coordinates are not updated within the preset time, the lost target is searched for in the subsequent scanned images through feature comparison. By using the feature information of the target, these features are relatively stable and unique. Even if the target is temporarily lost, it can still be repositioned based on feature matching in the subsequent images. The subsequent target with a feature comparison result greater than the similarity threshold and the highest similarity is determined as the lost target, and the trajectory data is spliced, and the missing part is supplemented with the difference, so that the trajectory of the target can continue after being lost, ensuring the integrity and coherence of the target trajectory. This is particularly important for long-term and long-distance target tracking tasks, which can effectively avoid wasting all previous efforts due to the temporary loss of the target, ensure the stability and continuity of the entire tracking process, and provide complete and accurate trajectory information for subsequent target analysis and strike guidance.

[0175] S207. Input the feature information and trajectory data into the target recognition model to obtain the identity of the target;

[0176] In some embodiments, after obtaining the feature information and trajectory data of the target, they are input into a pre-trained target recognition model. The model first preprocesses the input feature information, such as operations like normalization, feature selection, or dimensionality reduction, to improve the quality of the data and the processing efficiency of the model. Then, according to the internal structure and algorithm of the model (such as the hierarchical structure of the neural network, the branching rules of the decision tree, etc.), feature extraction and pattern recognition are performed on the feature information and trajectory data. The model calculates the probability distribution of the target belonging to different categories or identities by comparing and matching with the feature patterns and trajectory patterns of various targets learned during the training phase. Finally, according to the set probability threshold or decision rule, the identity category of the target is determined, and the corresponding recognition result is output.

[0177] In some embodiments, the training process of the model includes:

[0178] S2071. Use the historical feature information, historical trajectory data, and the corresponding historical identity as the data set;

[0179] S2072. Adopt clustering processing, and cluster the historical feature information and historical trajectory data into several data subsets according to the historical identity;

[0180] In some embodiments, after obtaining a data set including historical feature information, historical trajectory data, and historical identities, clustering processing begins. First, a suitable clustering algorithm is selected. Common ones include the K-Means clustering algorithm, hierarchical clustering algorithm, etc. (here, it is assumed that the K-Means clustering algorithm is selected for illustration). Then, the key feature dimensions for clustering are determined, that is, which specific indicators in the historical feature information and historical trajectory data are used as the basis for judging similarity. For example, the shape features of the target, average speed, etc. can be selected as the feature dimensions. Next, according to the historical identity labels in the data set, the data related to the targets with the same historical identity is used as a reference for initial clustering, and the clustering algorithm is started for calculation. Taking the K-Means algorithm as an example, K clustering centers are randomly initialized first (the value of K is determined according to the known number of historical identity categories or the estimated number of categories), and then the distance from each data point (i.e., the point corresponding to each set of historical feature information and historical trajectory data) to these clustering centers is calculated (the distance metric can be selected as Euclidean distance, cosine distance, etc. according to the data characteristics), and the data points are assigned to the category represented by the nearest clustering center. After that, the clustering centers are continuously updated, and the process of assigning data points is repeated until the clustering centers no longer change significantly or reach the preset number of iterations. Finally, several data subsets clustered according to historical identities are formed, and these subsets can clearly reflect the data characteristics and trajectory features of different types of targets.

[0181] S2073. Train the target recognition model according to the data subsets in sequence, so that the target recognition model can identify historical identities based on historical feature information and historical trajectory data;

[0182] S2074. Train the target recognition model according to the data set, so that the target recognition model can identify multiple historical identities based on multiple historical feature information and multiple historical trajectory data.

[0183] In some embodiments, after the sequential training of the target recognition model based on the data subsets is completed, the target recognition model is then comprehensively trained using the entire dataset. The dataset is comprehensively preprocessed. Since the dataset contains data of multiple types of targets, there may be differences in data formats, magnitudes, etc., so unified processing is required. For example, all historical feature information is normalized so that its numerical range is within a suitable interval, and the historical trajectory data is regularized and aligned in time series to ensure that the trajectory data of different targets is comparable in the time dimension. Then, the processed dataset is divided into a training set and a validation set according to a certain ratio (such as the common 80% as the training set and 20% as the validation set). The training set is used for learning and adjusting the parameters of the model, and the validation set is used to evaluate the performance of the model on unseen data. Next, the historical feature information and historical trajectory data in the training set are input into the target recognition model, and the corresponding historical identities are used as correct labels for forward propagation calculation to obtain the prediction results of the model. The prediction error is measured by calculating a loss function (such as the mean squared error loss function, etc.), and then the backpropagation algorithm combined with an optimizer (such as the Adam optimizer, etc.) is used to adjust the parameters of the model according to the loss value. During the training process, the validation set is regularly used to validate the model, and the changes in evaluation metrics such as the accuracy rate and recall rate of the model on the validation set are observed. Based on these metrics, it is judged whether problems such as overfitting or underfitting occur in the model, and the training parameters (such as the learning rate, number of training epochs, etc.) are adjusted in a timely manner. The above training process is continuously carried out until the performance of the model on the validation set reaches the expected standard, such as the accuracy rate reaching more than 90%. At this time, the model can accurately identify multiple historical identities based on multiple historical feature information and multiple historical trajectory data, completing the training based on the entire dataset, enabling the model to have stronger generalization ability and comprehensive recognition ability for multiple targets in complex scenarios.

[0184] It can be seen that the historical feature information, historical trajectory data, and corresponding historical identities are integrated into a dataset and subjected to clustering processing. They are carefully clustered into multiple data subsets according to historical identities, which enables each subset to precisely cover the features and trajectory characteristics of specific types of targets. The target recognition model is trained for each subset respectively, and the model can deeply analyze and master the unique features and trajectory patterns of various targets, greatly improving the recognition accuracy and pertinence of single targets. Subsequently, the entire dataset is used for comprehensive training, enabling the model to quickly and accurately distinguish and identify the identities of each target based on the unique feature information and trajectory data of different targets when facing multi-target situations in complex scenarios. Even in complex situations where multiple targets are close to each other, occluded, or have different motion states, the probability of misjudgment and missed judgment can be effectively reduced, enhancing the accuracy and efficiency of multi-target recognition.

[0185] S208. Determine whether the target meets the tracking conditions according to the preset target list and the identity of the target;

[0186] In some embodiments, after obtaining the identity of the target, read the pre-stored target list and the set tracking conditions. Compare the identity of the target with the items in the target list to check whether the target belongs to the specified objects of interest in the list. At the same time, according to information such as the position and motion state of the target, combined with the set tracking conditions (such as whether the target enters a specific region of interest, whether it has specific motion characteristics, etc.), comprehensively determine whether the target meets the tracking requirements.

[0187] S209. If the target meets the tracking conditions, track the target;

[0188] In some embodiments, when it is determined that the target meets the tracking conditions, calculate the flight parameters that the aircraft needs to adjust, such as flight direction, speed, altitude, etc., according to the current position and motion state of the target, so that the aircraft can quickly approach the target and maintain within a suitable tracking distance and angle range.

[0189] In some embodiments, the following "target tracking and detection method for optoelectronic pods" can be used for tracking, specifically step S302. In this case, steps S202 and S209 are synchronized in time.

[0190] In other embodiments, before target recognition, step S102 of the above "control method for optoelectronic pods of unmanned aerial vehicles" can be used to obtain a scanned image.

[0191] S210. Output the position information of the target and the angular velocity information of the target to guide the execution of the strike process.

[0192] It can be seen that performing superpixel segmentation on the scanned image information to remove the background can effectively reduce the interference of irrelevant information and make subsequent target recognition more accurate and efficient. Then, calculate the three-dimensional coordinates by integrating the position and attitude of the sensor, the position and distance information of the target image, and lock the spatial position of the target. Obtain the target features and perform multi-hypothesis tracking in combination with the three-dimensional coordinates, which can handle complex motion states and ensure the stability and accuracy of target trajectory detection. Even if the target motion is variable, it is not easy to be lost. Input the features and trajectory data into the model to identify the target identity, and judge the tracking conditions according to the preset list to ensure that only key targets are tracked, avoiding waste of resources. Finally, output the target position and angular velocity information to guide the strike, improve the strike hit rate, and reduce the strike failure cases caused by inaccurate target positioning or tracking errors, overall improving the target tracking ability of autonomous or semi-autonomous aircraft in complex environments.

[0193] Please refer to Figure 3 , Figure 3is a flow chart of the target tracking and detection method of the optoelectronic pod in the embodiment of the present application; please refer to Figure 4 , Figure 4 It is a schematic diagram of a framework of a target tracking and detection method for an optoelectronic pod in an embodiment of the present application.

[0194] A target tracking and detection method for an optoelectronic pod, comprising:

[0195] S301, performing target detection on the video image to obtain a target detection result;

[0196] In some embodiments, a suitable target detection algorithm is selected. Common ones include deep learning-based target detection algorithms (such as the YOLO series, FasterR-CNN, etc.) and traditional feature-based target detection algorithms (such as Haar features + AdaBoost algorithm, etc.). Taking the YOLO algorithm as an example, the input video image is first preprocessed to adjust the image size, normalize the pixel values, etc., so that it meets the requirements of the algorithm input. Then, the processed image is sent to the pre-trained or trained YOLO network model. The network model will perform feature extraction, classification and positioning operations on the image, and finally output the target detection result through multiple convolutional layers and fully connected layers.

[0197] It should be noted that, in some embodiments, the above “target identification, tracking and guidance method of optoelectronic pod” can be used for detection. Specifically, steps S201 to S202, or S201 to S205, or S201 to S207;

[0198] In other embodiments, before target recognition, step S102 of the above “method for controlling an optoelectronic pod of a drone” may be used to acquire video images.

[0199] In some embodiments, step S301 specifically includes:

[0200] S3011, screening out a time-sensitive region from the video image;

[0201] In some embodiments, a preliminary motion analysis is performed on the video image, and the motion of pixels in the image is detected using algorithms such as the optical flow method. Areas where pixels move frequently and change greatly are screened out, because these areas are often where objects (possibly targets) are active and are time-sensitive areas.

[0202] S3012, extracting feature information of the time-sensitive area;

[0203] In some embodiments, for each time-sensitive region, a suitable combination of feature extraction methods is selected. For example, for color feature extraction, the image within the region can first be converted to a suitable color space (such as the RGB color space or the HSV color space, etc.), and then the color histogram in the corresponding color space is calculated to statistically analyze the pixel number distribution of different color components. For texture feature extraction, the method of gray-level co-occurrence matrix can be used, setting appropriate parameters (such as pixel distance, angle, etc.), calculating the co-occurrence relationship of pixel grayscales in different directions, and then obtaining relevant parameters reflecting texture characteristics. For shape feature extraction, an edge detection algorithm (such as the Canny edge detection algorithm) is used to obtain the edge contour of the object within the region, and then geometric parameters such as the perimeter, area, and circularity of the contour are calculated to describe its shape characteristics. At the same time, the spatial position features of the region (such as the center coordinates of the region, the upper-left and lower-right coordinates, etc.) are recorded. Then, these different types of feature information are sorted and stored to form a feature information set corresponding to each time-sensitive region for subsequent matching and comparison with the pre-stored target feature template.

[0204] S3013. Match and compare the feature information with the pre-stored target feature template;

[0205] In some embodiments, the feature information of the time-sensitive region and the feature information of the target feature template are organized into corresponding feature vector forms (assuming the feature vector of the time-sensitive region is, and the feature vector of the target feature template is, where is the dimension of the feature vector corresponding to different feature parameters). Then, the distance value between the two is calculated according to the formula for calculating the Euclidean distance. The smaller the distance value, the more similar the two are, and the greater the possibility that the time-sensitive region contains the target. Or, if the method of comparing the value ranges of feature parameters is adopted, the value-taking situations of the time-sensitive region and the target feature template in aspects such as the numerical values in each interval of the color histogram, shape geometric parameters (such as aspect ratio, perimeter, area, etc.), and texture feature parameters (such as contrast, correlation, etc.) are respectively compared to determine whether they are within a reasonable similarity range. For example, if the difference between the aspect ratio of the shape of the time-sensitive region and the aspect ratio of the target feature template is within a certain threshold, it is considered that the two are relatively matched in terms of shape features.

[0206] S3014. Determine the time-sensitive regions with matching results greater than the threshold as the target.

[0207] It can be seen that by screening out the time-sensitive regions from the video image, extracting their feature information for matching and comparison with the pre-stored target feature template, and determining the regions with matching results greater than the threshold as the target. The screening of time-sensitive regions can quickly focus on the key regions where target changes or new targets may appear, avoiding a comprehensive and time-consuming detection of the entire video image. Through the matching and comparison with the pre-stored template, using the existing target feature knowledge, the target can be efficiently identified, reducing the computational amount and time consumption of ineffective detection.

[0208] S302. During target detection, track the targets in the video image to obtain the target tracking results;

[0209] It should be noted that in some embodiments, the above "Target Recognition, Tracking, and Guidance Method for Electro-optical Pods" can be used for tracking. Specifically, it is steps S203 to S209, or S206 to S209, or S208 to S209;

[0210] In other embodiments, before target recognition, step S102 of the above "Control Method for Electro-optical Pods of Unmanned Aerial Vehicles" can be used to obtain the video image.

[0211] In some embodiments, according to the target initial position information obtained in the target detection step, determine the initial region of the target in the video image, and extract the features of the region (such as HOG features, etc.) as the feature representation of the target. Then, in subsequent video image frames, use a certain range where the target may appear (usually estimated according to the motion characteristics and prior knowledge of the target) as the search region, and use the correlation filtering algorithm to calculate the position in the search region that is most similar to the target features, determine the position of the target in the current frame, record the position information and information such as the speed and direction calculated by comparing with the position of the previous frame, form the target tracking results, and continuously update these results as the video image frames are continuously input to achieve continuous tracking of the target.

[0212] It should be clear that in the method adopted in this embodiment, the two operation processes of step S301 and step S302 are independent and carried out in parallel, so that target detection and target tracking can be started simultaneously.

[0213] In some specific embodiments, step S302 specifically includes: S30210. Extract features from adjacent frame images to obtain the position information of the target in adjacent frames;

[0214] It should be noted that feature extraction has been discussed in detail in step S3012, and the relevant steps can refer to step S3013, which will not be elaborated here.

[0215] S30211. Obtain the position change information according to the position information;

[0216] In some embodiments, sequentially take out the position coordinate information of the target in adjacent frames in chronological order, calculate the position change amounts in the horizontal and vertical directions respectively, that is, calculate the horizontal displacement and vertical displacement, and record these two displacement change amounts, which together constitute the position change information of the target between these two frames.

[0217] S30212. Determine the motion speed and motion direction of the target according to the position change information;

[0218] In some embodiments, the time interval between adjacent frames is determined according to the frame rate of the video image. Then, for the position change information (i.e., horizontal displacement and vertical displacement) between each pair of adjacent frames, the instantaneous velocity of the target is calculated by calculating the velocity components in the horizontal and vertical directions and using the Pythagorean theorem, so as to obtain the motion velocity of the target at different adjacent frame stages. Next, by calculating the angle between the target motion direction and the positive direction of the horizontal axis and using the arctangent function, the motion direction of the target at each stage is determined.

[0219] S30213. Predict the position range where the target will appear in the next frame according to the motion velocity and motion direction;

[0220] In some embodiments, according to the motion velocity and motion direction of the target, the theoretically displacement amount of the target within the time interval between adjacent frames is calculated. Based on the position coordinates of the target in the current frame, the theoretically central coordinate position of the target in the next frame is calculated, and a suitable rectangular or elliptical area is constructed as the position range where the target will appear in the next frame.

[0221] S30214. Perform target matching and tracking according to the texture intensity and phase characteristics of the target within the position range.

[0222] It can be seen that feature extraction is performed on adjacent frame images to obtain target position information, and then the position change information is obtained to determine the target motion velocity and direction. Next, the position range where the target will appear in the next frame is predicted, and target matching and tracking are performed in combination with the texture intensity and phase characteristics of the target. Through the feature analysis of adjacent frames and the prediction of the target motion state, the area where the target may appear can be planned in advance, reducing the blind search in the entire image space and improving the tracking efficiency. At the same time, using the texture intensity and phase characteristics for matching and tracking enhances the accuracy and stability of target recognition, especially when the appearance and pose of the target change, the target can be more accurately located.

[0223] In the actual use process, the computational amount for tracking the target in the video image is too large; therefore, in some embodiments, step S302 specifically includes:

[0224] S3021. Divide the video image into several regions;

[0225] In some embodiments, the uniform grid division method is used, that is, according to the preset number of rows and columns, by calculating the size of the image, it is divided in the horizontal and vertical directions at corresponding intervals to obtain small rectangular regions. Another way can be content-aware division. First, perform simple preprocessing on the image such as edge detection and object recognition, and then divide regions of different shapes and sizes according to the detected object contours or different semantic regions (such as after distinguishing the foreground object and the background region, further subdividing the range where the foreground object is located, etc.).

[0226] S3022. Compare the statistical characteristics of each region with the target feature range;

[0227] In some embodiments, the feature comparison method has been discussed in detail in step S3013 and will not be elaborated here.

[0228] S3023. If the statistical characteristics of the region do not belong to the target feature range, then perform constant false alarm detection technology to adjust the target feature range;

[0229] In some embodiments, when the statistical characteristics of a certain region do not belong to the target feature range after comparison, it is necessary to use constant false alarm detection technology to adjust the target feature range. Then, according to the set false alarm probability (generally preset according to the actual application scenario and the tolerance for misjudgment, such as setting the false alarm probability to 0.05), using constant false alarm detection algorithms (common ones include cell-averaging constant false alarm detection algorithm, ordered statistic constant false alarm detection algorithm, etc.), based on the pixel feature distribution within the region, adjust the relevant parameters in the target feature range. For example, if it is the cell-averaging constant false alarm detection algorithm, the region will first be divided into multiple small cells, the average value of pixel features in each cell will be statistically calculated, and then an appropriate adjustment threshold will be calculated according to the overall feature distribution and the false alarm probability, and then parameters such as the target color feature range and shape feature range will be modified accordingly, so that the target feature range can better adapt to the actual situation of this region, in order to determine again whether this region may contain the target later and improve the accuracy of target detection and tracking.

[0230] S3024. If the statistical characteristics of the region do not belong to the adjusted target feature range;

[0231] In some embodiments, after adjusting the target feature range by constant false alarm detection technology, it is necessary to compare the statistical characteristics of the region with the adjusted target feature range again to further confirm whether this region may contain the target. The operation method of this step is similar to the basic process of comparison in step S3022 and will not be elaborated here.

[0232] S3025. Then determine the region as a non-target region and perform suppression processing.

[0233] In some embodiments, when it is determined through step S3024 that the statistical characteristics of a region do not belong to the adjusted target feature range, it can be determined that the region is a non-target region and suppression processing is performed. First, a specific processing method is selected according to the set suppression processing strategy. For example, if the method of setting the region pixel value to zero is selected, all pixel points within the region are traversed and their pixel values are modified to 0, making it visually black to achieve the purpose of eliminating the influence of the region; if the method of reducing the pixel weight is selected, a reasonable weight reduction coefficient needs to be set (such as setting it to 0.1, indicating that the region pixel weight is reduced to one-tenth of the original). Then, according to the coordinate information of the region, in subsequent calculations related to target tracking, the pixels within the region are weighted according to the set weight reduction coefficient, so that the influence of the region in the overall calculation is greatly reduced.

[0234] It can be seen that by dividing the video image into several regions and comparing them with the target feature range, and using the constant false alarm detection technology to adjust the target feature range for non-conforming regions, this method of regional processing and dynamic adjustment reduces unnecessary computational complexity. Because there is no need to perform comprehensive and complex calculations on the entire image, but to focus on the regions where the target may exist and the precise definition of the target feature range, avoiding the high computational complexity of processing the entire image without discrimination like the TLD algorithm. Thus, target tracking can be achieved more efficiently, reducing the demand for computing resources, improving the real-time performance of the algorithm, ensuring that the target can be tracked quickly and accurately during the operation of the aircraft, and adapting to complex and changing actual scenarios.

[0235] In some embodiments, after step S3025, the method further includes:

[0236] S3026. Determine the regions other than the non-target regions as potential target regions;

[0237] S3027. Score the potential target regions using a random fern classifier;

[0238] In some embodiments, to ensure that the random fern classifier has been adequately trained, the training data should include a large number of target samples and non-target samples, as well as their corresponding various image feature information. Through these data, the classifier learns the feature patterns of targets and non-targets. Then, for each potential target region, multiple feature information is extracted, such as calculating the color histogram of pixels within the region as the color feature, calculating the texture feature through the gray-level co-occurrence matrix, and extracting the shape feature using the edge detection algorithm, etc. These feature information are input into the trained random fern classifier. The classifier will comprehensively evaluate the region according to its internal decision function and the learned feature patterns, and output a score value indicating the likelihood that the region contains the target. This score value will serve as an important basis for subsequent determination of the tracking region, enabling the target tracking to more precisely focus on the region most likely to contain the target, and improving the accuracy and efficiency of the tracking.

[0239] S3028. Select the region of the preset data with the highest score as the tracking region;

[0240] In some embodiments, to sort the scores of all potential target regions, common sorting algorithms (such as the quicksort algorithm) can be used to arrange the potential target regions in descending order of scores. Then, according to the pre-set selection rules (such as selecting the top N regions, or selecting regions with scores higher than a certain threshold), the corresponding regions are extracted from the sorted region list, and these regions are the determined tracking regions.

[0241] S3029. Perform target tracking on the tracking region.

[0242] It can be seen that splitting the video image into several regions and comparing them with the target feature range, and using the constant false alarm detection technology to adjust the target feature range for the non-conforming regions. This way of regional processing and dynamic adjustment reduces unnecessary computational complexity. Because there is no need to perform comprehensive and complex calculations on the entire image, but focus on the regions where the target may exist and the precise definition of the target feature range, avoiding the high computational complexity of processing the entire image without discrimination like the TLD algorithm. Thus, the target tracking can be achieved more efficiently, reducing the demand for computing resources, improving the real-time performance of the algorithm, ensuring that the target can be tracked quickly and accurately during the operation of the aircraft, and adapting to complex and changeable actual scenarios.

[0243] In some embodiments, after step S306, the method further includes:

[0244] S308. In the case where the target with the highest confidence score is lower than the score threshold or there is no target; extract the feature data of the target in the most recent frame of the video image;

[0245] In some embodiments, after the confidence scores of all targets are completed, if it is found that the target with the highest confidence score has a score lower than the set score threshold, or if no target is detected in the video image at all, it is necessary to perform the step of extracting the feature data of the target in the most recent frame of the video image.

[0246] It should be noted that feature extraction has been discussed in detail in step S3012. For related steps, reference can be made to step S3013, and details will not be repeated here.

[0247] S309: Determine the search area based on the target speed and the target's position information;

[0248] In some embodiments, according to the target speed, combined with the frame rate of the video image and a certain time prediction step length, calculate the possible distance that the target may move within this period of time. Then, with the current position information of the target as the center, according to the calculated possible moving distance, expand a certain range around to determine the search area. For example, the search area can be set as a circular area with the target center coordinates as the center of the circle and a radius equal to the possible moving distance plus a certain margin, or set as a square area with the target position as the center and a side length equal to twice the possible moving distance, etc., so as to determine a reasonable search area to facilitate further narrowing the range and searching for the target in the subsequent steps.

[0249] S310: Narrow down the search area according to the movement direction of the target to obtain the region of interest;

[0250] In some embodiments, determine the corresponding angular range according to the movement direction of the target. For a circular search area, according to the angular range, determine the region of interest through the sector area corresponding to the central angle, remove the sector parts in the circle that do not conform to the movement direction, and retain the sector area where the target may appear as the region of interest.

[0251] S311: Extract the standby target in the subsequent video image according to the coordinates of the region of interest;

[0252] In some embodiments, the coordinate information of the region of interest in the video image is obtained, including the upper left corner coordinates and the lower right corner coordinates, so as to determine the specific position of the region of interest in the entire video image. Then, for subsequent video image frames (which can be the next frame or several subsequent frames, determined according to the actual situation), in these frame images, target detection operations are only performed on the pixel range corresponding to the region of interest. Using a suitable target detection algorithm, the image data within the region of interest is input into the target detection algorithm, and the algorithm will output the possible target information detected within this region. These detected targets are used as standby targets, and their relevant features are extracted at the same time. The standby targets and their feature information are sorted out to prepare for subsequent confidence score calculation, so as to further distinguish whether these standby targets are real targets.

[0253] S312. Perform confidence score calculation on the standby targets and the targets;

[0254] In some embodiments, relevant feature information of the original target needs to be obtained, including previously accumulated appearance features (such as color, shape, texture features, etc.), position features (such as the movement trajectory of the target, the position coordinates before disappearance, etc.), and the previous confidence score situation, etc. These information serve as the basic data for comparison. Then, for each standby target, its corresponding appearance, position and other features are also extracted. Using a suitable similarity calculation method, for example, for appearance features, the structural similarity index (SSIM) can be used to calculate the similarity with the original target, and for position features, quantitative evaluation can be carried out according to the distance deviation from the predicted position of the original target (such as calculating the Euclidean distance). Combining these similarity evaluation results, with a certain weight assignment (for example, setting the appearance feature similarity weight to 0.6 and the position feature similarity weight to 0.4 according to experience or experimental data), the confidence score of each standby target is calculated by means of weighted summation, etc., and these scores are recorded for subsequent judgment on whether the standby targets meet the requirements, so as to retrieve the possibly lost targets and ensure the continuity and accuracy of target tracking.

[0255] S313. Determine the standby targets with confidence scores higher than the score threshold as the lost targets.

[0256] It can be seen that when the target with the highest confidence score is lower than the score threshold or there is no target, by extracting the target feature data of the most recent frame, the search area is determined based on the target speed and position information, and the region of interest is obtained by narrowing it down according to the movement direction. Standby targets are extracted in this region and confidence scores are calculated to determine the lost targets. This method avoids the aimless search and calculation in the entire image space like the TLD algorithm, but instead narrows the search range in a targeted manner based on the movement characteristics of the target, reducing the computational amount. Ensure that the aircraft can still efficiently re-locate the target when the target is temporarily lost or difficult to detect.

[0257] It should be noted that in actual application scenarios, when it comes to determining whether a backup target is a truly lost target, the judgment of the response value is a crucial link. Since the target is extremely vulnerable to occlusion interference from background elements (such as other objects in the environment, objects similar in appearance to the target, etc.) in the actual environment, this poses higher requirements for the algorithm, that is, it needs to be able to accurately determine the occlusion situation.

[0258] Currently, the maximum response value obtained from related operations is usually directly used as the basis for judging the tracking confidence, so as to determine whether the target is lost or to determine whether the filter template needs to be updated for the current frame. However, from the actual situation, the related responses of the filter often do not show a standard Gaussian distribution form. This means that in the actual tracking process, there may be a situation where the maximum response value is at a relatively low level, but the target is actually still in the correct tracking state; conversely, there may also be an abnormal situation where the maximum response value is relatively high, but the target has been lost. Therefore, in some embodiments, after step S313, the method further includes:

[0259] S314. Calculate the average peak correlation energy for the backup target with a confidence score higher than the score threshold;

[0260] S315. If the average peak energy is less than the preset energy threshold, cancel the determination that the backup target with a confidence score higher than the score threshold is a lost target;

[0261] The calculation function of the average peak correlation energy is:

[0262]

[0263] In the formula, APCE is the average peak correlation energy, is the maximum response value, is the minimum response value, is the response value of pixel point w and pixel point h.

[0264] It can be seen that after determining the lost target, calculate its average peak correlation energy and compare it with the preset energy threshold to further confirm the authenticity of the lost target. This avoids learning a large amount of interference information. When the target is occluded by the background or interfered by similar targets, it can more accurately determine whether the target is truly lost, improves the accuracy and stability of tracking, ensures the continuous and effective tracking of the target by the aircraft in a complex environment, and reduces the tracking failure caused by misjudgment.

[0265] S303. Obtain positive samples and negative samples according to the target detection result, target tracking result, and video image. The positive samples are the data containing the target in the target detection result and target tracking result, and the negative samples are the data not containing the target in the video image;

[0266] In some embodiments, after obtaining the target detection result and the target tracking result, the acquisition of positive and negative samples begins. First, according to the position coordinate information of the target in the target detection result, the regional image data where the target is located is cropped from the video image. These cropped data are part of the positive samples. At the same time, the feature descriptions of the target within the region (such as color feature vectors, texture feature vectors, etc.) are extracted and together with the cropped image data form complete positive samples. Then, for the negative samples, the other regions in the video image except the region where the target is located are traversed, and part of the regional image data is selected as negative samples according to certain rules (such as uniform sampling, random sampling, etc.), or according to the characteristics of the target such as size and shape, the regions similar to the target features are excluded, and more representative non-target regional image data is selected as negative samples to ensure that the negative samples can fully reflect the characteristics of the background and non-target objects. The obtained positive and negative samples are sorted out and prepared for subsequent model training.

[0267] S304. Train the target detection update model according to the positive and negative samples;

[0268] In some embodiments, a suitable model structure needs to be selected. Commonly, it can be a structure based on a convolutional neural network (CNN). For example, classic network structures such as ResNet can be used as the basis, and then according to the actual requirements of the target detection task, some layers are added or modified on this basis (such as adding an output layer to output the target detection result, etc.). Then, the positive and negative samples are divided into a training set and a validation set according to a certain ratio (such as 80% positive samples and 20% negative samples, and the specific ratio can be adjusted according to the actual situation). The training set is used for the parameter learning of the model, and the validation set is used to evaluate the performance of the model on unseen data. The positive and negative sample image data in the training set are input into the model, and the prediction result of the model for the target is obtained through forward propagation calculation. Then, the difference between the prediction result and the true label (the positive sample is the target, and the negative sample is the non-target) is measured by calculating the loss function (such as the cross-entropy loss function, etc.). According to the loss value, the parameters of the model are adjusted using the backpropagation algorithm, such as the convolution kernel weights of the convolutional layer and the connection weights of the fully connected layer. The data in the training set is iteratively trained for multiple rounds, and at the same time, the performance of the model is verified regularly using the validation set, and the training parameters (such as the learning rate, etc.) are adjusted according to the verification result until the performance of the model on the validation set reaches the expected standard (such as the accuracy rate reaches a certain value, etc.), and the training and update of the model are completed.

[0269] S305. The target detection update model outputs feature data and a target model. The feature data is used to perform target detection on a new video image when receiving the new video image, and the target model is used to track the target of the new video image when receiving the new video image;

[0270] In some embodiments, after the target detection update model is trained, it can process newly input video images and output feature data and a target model. First, for a new video image, the model extracts features of the image according to its internal feature extraction layers (such as convolutional layers) to obtain a feature map of the image, and then processes the feature map through specific algorithms and layers (such as a combination of fully connected layers or pooling layers) to convert it into feature data that can be directly used for target detection. These feature data can be output in the form of vectors or matrices for subsequent target detection operations. At the same time, the model constructs a target model based on information such as the motion law and appearance change pattern of the target learned during the previous training process. This target model contains the feature parameters and motion parameters of the target in different states and can predict the position, pose, etc. of the target at the next moment according to the current state of the target when tracking the target in a new video image, so as to achieve continuous tracking of the target, and output the feature data and the target model for application in subsequent target detection and tracking processes.

[0271] S306. Perform a confidence score based on the target detection result and the target tracking result;

[0272] In some embodiments, a scoring method based on weighted summation is used. Obtain the detection confidence score of the target from the target detection result. For example, use the target category probability output by the deep learning target detection model as the detection confidence score. Calculate the tracking stability score of the target from the target tracking result. By calculating the variance of the position coordinates of the target in consecutive frames, if the variance is less than the set threshold, the tracking stability score is higher, otherwise it is lower. At the same time, calculate the consistency score of the target appearance features, and use an image feature extraction algorithm (such as SIFT features) to calculate the feature similarity of the target in different frames. The higher the similarity, the higher the score. Then, set weights for these three metrics and calculate the confidence score for each target.

[0273] S307. Determine the target with the highest confidence score as the correct target.

[0274] It can be seen that performing tracking while conducting target detection shortens the latency between detection and tracking. After obtaining the target detection results and target tracking results, positive and negative samples are acquired based on them and the video images. This method can accurately collect data of targets and non-targets, making the training samples more targeted and diverse. Then, the positive and negative samples are used to train the updated model for target detection, enabling the model to learn the changing characteristics of the target in real time and output more accurate feature data and target models for the detection and tracking of new video images. Finally, the correct target is determined through confidence scoring. Through the interaction of these technical features, the tracking and detection accuracy of the autonomous or semi-autonomous aircraft in complex and changing environments is improved, effectively solving the problems of difficult accurate tracking and easy target loss caused by target changes, and enhancing the stability and reliability of target tracking and detection.

[0275] The following introduces the exemplary aerospace detection device 500 provided by the embodiments of the present application. Figure 5 It is a schematic diagram of the exemplary hardware structure of the aerospace detection device 500 provided by the embodiments of the present application.

[0276] In some embodiments, the aerospace detection device 500 is a computer device or the aerospace detection device 500 includes a computer device. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with other external terminals or servers through a network connection. In some embodiments, the network interface can be a wired network interface, and in some embodiments, the network interface can also be a wireless network interface. When the computer program is executed by the processor, it realizes the method in the embodiments of the present application.

[0277] Those skilled in the art can understand that Figure 5 the structure shown in

[0278] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present application.

[0279] In the above embodiments, depending on the context, the term "when..." can be interpreted to mean "if...", or "after...", or "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if detecting (the stated condition or event)" can be interpreted to mean "if determining...", or "in response to determining...", or "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".

[0280] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media (such as solid-state drives), etc.

[0281] Those of ordinary skill in the art can understand all or part of the processes in the above method embodiments. The processes can be completed by relevant hardware instructed by a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage media include: ROM or random access memory RAM, magnetic disks, or optical disks, etc., which can store program codes.

Claims

1. A target tracking and detection method for an optoelectronic pod, characterized in that: include: Perform target detection on the video image to obtain the target detection result; During target detection, the target in the video image is tracked to obtain a target tracking result; wherein: the video image is divided into a plurality of regions; comparing the statistical characteristics of each of said regions to a target characteristic range; If the statistical characteristics of the area do not belong to the target feature range, a constant false alarm detection technology is performed to adjust the target feature range; If the statistical characteristics of the region do not fall within the adjusted target characteristic range; The area is determined as a non-target area and is subjected to suppression processing; determining an area other than the non-target area as a potential target area; The potential target area is scored using a random fern classifier; Select the area with the highest-scoring preset data as the tracking area; Performing target tracking on the tracking area; Acquire positive samples and negative samples according to the target detection result, the target tracking result, and the video image, wherein the positive sample is data containing the target in the target detection result and the target tracking result, and the negative sample is data not containing the target in the video image; Training the target detection update model according to the positive samples and negative samples; The target detection update model outputs feature data and a target model, wherein the feature data is used to perform target detection on the new video image when a new video image is received, and the target model is used to track the target of the new video image when the new video image is received; Performing confidence scoring according to the target detection result and the target tracking result; The target with the highest confidence score is identified as the correct target.

2. The method according to claim 1, characterized in that After the step of performing confidence scoring according to the target detection result and the target tracking result, the method further includes: When the target with the highest confidence score is lower than the score threshold or there is no target; extracting feature data of the target in the most recent frame of the video image; Determine a search area based on the target speed and the target position information; The search area is narrowed down according to the moving direction of the target to obtain an area of ​​interest; extracting a backup target in subsequent video images according to the coordinates of the region of interest; Performing confidence scoring on the backup target and the target; The backup targets having a confidence score higher than a score threshold are determined as lost targets.

3. The method according to claim 2, characterized in that After the step of determining the backup target having a confidence score higher than a score threshold as a lost target, the method further comprises: Calculate the average peak correlation energy of the backup targets whose confidence scores are higher than a score threshold; If the average peak energy is less than a preset energy threshold, the backup target with a confidence score higher than the score threshold is cancelled and determined as a lost target; The calculation function of the average peak correlation energy is: ; Wherein, APCE is the average peak correlation energy, is the maximum response value, is the minimum response value, is the response value of pixel w and pixel h.

4. The method according to claim 1, characterized in that The step of performing target detection on the video image to obtain the target detection result specifically includes: Filtering out a time-sensitive region from the video image; Extracting feature information of the time-sensitive area; Matching and comparing the feature information with a pre-stored target feature template; The time-sensitive area whose matching result is greater than the threshold is determined as the target.

5. The method according to claim 1, characterized in that The step of tracking the target in the video image to obtain the target tracking result during target detection specifically includes: Extract features from adjacent frame images to obtain the location information of the target in the adjacent frames; Obtaining position change information according to the position information; Determine the moving speed and moving direction of the target according to the position change information; Predicting a position range where the target will appear in the next frame according to the movement speed and the movement direction; Target matching and tracking are performed based on the texture intensity and phase characteristics of the target within the position range.

6. An aerospace detection device, characterized in that: The aerospace detection device comprises: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code comprises computer instructions, and the one or more processors call the computer instructions to enable the aerospace detection device to perform the method as described in any one of claims 1-5.

7. A computer program product comprising instructions, characterized in that When the computer program product is run on an aerospace detection device, the aerospace detection device is caused to perform the method according to any one of claims 1 to 5.

8. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on an aerospace detection device, the aerospace detection device is caused to perform the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Target tracking method and device, electronic equipment and storage medium

    CN110610510A

  • ROS-based airborne target detection and tracking method

    CN112164095A