A target identification, tracking and guidance method for an optoelectronic pod, a detection device, a program product and a storage medium
By superpixel segmentation and three-dimensional coordinate calculation of scanned image information, and multi-assumption tracking combined with feature information, the accuracy and stability problems of target tracking in complex environments are solved, and efficient and accurate target tracking is achieved.
Patent Information
- Application Number
- CN202510191721.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The prior art is difficult to achieve high-precision and stable target tracking in complex environments, especially when the target motion state suddenly changes.
By superpixel segmenting the scanned image information to remove the background, calculate the three-dimensional coordinates of the target in the three-dimensional space, and perform multi-assumption tracking based on feature information, input the target recognition model to judge the tracking conditions.
High-precision and stable tracking of the target in complex environments is achieved, tracking loss caused by changing target movements is reduced, and hitting hit rate and resource utilization efficiency are improved.
Smart Images

Figure CN119672325B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the general field of image data processing or generation technology, and in particular to a target recognition, tracking and guidance method, detection equipment, program product and storage medium for an optoelectronic pod. Background Art
[0002] In the civilian field, drones and tethered drones play a key role. The aerospace detection equipment they carry, such as optoelectronic seekers, optoelectronic pods, and three-light hybrid pods, are crucial for accurate detection of targets.
[0003] The target recognition, tracking and guidance method in the related art collects images of the target area through various sensors and extracts possible target areas. The target is identified by extracting some features of the target and comparing them with the pre-stored target feature library. When tracking the target, its position at the next moment is predicted, and the flight attitude of the aircraft and the direction of the detection equipment are adjusted accordingly to keep tracking the target.
[0004] However, this method has difficulty accurately predicting the target's position when the target's motion state changes suddenly, such as rapid turning, acceleration or deceleration, which can easily lead to tracking loss. As time goes by and the task continues, the calculation of the target's position will become increasingly inaccurate.
[0005] In summary, relevant technologies cannot track targets with high precision and stability in complex environments. Summary of the invention
[0006] The present application provides a target identification, tracking and guidance method, detection equipment, program product and storage medium for an optoelectronic pod, which are used to track targets with high precision and stability in complex environments.
[0007] In the first aspect, the present application provides a target recognition, tracking and guidance method for an optoelectronic pod, comprising: performing superpixel segmentation on the acquired scanning image information, removing the background to obtain a superpixel area; performing target recognition on the superpixel area to identify the target; obtaining the original position and posture information of the sensor; calculating the three-dimensional coordinates of the target in three-dimensional space based on the position information of the target in the scanning image information, using the distance information from the target to the sensor, the original position and posture information; obtaining feature information of each target; performing target trajectory detection based on multi-hypothesis tracking based on the feature information and the three-dimensional coordinates to obtain trajectory data; inputting the feature information and trajectory data into a target recognition model to obtain the identity of the target; judging whether the target meets the tracking conditions based on a preset target list and the identity of the target; if the target meets the tracking conditions, tracking the target; outputting the target's position information and the target's angular velocity information to guide the execution of the strike process.
[0008] By adopting the above technical solution, the scanned image information is segmented into superpixels to remove the background, which can effectively reduce the interference of irrelevant information and make subsequent target recognition more accurate and efficient. Then, the three-dimensional coordinates are calculated by integrating the sensor position and posture, the target image position and the distance information to lock the target spatial position. Obtaining the target features and combining the three-dimensional coordinates for multi-hypothesis tracking can cope with complex motion states and ensure the stability and accuracy of target trajectory detection, even if the target motion is changeable, it is not easy to lose. Input the feature and trajectory data into the model to identify the target identity, judge the tracking conditions according to the preset list, and ensure that only key targets are tracked to avoid resource waste. Finally, the target position and angular velocity information are output to guide the strike, improve the strike hit rate, reduce the strike failure caused by inaccurate target positioning or tracking errors, and overall improve the target tracking capability of autonomous or semi-autonomous aircraft in complex environments.
[0009] In combination with some embodiments of the first aspect, in some embodiments, target trajectory detection based on multi-hypothesis tracking is performed according to feature information and three-dimensional coordinates to obtain the step of trajectory data, specifically including: constructing multiple different motion trajectories for each target according to the dynamic change data generated by the feature information and three-dimensional coordinates over time; using the new three-dimensional coordinates received subsequently to verify the correlation of the motion hypothesis; and selecting the motion trajectory with the highest correlation as the trajectory data.
[0010] By adopting the above technical solution, multiple motion trajectory hypotheses are constructed based on feature information and dynamic changes of three-dimensional coordinates over time, taking into account various possibilities of target movement. Using subsequent new three-dimensional coordinates for correlation verification can timely eliminate unreasonable assumptions and ensure the accuracy of tracking. The trajectory with the highest correlation is selected as the final trajectory data. This method based on data-driven and probability statistics can effectively avoid tracking loss caused by sudden acceleration, deceleration, turning, etc. of the target when the target motion state is complex and changeable, improve the stability and reliability of target tracking, and provide a trajectory basis for subsequent target identification, decision-making and guidance operations.
[0011] In combination with some embodiments of the first aspect, in some embodiments, after the step of selecting the motion trajectory with the highest correlation as the trajectory data, the method also includes: when the three-dimensional coordinates of the current target are not updated within a preset time, feature comparison is performed on the feature information of the subsequent target in the subsequent scanned image information according to the feature information of the current target; the current target is any target, and the subsequent scanned image information is the scanned image information subsequently obtained of the scanned image information of the current target; the subsequent target whose feature comparison result is greater than the similarity threshold and has the highest similarity is determined as the lost target of the current target; the trajectory data of the current target and the trajectory data of the lost target are spliced, wherein the missing trajectory data is supplemented by the difference.
[0012] By adopting the above technical solution, when the three-dimensional coordinates of the target are not updated within the preset time, the lost target is found in the subsequent scanned images through feature matching. Using the feature information of the target, these features are relatively stable and unique. Even if the target is temporarily lost, it can still be relocated in the subsequent images based on feature matching. The subsequent target with the highest similarity and feature comparison result greater than the similarity threshold is determined as the lost target, and the trajectory data is spliced, and the missing part is supplemented with the difference, so that the trajectory of the target can be continued after it is lost, ensuring the integrity and continuity of the target trajectory. This is particularly important for long-term and long-distance target tracking tasks. It can effectively avoid the loss of previous efforts due to the temporary loss of the target, ensure the stability and continuity of the entire tracking process, and provide complete and accurate trajectory information for subsequent target analysis and strike guidance.
[0013] In combination with some embodiments of the first aspect, in some embodiments, based on feature information and dynamic change data generated by three-dimensional coordinates over time, the steps of constructing multiple different motion trajectories for each target specifically include: calculating the velocity components of the target in the directions of each coordinate axis according to the changes in the three-dimensional coordinates of the target at consecutive time points; calculating the acceleration components of the target in the directions of each coordinate axis according to the changes in the velocity components at different times; constructing a motion hypothesis space by changing the acceleration, speed and direction angle; extracting the next motion state parameters of the target from the motion hypothesis space to generate multiple motion hypotheses that conform to the probability distribution, the motion state parameters include acceleration, speed, and direction angle, wherein when extracting the speed, the current speed of the target is used as the midpoint, and the speed change amount generated by the current acceleration within a preset time step is used as the radius to determine the extraction range; when extracting the acceleration, the current acceleration of the target is used as the midpoint, and the acceleration change amplitude is used as the radius to determine the extraction range; when extracting the direction angle, the current direction angle of the target is used as the midpoint, and the maximum steering angle is used as the radius to determine the extraction range.
[0014] By adopting the above technical solution, the target's motion state and trend can be deeply understood by calculating the target's velocity and acceleration components in the direction of each coordinate axis. Based on the target's current motion parameters, the extraction range is reasonably determined to construct the motion hypothesis space. This adaptive construction method based on the target's current motion state makes the generated multiple motion hypotheses more in line with the target's actual possible motion conditions. When extracting velocity, acceleration and direction angle, the target's motion characteristics and physical limitations, such as the maximum steering angle, acceleration change amplitude and other factors, are fully considered to avoid generating unreasonable hypotheses. Therefore, when facing complex and diverse target motion scenes, it can more comprehensively and accurately cover the target's possible motion paths, improve the accuracy and efficiency of multi-hypothesis tracking, enhance the adaptability and responsiveness to changes in the target's motion state, ensure the accuracy and stability of target trajectory detection, and provide strong support for the aircraft to effectively track targets in complex environments.
[0015] In combination with some embodiments of the first aspect, in some embodiments, before the step of inputting feature information and trajectory data into a target recognition model to obtain the identity of the target, the method also includes: taking historical feature information, historical trajectory data and corresponding historical identities as a data set; using clustering processing to cluster the historical feature information and historical trajectory data into several data subsets according to the historical identities; training the target recognition model according to the data subsets in sequence, so that the target recognition model can identify the historical identity according to the historical feature information and historical trajectory data; training the target recognition model according to the data set, so that the target recognition model can identify multiple historical identities according to multiple historical feature information and multiple historical trajectory data.
[0016] By adopting the above technical solution, historical feature information, historical trajectory data and corresponding historical identities are integrated into a data set and clustered. The data is carefully clustered into multiple data subsets according to the historical identity, so that each subset accurately covers the characteristics and trajectory characteristics of a specific type of target. The target recognition model is trained for each subset separately. The model can deeply analyze and master the unique characteristics and trajectory patterns of each type of target, greatly improving the recognition accuracy and pertinence of a single target. Subsequently, the entire data set is used for comprehensive and integrated training, so that the model can quickly and accurately distinguish and identify the identity of each target based on the unique feature information and trajectory data of each target when facing multiple targets in complex scenes. Even in complex situations where multiple targets are close to each other, occluded or in different motion states, it can effectively reduce the probability of misjudgment and missed judgment, and enhance the accuracy and efficiency of multi-target recognition.
[0017] In combination with some embodiments of the first aspect, in some embodiments, target recognition is performed on a superpixel area, and the steps of identifying the target specifically include: calculating the gradient amplitude and direction of each pixel point in the superpixel area, the gradient amplitude is the intensity of the grayscale change of the pixel point, and the direction is the direction of the fastest grayscale change; performing non-maximum suppression on the gradient amplitude, retaining only the points with the largest amplitude in the direction, and obtaining the edge; determining a first threshold and a second threshold based on the statistical distribution of the gradient amplitude, the first threshold being greater than the second threshold; determining an edge with a gradient amplitude higher than the first threshold as a strong edge; determining an edge with a gradient amplitude lower than the second threshold as a weak edge; connecting the strong edge with the corresponding weak edge to obtain a time-sensitive target; and performing target detection of local image features on the time-sensitive target to obtain the target.
[0018] By adopting the above technical solution, the gradient amplitude and direction of the pixel points in the superpixel area are calculated, and the edge information of the image can be effectively extracted, because the edge is usually the dividing line between the target and the background. The edge is obtained by retaining the point with the largest amplitude in the direction through non-maximum suppression, which further refines the edge information and highlights the contour features of the target. According to the statistical distribution of the gradient amplitude, a reasonable threshold is determined, the edge is divided into strong edges and weak edges, and the strong edge is connected with the corresponding weak edge to obtain the time-sensitive target. This target extraction method based on edge features can quickly locate potential target areas under complex backgrounds and reduce interference from irrelevant areas. Target detection of local image features of time-sensitive targets is performed, and the local features of the target are used for identification, which improves the accuracy and speed of target recognition, especially for situations where the target appearance is changeable and partially occluded in complex environments.
[0019] In combination with some embodiments of the first aspect, in some embodiments, after the step of obtaining the original position and posture information of the sensor, the method also includes: fusing the thermal imaging image and the scanning image information to obtain a fused image; the step of calculating the three-dimensional coordinates of the target in the three-dimensional space based on the position information of the target in the scanning image information, using the distance information from the target to the sensor, and the original position and posture information, specifically includes: calculating the three-dimensional coordinates of the target in the three-dimensional space based on the position information of the target in the fused image, using the distance information from the target to the sensor, and the original position and posture information.
[0020] By adopting the above technical solution, the thermal imaging image and the scanning image information are fused to obtain a fused image, which gives full play to the advantages of thermal imaging images in detecting heating targets and penetrating smoke, as well as the advantages of scanning images in obtaining target details and shape information.
[0021] In a second aspect, the present application provides an aerospace detection device, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the aerospace detection device to perform the method described in the first aspect and any possible implementation method of the first aspect.
[0022] In a third aspect, the present application provides a computer program product comprising instructions, which, when executed on an aerospace detection device, enables the aerospace detection device to execute the method described in the first aspect and any possible implementation of the first aspect.
[0023] In a fourth aspect, the present application provides a computer-readable storage medium comprising instructions, which, when executed on an aerospace detection device, causes the aerospace detection device to execute the method described in the first aspect and any possible implementation of the first aspect.
[0024] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0025] 1. Performing super-pixel segmentation on the scanned image information to remove the background can effectively reduce the interference of irrelevant information and make subsequent target recognition more accurate and efficient. Then, the three-dimensional coordinates are calculated by integrating the sensor position and posture, the target image position and the distance information to lock the target spatial position. Obtaining the target features and combining the three-dimensional coordinates for multi-hypothesis tracking can cope with complex motion states and ensure the stability and accuracy of target trajectory detection, even if the target motion is changeable. Input the feature and trajectory data into the model to identify the target identity, judge the tracking conditions according to the preset list, and ensure that only key targets are tracked to avoid waste of resources. Finally, the target position and angular velocity information are output to guide the strike, improve the strike hit rate, reduce the strike failure caused by inaccurate target positioning or tracking errors, and overall improve the target tracking capability of autonomous or semi-autonomous aircraft in complex environments.
[0026] 2. By calculating the velocity and acceleration components of the target in each coordinate axis direction, we can gain an in-depth understanding of the target's motion state and trend. Based on the target's current motion parameters, we reasonably determine the extraction range to construct the motion hypothesis space. This adaptive construction method based on the target's current motion state makes the generated multiple motion hypotheses more in line with the target's actual possible motion conditions. When extracting velocity, acceleration, and azimuth, we fully consider the target's motion characteristics and physical limitations, such as the maximum steering angle, acceleration change amplitude, and other factors, to avoid generating unreasonable hypotheses. Therefore, when facing complex and diverse target motion scenes, we can more comprehensively and accurately cover the target's possible motion paths, improve the accuracy and efficiency of multi-hypothesis tracking, enhance the adaptability and responsiveness to changes in the target's motion state, ensure the accuracy and stability of target trajectory detection, and provide strong support for the aircraft to effectively track targets in complex environments.
[0027] 3. Integrate historical feature information, historical trajectory data, and corresponding historical identities into a data set and perform clustering. Clustering into multiple data subsets in detail according to historical identities enables each subset to accurately cover the features and trajectory characteristics of a specific type of target. The target recognition model is trained for each subset separately. The model can deeply analyze and master the unique features and trajectory patterns of each type of target, greatly improving the recognition accuracy and pertinence of a single target. Subsequently, the entire data set is used for comprehensive and integrated training, enabling the model to quickly and accurately distinguish and identify the identities of each target based on the unique feature information and trajectory data of each target when facing multiple targets in complex scenes. Even in complex situations where multiple targets are close to each other, occluded, or in different motion states, it can effectively reduce the probability of misjudgment and missed judgment, and enhance the accuracy and efficiency of multi-target recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a flow chart of a control method of an optoelectronic pod of a UAV in an embodiment of the present application;
[0029] Figure 2 It is a flow chart of a target identification, tracking and guidance method of an optoelectronic pod in an embodiment of the present application;
[0030] Figure 3 It is a flow chart of a target tracking and detection method of an optoelectronic pod in an embodiment of the present application;
[0031] Figure 4 It is a schematic diagram of a framework of a target tracking and detection method for an optoelectronic pod in an embodiment of the present application;
[0032] Figure 5 It is a schematic diagram of an exemplary hardware structure of aerospace detection equipment in an embodiment of the present application. DETAILED DESCRIPTION
[0033] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to be used as limitations to the present application. As used in the specification and appended claims of the present application, the singular expressions "one", "a kind of", "said", "above", "the" and "this" are intended to also include plural expressions, unless there is a clear indication to the contrary in the context. It should also be understood that the term "and / or" used in the present application refers to and includes any or all possible combinations of one or more listed items.
[0034] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as suggesting or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, unless otherwise specified, "plurality" means two or more.
[0035] See also Figure 1 , Figure 1 It is a flow chart of a control method of an optoelectronic pod of a UAV in an embodiment of the present application;
[0036] A control method for an optoelectronic pod of an unmanned aerial vehicle, comprising:
[0037] S101, when determining that the current position is in the mission area, guiding the autonomous or semi-autonomous aircraft to fly according to the planned route;
[0038] Among them, the mission area refers to a spatial area with clear boundaries and scope that is pre-set according to a specific aerospace exploration mission. Its boundaries are usually determined by parameters such as geographic coordinates and altitude range; the planned route refers to a flight path designed based on comprehensive considerations such as mission requirements, aircraft performance, environmental conditions (such as terrain, weather, etc.) and safety factors, which includes a series of continuous waypoints and flight trajectories connecting these waypoints.
[0039] In some embodiments, after the aircraft is started and lifted off, the real-time position information of the aircraft is continuously obtained. At the same time, the pre-stored mission area information is read, and the real-time position is compared with the boundary conditions of the mission area. Once it is determined that the current position is within the mission area, the aircraft retrieves the corresponding planned route data. Then, according to the starting point and initial flight direction in the planned route, the power system and flight attitude adjustment device of the aircraft are controlled to make the aircraft start flying along the planned route, ensuring that it can accurately enter the mission area and perform the detection mission according to the predetermined plan.
[0040] S102, scanning the task area on the planned route;
[0041] In some embodiments, when the aircraft flies along the planned route, the detection device is activated according to a preset scanning scheme. During the scanning process, the scanning control system monitors the position, speed and attitude information of the aircraft in real time, and adjusts the working parameters of the detection device according to this information to ensure the stability and accuracy of the scanning, such as adjusting the shooting frame rate of the camera according to the flight speed of the aircraft to ensure the resolution and clarity of the image.
[0042] It should be noted that the optical detection components of aerospace detection equipment are roughly divided into visible light television, infrared thermal imaging, laser rangefinders, and cameras.
[0043] In some embodiments, the seeker includes: an optical detection component, a seeker frame, and a servo platform. The optical detection component includes: a short-focus visible light TV, a long-focus visible light TV, an infrared thermal imaging device, and a laser rangefinder. The servo platform includes: a three-axis gyroscope, a driver, a motor, an angle measurement sensor, a servo computer board; a power board, an electrical connector, and a video tracking and recording board.
[0044] In other embodiments, the pod includes: an optical detection component, a seeker frame, and a servo platform. The optical detection component includes: a visible light television and an infrared thermal imager. The seeker frame includes: a gyroscope and a driver. The servo platform includes: a motor, an angle sensor, a servo computer board, a power board, a video tracking and recording board, and an electrical connector.
[0045] In some other embodiments, the optoelectronic pod includes: an optical detection component, an optoelectronic pod frame, and a servo platform. The optical detection component includes: a zoom camera, a wide-angle camera, an infrared thermal imager, and a laser rangefinder. The optoelectronic pod frame includes: a three-axis gyroscope and a driver. The servo platform includes: a motor, an angle sensor, a servo computer board, a power board, a video tracking and recording board, and an electrical connector.
[0046] The following is a description of the devices in the optical detection assembly:
[0047] In some embodiments, the task area is scanned using visible light;
[0048] Among them, the action distance function is:
[0049]
[0050] In the formula, is the distance to the target, H is the critical size of the target, is the focal length of the objective lens, is the limiting resolution of CMOS, n is the spatial frequency;
[0051] In some embodiments, the main function of the visible light television is to provide a stable daytime visible light image for the seeker. When the seeker is working, the visible light television provides real-time images. For changes in external light intensity, the visible light television adaptively adjusts the exposure time of the CMOS device to ensure real-time output of images suitable for observation and tracking.
[0052] The working principle of visible light television is as follows: the visible light reflected by the ground scenery is transmitted through the atmosphere, gathered on the focal plane of the detector through the optical system, converted into an image electrical signal by the detector, and the image electrical signal is stored through the image processing system, and then processed through defogging, target recognition, extraction, positioning information calculation, and identification to form an image product.
[0053] It should be noted that in order to complete the detection and identification of the target, while meeting the target imaging size (spatial resolution) requirements, the target imaging contrast requirements must also be met. The combination of the two can meet the detection and identification of the target.
[0054] The contrast function is:
[0055]
[0056] In the formula, is the contrast, Output signal contrast for TV sights. is the signal processing modulation transfer function, is the modulation transfer function of the display device, is the intrinsic contrast, is the atmospheric contrast transfer coefficient, is the modulation transfer function of the optical system, is the modulation transfer function.
[0057] In other embodiments, the visible light television is equipped with a focusing motor (in this case, it is called a zoom camera), and the rotation of the motor gear drives the rotation of the motor driving wheel, which in turn drives the focusing seat. The rotation of the focusing seat drives the front lens group to move forward and backward, and the movement of the front lens group changes the slight change in the focal length of the optical system to achieve the stability of the image plane, thereby achieving the purpose of focusing. When the focus is adjusted to the two extremes, there are automatic slipping mechanisms on the motor and a microcomputer switch in place protection mechanism to prevent the motor from being blocked and damaged. The circuit control part of the entire lens has a protection circuit to prevent damage to the entire machine due to high voltage and wrong connection, thereby improving the reliability of the entire machine. The motor and the lens are matched in a backpack type and distributed along the arc of the lens, with the smallest volume.
[0058] In other embodiments, the purpose of automatic dimming is to change the aperture size, exposure, gain value of the optical imaging system, and then change the luminous flux of the optical imaging system to adapt to the current ambient lighting and obtain a good image. Dimming mainly includes the following three parts:
[0059] Aperture adjustment The aperture controls the luminous flux per unit time. As the light increases, the aperture is increased by the motor drive, thereby increasing the luminous flux, and vice versa. However, improper aperture control will cause noise interference and affect the image quality. The aperture adjustment does not consider continuous adjustment, but it will consider adjustment by scene. Therefore, this system design selects a fixed aperture size and completes the dimming control by adjusting the two variables of gain value and exposure time.
[0060] The exposure time controls the amount of light entering by controlling the light entering time. The amount of light entering is the product of the light entering time and the light intensity.
[0061] Gain control adjusts the grayscale value of each pixel in the image through the internal amplification circuit of the optical imaging system. While amplifying the effective signal, it will also amplify the interference noise. Generally, exposure time control is given priority, followed by gain control.
[0062] It can be seen that when using visible light to scan the task area, the various parameters in the range function and contrast function, such as the focal length of the objective lens, the limiting resolution of the CMOS, the spatial frequency, the contrast, and the various transfer functions, cooperate with each other. The focal length of the objective lens and the limiting resolution of the CMOS determine the ability to capture the details of the target, and have a comprehensive impact on the clarity and contrast of the image, so that the target and the background can be more clearly distinguished during the scanning process, and the position and characteristics of the target can be determined more accurately, thereby improving the accuracy and efficiency of target detection.
[0063] In some embodiments, the mission area is scanned using an infrared thermal imaging camera solution;
[0064] Among them, the function of the minimum resolvable temperature difference is:
[0065]
[0066] Where MRTD is the minimum resolvable temperature difference at the target spatial frequency f, is the threshold signal-to-noise ratio, MTF(f) is the transfer function of the thermal imaging system, NETD is the noise equivalent temperature difference, is the instantaneous field of view of a single pixel of the detector, is the effective integration time of the human eye, is the frame rate of the thermal imaging system, is the circuit equivalent noise bandwidth, is the residence time;
[0067] In some embodiments, the infrared thermal imager can detect and monitor targets during the day and night, and is mainly used in the airborne field with high reliability requirements. It has the characteristics of small size, light weight and strong environmental adaptability. It is mainly composed of optical parts and mechanical parts.
[0068] The working principle of infrared thermal imager is: infrared radiation of target background is transmitted to infrared imager through atmosphere. The infrared optical system first receives infrared radiation and reflection of target and gathers energy on the photosensitive surface of infrared detector. The detector converts target radiation intensity into analog electrical signal, and image preprocessing circuit module buffers, amplifies and converts analog electrical signal into digital signal, realizes functions such as non-uniformity correction of image signal, bad element detection and bad element replacement, gain and bias control, and outputs analog and digital video signals.
[0069] The function of the action distance is:
[0070]
[0071] In the formula, T is the temperature difference between the target and the background after atmospheric attenuation, is the actual equivalent temperature difference between the target and the background, is the average atmospheric transmittance at the range R, MRTD is the minimum resolvable temperature difference at the target spatial frequency f, H is the characteristic size of the target, is the equivalent number of pixels occupied by the target on the detector.
[0072] It can be seen that when using the infrared thermal imager solution for scanning, the various parameters in the minimum resolvable temperature difference function and the range function work together. The transfer function of the thermal imaging system, the noise equivalent temperature difference, the instantaneous field of view of a single pixel of the detector, the effective integration time of the human eye, the frame rate of the thermal imaging system, the circuit equivalent noise bandwidth, the dwell time, and the atmospheric transmittance and other factors make it possible to detect the target and determine its range in different temperature environments and target thermal radiation characteristics. This makes up for the shortcomings of visible light scanning, further improves the comprehensiveness and accuracy of target detection, and reduces the omission of targets caused by environmental factors.
[0073] S103, running a target detection algorithm during the scanning process to detect targets in the task area;
[0074] In some embodiments, the data is preprocessed, such as grayscale, noise reduction, enhancement, normalization and other operations of the image, as well as filtering, amplification, demodulation and other processing of the signal, to improve the quality of the data and the recognizability of the features. Then, according to the type of target detection algorithm (such as traditional algorithms based on feature extraction or deep learning algorithms based on neural networks), feature extraction and pattern recognition are performed. For traditional algorithms, the shape, color, texture, edge and other features of the target may be extracted and compared and matched with the pre-established target feature library; for deep learning algorithms, the preprocessed data is input into the trained neural network model, and the model automatically learns and extracts the target features in the data, and outputs information such as the category, location and confidence of the target. Once determined as a target, the algorithm will record the detailed information of the target and mark it in the original data for subsequent tracking and analysis.
[0075] It should be noted that, in some embodiments, the following “target identification, tracking and guidance method for optoelectronic pod” may be used for detection. Specifically, steps S201 to S202, or S201 to S205, or S201 to S207;
[0076] In other embodiments, the following “target tracking detection method of optoelectronic pod” may be used for detection, specifically step S301. In this case, steps S103 and S104 are synchronized in time.
[0077] S104, tracking the detected target;
[0078] In some embodiments, after the target detection algorithm determines the target. The aerospace detection equipment will first quickly adjust the direction of the autonomous or semi-autonomous aircraft according to the initial position information of the target so that it is aimed at the target. Then, as the aircraft continues to fly and the target moves, the aerospace detection equipment will obtain the target's motion information in a variety of ways. For example, for optical image tracking, by comparing continuous image frames, using image feature matching algorithms (such as optical flow method, SIFT feature matching, etc.) to calculate the displacement of the target in the image, combined with the flight parameters of the aircraft (such as speed, altitude, attitude angle, etc.), the actual movement speed and direction of the target are calculated; for radar tracking, by analyzing the frequency change, delay change and other parameters of the radar echo signal, the distance, speed and angle change of the target are calculated in real time. According to this target motion information, the aerospace detection equipment controls the aircraft to adjust the flight attitude and speed of the aircraft accordingly, so that the aircraft can maintain the relative position and angle with the target, ensure that the target is always within the effective observation range of the detection equipment, and achieve stable tracking of the target.
[0079] It should be noted that, in some embodiments, the following “target identification, tracking and guidance method for optoelectronic pod” may be used for tracking. Specifically, steps S203 to S209, or S206 to S209, or S208 to S209;
[0080] In other embodiments, the following “target tracking and detection method of optoelectronic pod” may be used for tracking, specifically step S302. In this case, steps S103 and S104 are synchronized in time.
[0081] In some embodiments, step S104 specifically includes:
[0082] S1041. When a target is detected, track the target;
[0083] S1042. When multiple targets are detected, rank the threat levels of the multiple targets;
[0084] Among them, threat ranking refers to evaluating and ranking the potential threat levels of multiple targets based on pre-set evaluation criteria.
[0085] In some embodiments, the types of multiple detected targets are identified, and the target types are determined by comparing with a known target database. Different basic threat values are assigned to different types of targets, and the basic threat values are weighted according to the motion parameters of the targets. Finally, the threat level of each target is obtained and ranked, which is not calculated here.
[0086] S1043, determine the target with the greatest threat, and track the target with the greatest threat;
[0087] S1044. If no target is detected, return to step S103.
[0088] It can be seen that when multiple targets are detected, the threat level is ranked and the target with the greatest threat is determined for tracking. This ensures that the aircraft focuses limited resources on the most critical target, avoiding the waste of resources and omission of key information caused by indiscriminate tracking of multiple targets. This enables the most threatening target to be quickly locked in when facing complex mission scenarios.
[0089] In some embodiments, step S104 specifically includes:
[0090] S1045, receiving a command input by a user in a manual tracking mode;
[0091] S1046. Perform azimuth and / or pitch movement according to the instruction.
[0092] S1047. In the automatic tracking mode, according to the position information of the target in the image, determine the azimuth parameter and / or the pitch parameter according to the position information and the position information of the image center, and perform azimuth and / or pitch movement according to the azimuth parameter and / or the pitch parameter.
[0093] It can be seen that receiving the instructions input by the user and performing azimuth and / or elevation movements according to the instructions gives the operator direct control over target tracking in special circumstances. When the automatic tracking system may deviate or require more manual judgment, the operator can flexibly adjust the tracking direction based on his own experience and understanding of the task to ensure that the target is always within the monitoring range.
[0094] It can be seen that in the automatic tracking mode, the azimuth and pitch parameters are determined according to the position information of the target in the image and the position information of the image center, and the movement is performed accordingly. Through this automatic adjustment mechanism based on image information, the system can respond to the movement of the target quickly and accurately and maintain stable tracking of the target. It enhances the autonomy and stability of autonomous or semi-autonomous aircraft when performing detection tasks, reduces the dependence on continuous manual operation, and thus reduces labor costs and the risk of operational errors.
[0095] S105, starting the laser ranging strategy after tracking to obtain the azimuth, pitch angle information and distance information of the target;
[0096] In some embodiments, after the aircraft has stably tracked the target, the laser rangefinder will start working. First, the laser rangefinder will emit a laser beam of specific wavelength, power and pulse width to the target. After encountering the target, the laser beam will be reflected back and received by the laser rangefinder on the aircraft. The time interval between laser emission and reception is measured by a high-precision time measurement device, and the distance between the target and the aircraft can be calculated according to the speed of light formula. At the same time, the angle measurement sensor carried on the aircraft is used to obtain the current attitude angle of the aircraft (including roll angle, pitch angle and yaw angle) and the pointing angle (azimuth and pitch angle) of the detection equipment in real time, and the azimuth and pitch angle of the target relative to the aircraft coordinate system are determined through coordinate transformation and geometric calculation.
[0097] In some embodiments, the laser ranging strategy includes:
[0098] The action distance equation is:
[0099]
[0100] T=exp(-αδR)
[0101] In the formula, is the minimum detectable power at the measurement range R; is the transmitting power; is the receiving area; is the effective reflection area of the target; is the effective reflection area of the target; is the transmittance of the optical system; is the two-way atmospheric transmittance; is the laser beam angle; R is the range, α is the atmospheric attenuation coefficient, and δ is the altitude correction factor.
[0102] It can be seen that laser ranging is based on the range equation. The parameters such as transmission power, receiving area, target effective reflection area, optical system transmittance, two-way atmospheric transmittance, laser emission beam angle, atmospheric attenuation coefficient and elevation correction factor work together to calculate the range and minimum detectable power according to different environmental conditions and target characteristics, thereby achieving high-precision measurement of target distance in a complex and changeable aerospace environment, ensuring that the acquired target azimuth, pitch angle information and distance information are highly matched, providing a reliable data basis for subsequent flight path adjustment, target analysis, etc., and improving the accuracy of target positioning by detection equipment.
[0103] S106: Report the azimuth, pitch angle information, and distance information.
[0104] It can be seen that after determining that the aircraft is in the mission area, flying according to the planned route ensures the orderliness and pertinence of the detection. During the flight, the mission area is scanned and the target detection algorithm is run to search for the target. After the target is detected, it is tracked to maintain continuous attention to the target and obtain more detailed and accurate target dynamic information. After that, the laser ranging strategy is started to accurately determine the azimuth, pitch angle and distance information of the target. This information complements and works together to provide comprehensive and accurate data support for subsequent decisions, avoiding decision-making difficulties caused by missing or inaccurate data, while reducing the reliance on manual repeated confirmation and supplementary data, reducing labor costs, and improving the execution efficiency of the entire detection mission.
[0105] See also Figure 2 , Figure 2 It is a flow chart of a target identification, tracking and guidance method of an optoelectronic pod in an embodiment of the present application;
[0106] A target identification, tracking and guidance method for an optoelectronic pod, comprising:
[0107] S201, performing super-pixel segmentation on the acquired scanned image information, removing the background to obtain a super-pixel region;
[0108] Among them, scanned image information refers to the image data obtained by scanning the mission area through aerospace detection equipment. These data contain visual information of the target object and the surrounding environment. Superpixel segmentation is an image processing technology that divides an image into many small areas with similar features. These small areas are called superpixels. Compared with traditional pixel-level segmentation, superpixel segmentation can better preserve the edge and regional information of the image.
[0109] In some embodiments, when the aircraft arrives at the predetermined mission area and begins to perform the target detection mission, once sufficient image data is obtained, the image is divided into multiple initial superpixel blocks. Then, these superpixel blocks are merged and adjusted through an optimization algorithm to ensure that the pixels in each superpixel area have higher similarity and consistency. In this process, the boundaries between the superpixel area and the surrounding area are continuously evaluated, and those superpixel areas that are obviously background are removed, and finally the superpixel areas after the background is removed are obtained, which will serve as the basis for subsequent target recognition.
[0110] In some embodiments, in some embodiments, the target and the background are distinguished by monitoring and analyzing the pixel change and movement state of the target area. When the target in a certain area is determined not to have a pixel change greater than a predetermined threshold and has not moved at any position, the area is determined to be a background area.
[0111] S202, performing target recognition on the super-pixel area to identify the target;
[0112] In some embodiments, the following “target tracking and detection method for optoelectronic pod” may be used for identification, specifically step S301. In this case, steps S202 and S209 are synchronized in time.
[0113] In other embodiments, before target recognition, step S102 of the above “method for controlling an optoelectronic pod of a drone” may be used to acquire a scanned image.
[0114] A target is an object that is identified as having specific meaning and value through the target recognition process and requires subsequent tracking and processing.
[0115] In some embodiments, various feature information is extracted from the superpixel region, including color features (such as color histograms, color moments, etc.), texture features (such as gray-level co-occurrence matrices, local binary patterns, etc.), shape features (such as contour shape, area, perimeter, etc.), and deep learning-based features (such as high-level features extracted by convolutional neural networks), etc. Then, these extracted features are compared and matched with target feature templates pre-stored in the aircraft database or trained target recognition models. If the feature similarity exceeds a preset threshold, the object in the superpixel region is determined to be a target, and its position, size, category, and other information are recorded to provide basic data for subsequent target tracking and processing.
[0116] In some other embodiments, the target is determined to be a time-sensitive target, that is, step S202 includes:
[0117] S2021, calculating the gradient amplitude and direction of each pixel in the superpixel area, where the gradient amplitude is the intensity of grayscale change of the pixel, and the direction is the direction in which the grayscale changes fastest;
[0118] In some embodiments, for each pixel point in the superpixel area, its gradient magnitude and direction are calculated by a specific gradient operator (such as Sobel operator, Prewitt operator, etc.). Taking the Sobel operator as an example, it performs differential operations on the grayscale values of the pixel points in the horizontal and vertical directions, and then calculates the gradient magnitude through a certain mathematical formula (such as the Pythagorean theorem), and determines the direction with the fastest grayscale change as the gradient direction based on the differential result. In this way, the gradient magnitude and direction information of each pixel point in the superpixel area can be obtained.
[0119] S2022, performing non-maximum suppression on the gradient amplitude, retaining only the point with the largest amplitude in the direction, and obtaining the edge;
[0120] In some embodiments, for each pixel point, its gradient direction is first determined, and then the gradient amplitude of its adjacent pixel points in this direction is checked. If the gradient amplitude of the pixel point is not the maximum value among the adjacent pixel points in its gradient direction, the gradient amplitude of the pixel point is set to zero, that is, the point is suppressed; only the pixel point with the largest amplitude in the gradient direction is retained. By performing such processing on all pixel points in the superpixel area, an edge image after non-maximum suppression can be obtained. These edge points can more accurately reflect the true contour of the target object and reduce edge blur and misjudgment.
[0121] S2023, determining a first threshold and a second threshold according to the statistical distribution of the gradient amplitude, wherein the first threshold is greater than the second threshold;
[0122] In some embodiments, the gradient amplitude of all pixels in the superpixel area is counted, and the maximum value, minimum value, mean value, variance and other statistics are calculated. Then, the first threshold and the second threshold are determined based on these statistics and preset rules. A common method is to adopt the idea of Otsu's method, and regard the histogram of the gradient amplitude as consisting of two or more pixels with different distributions (corresponding to the target and the background), and determine the first threshold and the second threshold by finding the threshold that maximizes the inter-class variance. The threshold determined in this way can better adapt to the characteristics of different images, reasonably divide the edge points into strong edges and weak edges, and provide more targeted edge information for subsequent target extraction and recognition.
[0123] In some specific embodiments, a histogram of the gradient amplitude is statistically calculated and its distribution is observed. If the histogram shows an obvious bimodal feature (i.e., there are two peaks, corresponding to the gradient amplitude concentration areas of the target edge and the background, respectively), the first threshold and the second threshold are determined at the valley bottom position between the two peaks. For example, the first threshold can be determined by finding the minimum value point between the two peaks in the histogram, and then the second threshold is determined according to a certain proportional relationship (such as half of the first threshold), which is not limited here.
[0124] S2024, determining an edge whose gradient amplitude is higher than a first threshold as a strong edge;
[0125] In some embodiments, if the gradient magnitude of a pixel point is higher than a first threshold, the pixel point is marked as a strong edge point, and these strong edge points will be retained for subsequent target construction and recognition steps.
[0126] S2025, determining an edge whose gradient amplitude is lower than a second threshold as a weak edge;
[0127] In some embodiments, for those pixels whose gradient magnitude is lower than the second threshold, they are marked as weak edge points. Unlike the processing of strong edge points, weak edge points may not be directly used to construct the main contour of the target like strong edge points, but they will be used as auxiliary information to be associated and integrated with strong edge points in subsequent steps to improve the edge description and feature extraction of the target object.
[0128] S2026, connecting the strong edge with the corresponding weak edge to obtain a time-sensitive target;
[0129] In some embodiments, the strong edge is analyzed to determine its continuity and direction characteristics, and based on these characteristics, a weak edge that may be connected to it is searched within a certain range around it. For example, a weak edge point whose gradient amplitude is lower than a first threshold but higher than a second threshold and whose direction has a certain consistency with the strong edge can be searched within a set neighborhood along the tangent direction or normal direction of the strong edge. Once a suitable weak edge point is found, it is connected to the strong edge. In this way, the edge contour of the target is gradually improved, making the description of the time-sensitive target more accurate and complete.
[0130] S2027. Perform target detection based on local image features on the time-sensitive target to obtain the target.
[0131] In some embodiments, multiple representative local areas are selected from the image area of the time-sensitive target, and these areas can be determined based on the shape, edge distribution and prior knowledge of the target. Then, for each local area, its corresponding local image features are extracted, such as the texture, shape and color features mentioned above. Then, these extracted local features are compared and matched with the local feature templates of various types of targets pre-stored in the aircraft database or the trained target recognition model. By calculating the similarity between the features (such as Euclidean distance, cosine similarity, etc.), if the similarity exceeds a preset threshold, the time-sensitive target is determined to be a specific target category, and its position in the image is further accurately located, and its center coordinates, size and direction and other information are recorded.
[0132] It can be seen that calculating the gradient amplitude and direction of the pixel points in the superpixel area can effectively extract the edge information of the image, because the edge is usually the boundary between the target and the background. The edge is obtained by retaining the point with the largest amplitude in the direction through non-maximum suppression, which further refines the edge information and highlights the contour features of the target. According to the statistical distribution of the gradient amplitude, a reasonable threshold is determined, the edge is divided into strong edges and weak edges, and the strong edge is connected with the corresponding weak edge to obtain the time-sensitive target. This target extraction method based on edge features can quickly locate potential target areas under complex backgrounds and reduce interference from irrelevant areas. Target detection of local image features for time-sensitive targets is performed, and the local features of the target are used for recognition, which improves the accuracy and speed of target recognition, especially for situations where the target appearance is changeable and partially occluded in complex environments.
[0133] S203, obtaining the original position and posture information of the sensor;
[0134] In some embodiments, the aerospace detection equipment will directly read the latest measurement values from the registers or data buffers of each sensor. For example, for a GPS sensor, the current position coordinate information is obtained from the GPS receiver through a specific communication protocol and data parsing algorithm; for an IMU, the acceleration and angular velocity data measured by it are read, and the attitude angle of the aircraft is preliminarily calculated through integration operations. After these raw data are obtained, they will be stored in the aircraft's memory for use in subsequent steps, and will not be repeated here.
[0135] S204, calculating the three-dimensional coordinates of the target in the three-dimensional space based on the position information of the target in the scanned image information, using the distance information from the target to the sensor, the original position and the posture information;
[0136] In some embodiments, according to the pixel coordinates of the target in the image, combined with the intrinsic parameters of the camera (such as focal length, optical center position, etc.) and extrinsic parameters (such as the rotation and translation matrix of the camera, which can be calculated through the attitude information and position information of the aircraft), the pixel coordinates of the target are converted into two-dimensional coordinates in the camera coordinate system through the perspective projection transformation formula. Then, using the distance information from the target to the sensor, the two-dimensional coordinates in the camera coordinate system are expanded into three-dimensional coordinates. Finally, according to the original position information of the aircraft, the three-dimensional coordinates of the target in the camera coordinate system are converted into three-dimensional coordinates with reference to the earth coordinate system or other global coordinate systems, thereby obtaining the position information of the target in three-dimensional space.
[0137] In some embodiments, step S204 specifically includes:
[0138] S2041, fusing the thermal imaging image and the scanning image information to obtain a fused image;
[0139] Among them, the fused image is a new image obtained by integrating the thermal imaging image and the scanning image information through specific algorithms and technical means. It combines the advantages of both, including the thermal characteristic information of the object in the thermal imaging image, and the details and position of the target in the scanning image.
[0140] In some embodiments, necessary preprocessing is performed on the thermal imaging image and the scanned image information in terms of spatial coordinates, resolution, data format, etc., so that they meet the basic requirements for fusion. For example, the resolution of the thermal imaging image is adjusted to match the resolution of the scanned image; the coordinate systems of the two are calibrated so that they can correspond to each other in space; and a suitable fusion algorithm is selected. Common ones include pixel-level fusion algorithms (such as weighted average method, wavelet transform method, etc.) and feature-level fusion algorithms (such as principal component analysis method, feature fusion matching method, etc.). Taking the weighted average method as an example, for the pixel points at the corresponding positions in the thermal imaging image and the scanned image, a weighted sum is performed according to a certain weight (the weight can be pre-set according to the image characteristics and actual needs, for example, when the thermal imaging image focuses more on temperature information, its weight can be appropriately increased), and the obtained new pixel value is assigned to the corresponding pixel position in the fused image. By performing such operations on all pixel points in sequence, a fused image can be obtained.
[0141] S2042: Calculate the three-dimensional coordinates of the target in the three-dimensional space based on the position information of the target in the fused image, using the distance information from the target to the sensor, the original position and the posture information.
[0142] It should be noted that the principle and process of this step are similar to those of step S204. The relevant principles and processes can refer to step S204 and are not limited here.
[0143] It can be seen that the fusion image obtained by fusing the thermal imaging image and the scanning image information fully utilizes the advantages of thermal imaging images in detecting hot targets and penetrating smoke, as well as the strengths of scanning images in obtaining target details and shape information.
[0144] S205, obtaining characteristic information of each target;
[0145] Among them, the target refers to various information that can describe the unique attributes and characteristics of the target object, including but not limited to the target's visual characteristics (such as color, shape, texture, edge, etc.), physical characteristics (such as size, mass, material, etc.), and motion characteristics (such as speed, acceleration, direction of movement, etc.). These feature information can be used to distinguish different target objects.
[0146] In some embodiments, after the position of the target is determined and its three-dimensional coordinates are calculated, the feature extraction module of the aircraft starts to work. For each target, its visual features are first extracted from the scanned image, such as the color histogram, shape descriptor (such as Hu moment, Fourier descriptor, etc.), texture features (such as gray-level co-occurrence matrix features) of the target are calculated through image segmentation and feature extraction algorithms. At the same time, combined with the motion information of the target, such as calculating its velocity and acceleration vector by the position change of the target in continuous image frames, and using other sensors (such as radar, infrared sensor, etc.) to measure the electromagnetic or thermal radiation characteristics of the target. These different types of feature information are integrated and quantified to form a feature vector for each target, which is stored in the memory of the aircraft for subsequent target trajectory detection and identity recognition tasks.
[0147] S206, performing target trajectory detection based on multi-hypothesis tracking according to the feature information and the three-dimensional coordinates to obtain trajectory data;
[0148] Among them, target trajectory detection based on multi-hypothesis tracking is a tracking method that simultaneously considers multiple possible motion states and trajectories of the target. It will infer multiple possible motion trajectory hypotheses of the target in the future based on the target's historical position (three-dimensional coordinates) and feature information. Then, as new observation data (such as three-dimensional coordinates and feature information at subsequent moments) are acquired, these hypotheses are verified and screened, and the trajectory that best matches the actual motion of the target is gradually determined. Multi-hypothesis tracking will take these different motion possibilities into account, generate multiple potential driving trajectory hypotheses, and then determine the most likely actual driving trajectory of the vehicle based on the new position information.
[0149] In some embodiments, step S206 specifically includes:
[0150] S2061, constructing a plurality of different motion trajectories for each target according to the feature information and the dynamic change data of the three-dimensional coordinates over time;
[0151] In some embodiments, when the aircraft continuously monitors the target, as time goes by, the characteristic information and three-dimensional coordinate data of the target are continuously collected, and its dynamic changes are recorded. For each target, its historical characteristic information and three-dimensional coordinate change trend are first analyzed. For example, if the speed of the target has shown a gradual increase over the past period of time, and the color characteristics remain stable, based on this information, combined with the target's possible motion mode (such as uniformly accelerated linear motion, uniform circular motion, etc.) and environmental constraints (such as terrain, obstacle distribution), multiple different motion trajectory hypotheses are generated through mathematical models and algorithms. These hypotheses not only take into account the range of variation of the target's speed, acceleration and other motion parameters, but also construct a series of trajectory models covering various possible motion paths of the target, providing multiple candidate solutions for subsequent tracking and verification to cope with the complex and changeable motion state of the target and improve the reliability and accuracy of target tracking.
[0152] In some embodiments, step S2061 specifically includes:
[0153] S20611, calculating the velocity component of the target in each coordinate axis direction according to the three-dimensional coordinate changes of the target at consecutive time points;
[0154] In some embodiments, when the aircraft continuously acquires the three-dimensional coordinate information of the target and accumulates the coordinate data of a certain number of consecutive time points, the velocity component calculation begins. First, the three-dimensional coordinates of the target corresponding to two adjacent time points are selected, and then the velocity components in each coordinate axis direction are obtained by calculating the ratio of the coordinate difference to the time interval.
[0155] S20612, calculating the acceleration component of the target in each coordinate axis direction according to the change of the velocity component at different times;
[0156] In some embodiments, after obtaining the data of the velocity component of the target in each coordinate axis direction over time, the acceleration component is calculated. The velocity component values corresponding to at least three consecutive time points are selected. The acceleration component is obtained by calculating the ratio of the difference between adjacent velocity components and the time interval;
[0157] In some embodiments, in order to improve accuracy, optimization may be performed by taking an average value through multiple calculations.
[0158] S20613, constructing a motion hypothesis space by changing acceleration, velocity and direction angle;
[0159] In some embodiments, after obtaining the acceleration component and velocity component of the target in each coordinate axis direction and determining the relevant description of the direction angle, the motion hypothesis space is constructed. The acceleration, velocity, and direction angle are combined within their respective value ranges to form a multidimensional space, which is the motion hypothesis space, which contains various possible combinations of the target's motion states in the future, providing rich possibilities for the subsequent extraction of specific motion hypotheses.
[0160] S20614. Extract the next motion state parameters of the target from the motion hypothesis space to generate multiple motion hypotheses that conform to the probability distribution. The motion state parameters include acceleration, speed, and direction angle. When extracting speed, the current speed of the target is used as the midpoint, and the speed change amount generated by the current acceleration within the preset time step is used as the radius to determine the extraction range. When extracting acceleration, the current acceleration of the target is used as the midpoint, and the acceleration change amplitude is used as the radius to determine the extraction range. When extracting direction angle, the current direction angle of the target is used as the midpoint, and the maximum steering angle is used as the radius to determine the extraction range.
[0161] In some embodiments, after the motion hypothesis space is constructed, it is necessary to extract the next motion state parameters of the target from this space to generate multiple motion hypotheses that conform to the probability distribution. First, according to a predetermined probability distribution model (this model can be constructed based on the type of target, historical motion data, and environmental factors, for example, for some common target motion modes, a higher probability is given; for some extreme cases that do not conform to physical laws or target routine operations, a lower probability is given), the three motion state parameters of acceleration, speed, and direction angle are extracted in the motion hypothesis space. When extracting speed, the current speed of the target is taken as the midpoint, and the speed change amount generated by the current acceleration within the preset time step is used as the radius to determine the extraction range. For extracting acceleration, the current acceleration of the target is taken as the midpoint, and the acceleration change amplitude is used as the radius to determine the extraction range. When extracting the direction angle, the current direction angle of the target is taken as the midpoint, and the maximum steering angle is used as the radius to determine the extraction range. In this way, each time a set of motion state parameters such as acceleration, velocity, and angular direction are extracted, a motion hypothesis is generated. By repeating the extraction process multiple times, multiple motion hypotheses that conform to the probability distribution can be obtained. These motion hypotheses cover various possible motion situations of the target in the future and take into account the probability of their occurrence, which is more conducive to the subsequent tracking and judgment of the actual motion trajectory of the target.
[0162] It can be seen that by calculating the velocity and acceleration components of the target in each coordinate axis direction, the motion state and trend of the target can be deeply understood. Based on the current motion parameters of the target, the extraction range is reasonably determined to construct the motion hypothesis space. This adaptive construction method based on the current motion state of the target makes the generated multiple motion hypotheses more in line with the actual possible motion of the target. When extracting velocity, acceleration and direction angle, the motion characteristics and physical limitations of the target, such as the maximum steering angle, acceleration change amplitude and other factors, are fully considered to avoid generating unreasonable hypotheses. Therefore, when facing complex and diverse target motion scenes, it can more comprehensively and accurately cover the possible motion paths of the target, improve the accuracy and efficiency of multi-hypothesis tracking, enhance the adaptability and responsiveness to changes in the target motion state, ensure the accuracy and stability of target trajectory detection, and provide strong support for the aircraft to effectively track targets in complex environments.
[0163] S2062, using the new three-dimensional coordinates received subsequently, verifying the relevance of the motion hypothesis;
[0164] In some embodiments, an appropriate distance measurement method (such as Euclidean distance) is used to calculate the error value between the predicted coordinates and the newly received actual coordinates. The smaller the error value, the higher the correlation between the motion hypothesis and the actual motion trajectory of the target.
[0165] S2063: Select the motion trajectory with the highest correlation as trajectory data.
[0166] It should be noted that, although a certain motion trajectory presents the highest correlation among the multiple constructed trajectories, if its correlation value is lower than a preset threshold, then the motion trajectory still needs to be corrected, or a new motion trajectory needs to be directly regenerated.
[0167] It can be seen that multiple motion trajectory hypotheses are constructed based on feature information and dynamic changes of three-dimensional coordinates over time, taking into account multiple possibilities of target movement. Using subsequent new three-dimensional coordinates for correlation verification can timely eliminate unreasonable hypotheses and ensure the accuracy of tracking. The trajectory with the highest correlation is selected as the final trajectory data. This method based on data-driven and probability statistics can effectively avoid tracking loss caused by sudden acceleration, deceleration, turning, etc. of the target when the target motion state is complex and changeable, improve the stability and reliability of target tracking, and provide a trajectory basis for subsequent target identification, decision-making and guidance operations.
[0168] S2064, when the three-dimensional coordinates of the current target are not updated within a preset time, feature comparison is performed on feature information of subsequent targets in subsequent scanned image information according to feature information of the current target; the current target is any target, and the subsequent scanned image information is scanned image information obtained subsequently to the scanned image information of the current target;
[0169] In some embodiments, when the aircraft finds that the three-dimensional coordinates of the current target have not been updated within a preset time, in order to avoid the loss of the target, the feature information of the target is used for subsequent processing. First, the feature information of the current target is extracted from the previously stored target information, including color, shape, texture, and motion features. Then, after the subsequent scanned image information is obtained, the feature extraction is performed on each potential target area in the image, and its color, shape, texture and other feature information are also obtained. Next, a specific feature comparison algorithm is used to compare the feature information of the current target with the feature information of each potential target in the subsequent image one by one, and the similarity between them is calculated. For example, for color features, the histogram intersection method can be used to calculate the similarity; for shape features, the shape context algorithm can be used to calculate the matching degree. Through these methods, subsequent targets with high similarity to the current target features are found, which provide a basis for further determining whether the target is lost and relocating the target, ensuring that in the case of a short-term loss of the target, there is still a chance to retrieve it through the feature information, and maintain the continuity and stability of target tracking.
[0170] S2065, determining the subsequent target whose feature comparison result is greater than the similarity threshold and has the highest similarity as the lost target of the current target;
[0171] S2066: splicing the trajectory data of the current target and the trajectory data of the lost target, wherein the missing trajectory data is supplemented by using the difference.
[0172] In some embodiments, after determining the lost target of the current target, the trajectory data of the two need to be spliced to restore the complete motion trajectory of the target. First, analyze the last known position and state information of the current target before it is lost, as well as the first position and state information of the lost target after it reappears, including three-dimensional coordinates, speed, direction, etc. Then, based on the known motion characteristics and environmental information of the target, estimate the possible motion path and state changes of the target during the loss period, and calculate the difference information of the missing trajectory part. For example, using the historical speed data and loss time of the target, the kinematic formula is used to calculate the distance and direction changes that the target may move during the loss period, thereby obtaining a series of intermediate position points and state parameters for supplementing the missing trajectory. Then, these differences are added to the trajectory data of the lost target so that it can be reasonably connected with the trajectory data of the current target in time and space to form a continuous and complete target motion trajectory.
[0173] It can be seen that when the three-dimensional coordinates of the target are not updated within the preset time, the lost target is found in the subsequent scanned images through feature matching. Using the feature information of the target, these features are relatively stable and unique. Even if the target is temporarily lost, it can still be relocated in the subsequent images based on feature matching. The subsequent target with the highest similarity and feature comparison result greater than the similarity threshold is determined as the lost target, and the trajectory data is spliced, and the missing part is supplemented with the difference, so that the trajectory of the target can be continued after it is lost, ensuring the integrity and coherence of the target trajectory. This is particularly important for long-term and long-distance target tracking tasks. It can effectively avoid the loss of previous efforts due to the temporary loss of the target, ensure the stability and continuity of the entire tracking process, and provide complete and accurate trajectory information for subsequent target analysis and strike guidance.
[0174] S207, inputting the feature information and trajectory data into a target recognition model to obtain the identity of the target;
[0175] In some embodiments, after the feature information and trajectory data of the target are obtained, they are input into a pre-trained target recognition model. The model first pre-processes the input feature information, such as normalization, feature selection or dimensionality reduction, to improve the quality of the data and the processing efficiency of the model. Then, according to the internal structure and algorithm of the model (such as the hierarchy of the neural network, the branching rules of the decision tree, etc.), feature extraction and pattern recognition are performed on the feature information and trajectory data. The model compares and matches the feature patterns and trajectory patterns of various targets learned in the training phase, and calculates the probability distribution of the target belonging to different categories or identities. Finally, according to the set probability threshold or decision rule, the identity category of the target is determined, and the corresponding recognition result is output.
[0176] In some embodiments, the model training process includes:
[0177] S2071, taking historical feature information, historical trajectory data, and corresponding historical identities as a data set;
[0178] S2072, clustering the historical feature information and historical trajectory data into a number of data subsets according to historical identities;
[0179] In some embodiments, after obtaining a data set containing historical feature information, historical trajectory data, and historical identity, clustering processing is started. First, select a suitable clustering algorithm, common ones include K-Means clustering algorithm, hierarchical clustering algorithm, etc. (here it is assumed that the K-Means clustering algorithm is selected as an example for explanation). Then, determine the key feature dimensions of clustering, that is, which specific indicators in the historical feature information and historical trajectory data are used as the basis for judging similarity. For example, the shape characteristics and average speed of the target can be selected as feature dimensions. Next, according to the historical identity labels in the data set, the target-related data with the same historical identity is used as a reference for the initial clustering, and the clustering algorithm is started for calculation. Taking the K-Means algorithm as an example, K cluster centers are randomly initialized first (the value of K is determined according to the known number of historical identity categories or the estimated number of categories), and then the distance from each data point (i.e., the point corresponding to each set of historical feature information and historical trajectory data) to these cluster centers is calculated (the distance measurement method can be selected from Euclidean distance, cosine distance, etc. according to the characteristics of the data), and the data points are assigned to the category represented by the nearest cluster center. The cluster centers are then continuously updated, and the process of assigning data points is repeated until the cluster centers no longer change significantly or the preset number of iterations is reached. Finally, several data subsets clustered according to historical identities are formed. These subsets can clearly reflect the data characteristics and trajectory characteristics of different types of targets.
[0180] S2073, training the target recognition model in sequence according to the data subsets, so that the target recognition model can recognize the historical identity according to the historical feature information and the historical trajectory data;
[0181] S2074. Train a target recognition model according to the data set, so that the target recognition model can recognize multiple historical identities according to multiple historical feature information and multiple historical trajectory data.
[0182] In some embodiments, after completing the sequential training of the target recognition model based on the data subset, the target recognition model is then comprehensively trained using the entire data set. The data set is comprehensively preprocessed. Since the data set contains data on multiple types of targets, the data format, magnitude, etc. may be different, so unified processing is required, such as normalizing all historical feature information so that its numerical range is within an appropriate interval, and regularizing and aligning the historical trajectory data in time series to ensure that the trajectory data of different targets are comparable in the time dimension. Then, the processed data set is divided into a training set and a validation set according to a certain ratio (such as the common 80% as a training set and 20% as a validation set). The training set is used for parameter learning and adjustment of the model, and the validation set is used to evaluate the performance of the model on unseen data. Next, the historical feature information and historical trajectory data in the training set are input into the target recognition model, and the corresponding historical identity is used as the correct label for forward propagation calculation to obtain the prediction result of the model. The prediction error is measured by calculating the loss function (such as the mean square error loss function, etc.), and then the back propagation algorithm is combined with the optimizer (such as the Adam optimizer, etc.) to adjust the model parameters according to the loss value. During the training process, the model is regularly verified using the validation set to observe the changes in the model's accuracy, recall rate and other evaluation indicators on the validation set. Based on these indicators, it is judged whether the model has problems such as overfitting or underfitting, and the training parameters (such as learning rate, number of training rounds, etc.) are adjusted in time. The above training process is continued until the performance of the model on the validation set reaches the expected standard, such as the accuracy rate reaches more than 90%. At this time, the model can accurately identify multiple historical identities based on multiple historical feature information and multiple historical trajectory data, and completes the training based on the entire data set, so that the model has stronger generalization ability and comprehensive recognition ability for multiple targets in complex scenarios.
[0183] It can be seen that the historical feature information, historical trajectory data and corresponding historical identities are integrated into a data set and clustered. The data are carefully clustered into multiple data subsets according to the historical identity, so that each subset accurately covers the characteristics and trajectory characteristics of a specific type of target. The target recognition model is trained for each subset separately. The model can deeply analyze and master the unique characteristics and trajectory patterns of various types of targets, greatly improving the recognition accuracy and pertinence of a single target. Subsequently, the entire data set is used for comprehensive and integrated training, so that the model can quickly and accurately distinguish and identify the identities of each target based on the unique feature information and trajectory data of each target when facing multiple targets in complex scenes. Even in complex situations where multiple targets are close to each other, blocked or in different motion states, it can effectively reduce the probability of misjudgment and missed judgment, and enhance the accuracy and efficiency of multi-target recognition.
[0184] S208, judging whether the target meets the tracking condition according to the preset target list and the identity of the target;
[0185] In some embodiments, after obtaining the identity of the target, the pre-stored target list and the set tracking conditions are read. The identity of the target is compared with the items in the target list to check whether the target belongs to the object of interest specified in the list. At the same time, based on the location, motion state and other information of the target, combined with the set tracking conditions (such as whether the target enters a specific area of interest, whether it has specific motion characteristics, etc.), a comprehensive judgment is made as to whether the target meets the tracking requirements.
[0186] S209, if the target meets the tracking condition, track the target;
[0187] In some embodiments, after determining that the target meets the tracking conditions, the flight parameters that need to be adjusted by the aircraft, such as flight direction, speed, altitude, etc., are calculated based on the current position and motion state of the target, so that the aircraft can quickly approach the target and remain within the appropriate tracking distance and angle range.
[0188] In some embodiments, the following “target tracking and detection method for optoelectronic pod” may be used for tracking, specifically step S302. In this case, steps S202 and S209 are synchronized in time.
[0189] In other embodiments, before target recognition, step S102 of the above “method for controlling an optoelectronic pod of a drone” may be used to acquire a scanned image.
[0190] S210: Output the target's position information and the target's angular velocity information to guide the execution of the striking process.
[0191] It can be seen that super-pixel segmentation of scanned image information to remove background can effectively reduce the interference of irrelevant information and make subsequent target recognition more accurate and efficient. Then, the three-dimensional coordinates are calculated by integrating the sensor position and posture, target image position and distance information to lock the target spatial position. Obtaining target features and combining them with three-dimensional coordinates for multi-hypothesis tracking can cope with complex motion states and ensure the stability and accuracy of target trajectory detection, even if the target motion is changeable. Input the feature and trajectory data into the model to identify the target identity, judge the tracking conditions according to the preset list, and ensure that only key targets are tracked to avoid waste of resources. Finally, the target position and angular velocity information are output to guide the strike, improve the strike hit rate, reduce the strike failure caused by inaccurate target positioning or tracking errors, and overall improve the target tracking capability of autonomous or semi-autonomous aircraft in complex environments.
[0192] See also Figure 3 , Figure 3It is a flow chart of a target tracking and detection method of an optoelectronic pod in an embodiment of the present application;
[0193] A target tracking and detection method for an optoelectronic pod, comprising:
[0194] S301, performing target detection on the video image to obtain a target detection result;
[0195] In some embodiments, a suitable target detection algorithm is selected. Common ones include deep learning-based target detection algorithms (such as the YOLO series, FasterR-CNN, etc.) and traditional feature-based target detection algorithms (such as Haar features + AdaBoost algorithm, etc.). Taking the YOLO algorithm as an example, the input video image is first preprocessed to adjust the image size, normalize the pixel values, etc., so that it meets the requirements of the algorithm input. Then, the processed image is sent to the pre-trained or trained YOLO network model. The network model will perform feature extraction, classification and positioning operations on the image, and finally output the target detection result through multiple convolutional layers and fully connected layers.
[0196] It should be noted that, in some embodiments, the above “target identification, tracking and guidance method of optoelectronic pod” can be used for detection. Specifically, steps S201 to S202, or S201 to S205, or S201 to S207;
[0197] In other embodiments, before target recognition, step S102 of the above “method for controlling an optoelectronic pod of a drone” may be used to acquire video images.
[0198] In some embodiments, step S301 specifically includes:
[0199] S3011, screening out a time-sensitive region from the video image;
[0200] In some embodiments, a preliminary motion analysis is performed on the video image, and the motion of pixels in the image is detected using algorithms such as the optical flow method. Areas where pixels move frequently and change greatly are screened out, because these areas are often where objects (possibly targets) are active and are time-sensitive areas.
[0201] S3012, extracting feature information of the time-sensitive area;
[0202] In some embodiments, for each time-sensitive region, a suitable feature extraction method combination is selected. For example, for color feature extraction, the image in the region can be first converted into a suitable color space (such as RGB color space or HSV color space, etc.), and then the color histogram under the corresponding color space is calculated to count the distribution of the number of pixels of different color components; for texture feature extraction, the grayscale co-occurrence matrix method can be used to set suitable parameters (such as pixel distance, angle, etc.), calculate the symbiotic relationship of pixel grayscale in different directions, and then obtain relevant parameters reflecting texture characteristics; for shape feature extraction, an edge detection algorithm (such as the Canny edge detection algorithm) is used to obtain the edge contour of the object in the region, and then its shape characteristics are described by calculating the perimeter, area, circularity and other geometric parameters of the contour; at the same time, the spatial position characteristics of the region are recorded (such as the center coordinates of the region, the coordinates of the upper left corner and the lower right corner, etc.). Then, these different types of feature information are sorted and stored to form a feature information set corresponding to each time-sensitive region, so as to be matched and compared with the pre-stored target feature template later.
[0203] S3013, matching and comparing the feature information with a pre-stored target feature template;
[0204] In some embodiments, the feature information of the time-sensitive region and the feature information of the target feature template are organized into corresponding feature vector forms (assuming that the feature vector of the time-sensitive region is , and the feature vector of the target feature template is , where is the dimension of the feature vector, corresponding to different feature parameters), and then the distance value between the two is calculated according to the Euclidean distance calculation formula. The smaller the distance value, the more similar the two are, and the greater the possibility that the time-sensitive region contains the target. Alternatively, if the method of comparing the range of feature parameter values is adopted, the values of the time-sensitive region and the target feature template in each interval of the color histogram, shape geometric parameters (such as aspect ratio, perimeter, area, etc.), texture feature parameters (such as contrast, correlation, etc.) are compared to determine whether they are within a reasonable similarity range. For example, if the difference between the shape aspect ratio of the time-sensitive region and the aspect ratio of the target feature template is within a certain threshold, it is considered that the two are more matched in terms of shape features.
[0205] S3014. Determine the time-sensitive area whose matching result is greater than the threshold as the target.
[0206] It can be seen that the time-sensitive area is screened out from the video image, its feature information is extracted and matched with the pre-stored target feature template, and the area with a matching result greater than the threshold is determined as the target. The screening of the time-sensitive area can quickly focus on key areas where target changes or new targets may appear, avoiding comprehensive and time-consuming detection of the entire video image. By matching and comparing with the pre-stored template, the existing target feature knowledge is used to efficiently identify the target, reducing the computational complexity and time consumption of invalid detection.
[0207] S302, during target detection, tracking the target in the video image to obtain a target tracking result;
[0208] It should be noted that, in some embodiments, the above “target identification, tracking and guidance method for optoelectronic pod” can be used for tracking. Specifically, steps S203 to S209, or S206 to S209, or S208 to S209;
[0209] In other embodiments, before target recognition, step S102 of the above “method for controlling an optoelectronic pod of a drone” may be used to acquire video images.
[0210] In some embodiments, according to the target initial position information obtained in the target detection step, the initial region of the target is determined in the video image, and the features of the region (such as HOG features, etc.) are extracted as the feature representation of the target. Then, in the subsequent video image frames, a certain range where the target may appear (usually estimated based on the motion characteristics of the target and prior knowledge) is used as the search area, and the position of the target in the search area that is most similar to the target feature is calculated using a correlation filtering algorithm, the position of the target in the current frame is determined, and the position information and the speed, direction and other information calculated by comparing with the position of the previous frame are recorded to form the target tracking results, and these results are continuously updated with the continuous input of video image frames to achieve continuous tracking of the target.
[0211] It should be clear that in the method adopted in this embodiment, the two operation processes of step S301 and step S302 are carried out independently and in parallel, so that target detection and target tracking can be started at the same time.
[0212] In some specific embodiments, step S302 specifically includes: S30210, extracting features from adjacent frame images to obtain location information of the target in the adjacent frames;
[0213] It should be noted that feature extraction has been discussed in detail in step S3012, and related steps can refer to step S3013, which will not be repeated here.
[0214] S30211. Obtaining location change information according to the location information;
[0215] In some embodiments, the position coordinate information of the target in adjacent frames is taken out in chronological order, and the position changes in the horizontal and vertical directions are calculated respectively, that is, the horizontal displacement and the vertical displacement are calculated, and these two displacement changes are recorded, which together constitute the position change information of the target between the two frames.
[0216] S30212, determining the moving speed and moving direction of the target according to the position change information;
[0217] In some embodiments, the time interval between adjacent frames is determined according to the frame rate of the video image, and then, for each pair of position change information between adjacent frames (i.e., horizontal displacement and vertical displacement), the instantaneous speed of the target is calculated by calculating the speed component in the horizontal direction and the speed component in the vertical direction, and the Pythagorean theorem is used to calculate the target's motion speed at different adjacent frame stages. Then, the target's motion direction at each stage is determined by calculating the angle between the target's motion direction and the positive direction of the horizontal axis and using the inverse tangent function.
[0218] S30213, predicting the position range where the target will appear in the next frame according to the movement speed and movement direction;
[0219] In some embodiments, the theoretical displacement of the target in the time interval between adjacent frames is calculated based on the target's movement speed and direction, and the theoretical center coordinate position of the target in the next frame is calculated based on the target's position coordinates in the current frame, and an appropriate rectangular or elliptical area is constructed as the position range of the target in the next frame.
[0220] S30214. Target matching and tracking are performed based on the texture intensity and phase characteristics of the target within the position range.
[0221] It can be seen that the target position information is obtained by extracting features from adjacent frame images, and then the position change information is obtained to determine the target movement speed and direction, and then the target position range in the next frame is predicted, and the target texture intensity and phase features are combined for target matching and tracking. Through feature analysis of adjacent frames and prediction of target motion state, the area where the target may appear can be planned in advance, reducing blind searches in the entire image space and improving tracking efficiency. At the same time, matching and tracking using texture intensity and phase features enhances the accuracy and stability of target recognition, especially when the target appearance and posture change, and can locate the target more accurately.
[0222] In actual use, the amount of calculation required to track the target in the video image is too large; therefore, in some embodiments, step S302 specifically includes:
[0223] S3021, dividing the video image into a plurality of regions;
[0224] In some embodiments, a uniform grid division method is used, that is, according to a preset number of rows and columns, the size of the image is calculated, and the image is divided horizontally and vertically at corresponding intervals to obtain small rectangular areas. Another method can be content-aware division, which first performs simple edge detection, object recognition and other preprocessing on the image, and then divides the image into areas of different shapes and sizes according to the detected object contours or different semantic areas (such as distinguishing the foreground object from the background area, and further subdividing the range of the foreground object, etc.).
[0225] S3022, comparing the statistical characteristics of each region with the target characteristic range;
[0226] In some embodiments, the feature comparison method has been discussed in detail in step S3013 and will not be repeated here.
[0227] S3023, if the statistical characteristics of the region do not belong to the target feature range, a constant false alarm detection technology is used to adjust the target feature range;
[0228] In some embodiments, when the statistical characteristics of a certain area do not belong to the target feature range after comparison, it is necessary to use constant false alarm detection technology to adjust the target feature range. Then, according to the set false alarm probability (generally pre-set according to the actual application scenario and the tolerance for misjudgment, such as setting the false alarm probability to 0.05), a constant false alarm detection algorithm (common ones include constant false alarm detection algorithms based on unit average, ordered statistical constant false alarm detection algorithms, etc.) is used to adjust the relevant parameters in the target feature range based on the pixel feature distribution in the area. For example, if it is a constant false alarm detection algorithm based on unit average, the area will be divided into multiple small units first, the average value of the pixel features in each unit will be counted, and then the appropriate adjustment threshold will be calculated based on the overall feature distribution and false alarm probability, and then the parameters such as the target color feature range and shape feature range will be modified accordingly, so that the target feature range can better adapt to the actual situation of this area, so as to judge whether the area may contain the target again in the future, and improve the accuracy of target detection and tracking.
[0229] S3024. If the statistical characteristics of the region do not fall within the adjusted target characteristic range;
[0230] In some embodiments, after adjusting the target feature range by the constant false alarm detection technology, it is necessary to compare the statistical characteristics of the region with the adjusted target feature range again to further confirm whether the region is likely to contain the target. The operation of this step is similar to the basic process of comparison in step S3022, and will not be repeated here.
[0231] S3025: The area is determined as a non-target area and suppression processing is performed.
[0232] In some embodiments, when it is determined through step S3024 that the statistical characteristics of the region do not belong to the adjusted target feature range, the region can be determined as a non-target region and suppression processing can be performed. First, a specific processing method is selected according to the set suppression processing strategy. For example, if the method of setting the regional pixel value to zero is selected, all pixel points in the region are traversed and their pixel values are modified to 0, so that they become black visually, thereby achieving the purpose of eliminating the influence of the region; if the method of reducing the pixel weight is selected, a reasonable weight reduction coefficient needs to be set (such as setting it to 0.1, indicating that the regional pixel weight is reduced to one tenth of the original), and then, according to the coordinate information of the region, in the subsequent target tracking related calculations, the pixels in the region are weighted according to the set weight reduction coefficient, so that the influence of the region in the overall calculation is greatly reduced.
[0233] It can be seen that the video image is divided into several areas and compared with the target feature range. For the areas that do not meet the requirements, the constant false alarm detection technology is used to adjust the target feature range. This regional processing and dynamic adjustment method reduces unnecessary calculations. Because there is no need to perform comprehensive and complex calculations on the entire image, but to focus on the areas where the target may exist and the precise definition of the target feature range, it avoids the indiscriminate high-computation processing of the entire image like the TLD algorithm, so that target tracking can be achieved more efficiently, reducing the demand for computing resources, improving the real-time performance of the algorithm, ensuring that the target can be tracked quickly and accurately during the operation of the aircraft, and adapting to complex and changing actual scenarios.
[0234] In some embodiments, after step S3025, the method further includes:
[0235] S3026, determining areas other than the non-target areas as potential target areas;
[0236] S3027, using a random fern classifier to score the potential target area;
[0237] In some embodiments, to ensure that the random fern classifier has been fully trained, the training data should contain a large number of target samples and non-target samples, as well as their corresponding various image feature information, so that the classifier can learn the characteristic patterns of targets and non-targets through these data. Then, for each potential target area, extract its various feature information, such as calculating the color histogram of pixels in the area as color features, calculating texture features through grayscale co-occurrence matrix, and extracting shape features using edge detection algorithm. Input these feature information into the trained random fern classifier, the classifier will conduct a comprehensive evaluation of the area based on its internal decision function and learned feature patterns, and output a score value, indicating the possibility that the area contains the target. This score value will serve as an important basis for determining the tracking area in the future, so that target tracking can be more accurately focused on the area most likely to contain the target, improving the accuracy and efficiency of tracking.
[0238] S3028, selecting the area with the highest-scoring preset data as the tracking area;
[0239] In some embodiments, the scores of all potential target areas are sorted, and a common sorting algorithm (such as a quick sorting algorithm) can be used to sort the potential target areas in descending order of scores. Then, according to a pre-set selection rule (such as selecting the first N areas, or selecting areas with scores higher than a certain threshold), the corresponding areas are extracted from the sorted area list, and these areas are the determined tracking areas.
[0240] S3029: Track the target in the tracking area.
[0241] It can be seen that the video image is divided into several areas and compared with the target feature range. For the areas that do not meet the requirements, the constant false alarm detection technology is used to adjust the target feature range. This regional processing and dynamic adjustment method reduces unnecessary calculations. Because there is no need to perform comprehensive and complex calculations on the entire image, but to focus on the areas where the target may exist and the precise definition of the target feature range, it avoids the indiscriminate high-computation processing of the entire image like the TLD algorithm, so that target tracking can be achieved more efficiently, reducing the demand for computing resources, improving the real-time performance of the algorithm, ensuring that the target can be tracked quickly and accurately during the operation of the aircraft, and adapting to complex and changing actual scenarios.
[0242] In some embodiments, after step S306, the method further includes:
[0243] S308, when the target with the highest confidence score is lower than the score threshold or there is no target; extracting feature data of the target in the most recent frame of the video image;
[0244] In some embodiments, after completing the confidence scoring of all targets, if it is found that the target with the highest confidence score has a score lower than the set scoring threshold, or if no target is detected in the video image at all, it is necessary to perform the step of extracting the feature data of the target in the most recent frame of the video image.
[0245] It should be noted that feature extraction has been discussed in detail in step S3012, and related steps can refer to step S3013, which will not be repeated here.
[0246] S309, determining a search area based on target speed and target location information;
[0247] In some embodiments, the possible moving distance of the target during this period is calculated based on the speed of the target, the frame rate of the video image, and a certain time prediction step. Then, with the current position information of the target as the center, according to the calculated possible moving distance, a certain range is extended to the surroundings to determine the search area. For example, the search area can be set as a circular area with the target center coordinates as the center and a radius equal to the possible moving distance plus a certain margin, or set as a square area with the target position as the center and a side length equal to twice the possible moving distance, etc., so as to determine a reasonable search area, which is convenient for further narrowing the range and finding the target in the future.
[0248] S310, narrowing the search area according to the moving direction of the target to obtain an area of interest;
[0249] In some embodiments, the corresponding angle range is determined according to the direction of movement of the target. For a circular search area, the region of interest can be determined based on the angle range and the fan-shaped area corresponding to the center angle of the circle. The fan-shaped part of the circle that does not conform to the direction of movement is removed, and the fan-shaped area where the target may appear is retained as the region of interest.
[0250] S311, extracting a backup target in subsequent video images according to the coordinates of the region of interest;
[0251] In some embodiments, the coordinate information of the region of interest in the video image is obtained, including the coordinates of the upper left corner and the lower right corner, so as to determine the specific position of the region of interest in the entire video image. Then, for subsequent video image frames (which may be the next frame or the next several frames, determined according to actual conditions), in these frame images, only the pixel range corresponding to the region of interest is subjected to target detection operation, and the image data in the region of interest is input into the target detection algorithm using a suitable target detection algorithm. The algorithm will output the possible target information detected in the region, and these detected targets are used as backup targets. At the same time, their relevant features are extracted, and the backup targets and their feature information are sorted out to prepare for the subsequent confidence score, so as to further identify whether these backup targets are real targets.
[0252] S312, scoring the confidence of the backup target and the target;
[0253] In some embodiments, it is necessary to obtain relevant feature information of the original target, including previously accumulated appearance features (such as color, shape, texture features, etc.), position features (such as the target's motion trajectory, position coordinates before disappearance, etc.) and previous confidence scores, etc., and these information are used as basic data for comparison. Then, for each backup target, its corresponding appearance, position and other features are also extracted, and a suitable similarity calculation method is used. For example, for appearance features, the structural similarity index (SSIM) can be used to calculate the similarity with the original target, and for position features, the distance deviation from the predicted position of the original target can be quantitatively evaluated (such as calculating the Euclidean distance). These similarity evaluation results are combined with certain weight allocations (for example, the appearance feature similarity weight is set to 0.6 and the position feature similarity weight is set to 0.4 based on experience or experimental data), and the confidence score of each backup target is calculated by weighted summation and other methods. These scores are recorded for subsequent judgment of whether the backup target meets the requirements, so as to retrieve the target that may be lost and ensure the continuity and accuracy of target tracking.
[0254] S313: Determine the backup target whose confidence score is higher than the score threshold as a lost target.
[0255] It can be seen that when the target with the highest confidence score is below the score threshold or there is no target, the search area is determined based on the target speed and position information by extracting the target feature data of the most recent frame, and the region of interest is narrowed down based on the direction of motion. In this area, the backup target is extracted and the confidence score is performed to determine the lost target. This method avoids aimless search and calculation in the entire image space like the TLD algorithm, but narrows the search range in a targeted manner based on the target's motion characteristics, reducing the amount of calculation. It ensures that the aircraft can still efficiently relocate the target when the target is temporarily lost or difficult to detect.
[0256] It should be noted that in actual application scenarios, when it comes to determining whether the backup target is a truly lost target, the judgment of the response value is a crucial link. Since the target in the actual environment is easily blocked by background elements (such as other objects in the environment, objects similar in appearance to the target, etc.), this puts higher requirements on the algorithm, that is, it needs to be able to accurately determine the occlusion situation.
[0257] At present, the maximum response value obtained by the relevant operation is usually directly used as the basis for determining the tracking confidence, so as to judge whether the target is lost or determine whether the current frame needs to perform an update operation on the filter template. However, from a practical point of view, the relevant response of the filter often does not present a standard Gaussian distribution. This means that in the actual tracking process, the maximum response value may be at a low level, but the target is actually still in a correct tracking state; conversely, there may also be an abnormal situation where the maximum response value is high, but the target has been lost. Therefore, in some embodiments, after step S313, the method further includes:
[0258] S314, calculating the average peak correlation energy of the backup targets whose confidence scores are higher than the score threshold;
[0259] S315, if the average peak energy is less than the preset energy threshold, canceling the backup target whose confidence score is higher than the score threshold and determining it as a lost target;
[0260] The calculation function of the average peak correlation energy is:
[0261]
[0262] Where APCE is the average peak correlation energy, is the maximum response value, is the minimum response value, is the response value of pixel w and pixel h.
[0263] It can be seen that after determining the lost target, its average peak correlation energy is calculated and compared with the preset energy threshold to further confirm the authenticity of the lost target. This avoids learning a large amount of interference information. When the target is blocked by the background or interfered by similar targets, it can more accurately determine whether the target is truly lost, improve the accuracy and stability of tracking, ensure the aircraft's continuous and effective tracking of the target in complex environments, and reduce tracking failures caused by misjudgment.
[0264] S303, obtaining positive samples and negative samples according to the target detection results, target tracking results, and video images, where the positive samples are data containing the target in the target detection results and target tracking results, and the negative samples are data not containing the target in the video images;
[0265] In some embodiments, after obtaining the target detection results and target tracking results, the acquisition of positive samples and negative samples begins. First, based on the position coordinate information of the target in the target detection result, the image data of the area where the target is located is cropped from the video image. These cropped data are part of the positive sample, and the feature description of the target in the area (such as color feature vector, texture feature vector, etc.) is extracted, and together with the cropped image data, a complete positive sample is formed. Then, for negative samples, other areas in the video image except the area where the target is located are traversed, and image data of some areas are selected as negative samples according to certain rules (such as uniform sampling, random sampling, etc.), or according to the size, shape and other features of the target, those areas with similar target features are excluded, and more representative non-target area image data are selected as negative samples to ensure that the negative sample can fully reflect the characteristics of the background and non-target objects, and the acquired positive and negative samples are sorted out and prepared for subsequent model training.
[0266] S304, training the target detection update model according to the positive samples and the negative samples;
[0267] In some embodiments, to select a suitable model structure, a common one may be a structure based on a convolutional neural network (CNN), such as a classic network structure such as ResNet, and then add or modify some layers (such as adding an output layer for outputting target detection results, etc.) based on the actual target detection task requirements. Next, the positive samples and negative samples are divided into a training set and a validation set according to a certain ratio (such as 80% positive samples and 20% negative samples, and the specific ratio can be adjusted according to the actual situation). The training set is used for model parameter learning, and the validation set is used to evaluate the performance of the model on unseen data. The positive and negative sample image data in the training set are input into the model, and the model's prediction results for the target are obtained through forward propagation calculation. The difference between the prediction results and the true labels (positive samples are targets, negative samples are non-targets) is measured by calculating the loss function (such as the cross entropy loss function, etc.). The back propagation algorithm is used to adjust the model parameters according to the loss value, such as the convolution kernel weights of the convolution layer and the connection weights of the fully connected layer. Multiple rounds of iterative training are performed on the data in the training set. At the same time, the validation set is used regularly to verify the model performance. The training parameters (such as the learning rate, etc.) are adjusted according to the verification results until the performance of the model on the validation set reaches the expected standard (such as the accuracy reaches a certain value), and the model training update is completed.
[0268] S305, the target detection update model outputs feature data and a target model, the feature data is used to perform target detection on the new video image when a new video image is received, and the target model is used to track the target of the new video image when a new video image is received;
[0269] In some embodiments, when the target detection update model has completed training, it can process the newly input video image and output feature data and target model. First, for a new video image, the model will extract features from the image according to its internal feature extraction layer (such as a convolutional layer) to obtain a feature map of the image, and then process the feature map through a specific algorithm and layer (such as a combination of a fully connected layer or a pooling layer) to convert it into feature data that can be directly used for target detection. These feature data can be output in the form of vectors or matrices for subsequent target detection operations. At the same time, the model will construct a target model based on the target's motion laws, appearance change patterns and other information learned during the previous training process. This target model contains the target's feature parameters and motion parameters in different states. When tracking the target in a new video image, it can predict its position, posture and other information at the next moment based on the target's current state, thereby achieving continuous tracking of the target and outputting the feature data and target model for application in subsequent target detection and tracking processes.
[0270] S306, performing confidence scoring according to the target detection result and the target tracking result;
[0271] In some embodiments, a scoring method based on weighted summation is used. A detection confidence score of the target is obtained from the target detection result, for example, the target category probability output by a deep learning target detection model is used as the detection confidence score. The tracking stability score of the target is calculated from the target tracking result by calculating the variance of the position coordinates of the target in consecutive frames. If the variance is less than a set threshold, the tracking stability score is high, otherwise it is low. At the same time, the consistency score of the target appearance features is calculated, and the feature similarity of the target in different frames is calculated using an image feature extraction algorithm (such as SIFT features). The higher the similarity, the higher the score. Then, weights are set for these three indicators to calculate the confidence score for each target.
[0272] S307: Determine the target with the highest confidence score as the correct target.
[0273] It can be seen that tracking while detecting the target shortens the delay time between detection and tracking. After obtaining the target detection results and target tracking results, positive samples and negative samples are obtained based on them and video images. This method can accurately collect target and non-target data, making the training samples more targeted and diverse. Then, the target detection update model is trained using positive and negative samples, so that the model can learn the changing characteristics of the target in real time, and output more accurate feature data and target models for detection and tracking of new video images. Finally, the correct target is determined by confidence scoring. The interaction of these technical features is combined to improve the tracking and detection accuracy of autonomous or semi-autonomous aircraft in complex and changing environments, effectively solving the problem of difficult to accurately track and easy to lose targets due to target changes, and enhancing the stability and reliability of target tracking detection.
[0274] An exemplary aerospace detection device 500 provided in an embodiment of the present application is introduced below. Figure 5 It is a schematic diagram of an exemplary hardware structure of an aerospace detection device 500 provided in an embodiment of the present application.
[0275] In some embodiments, the aerospace detection device 500 is a computer device or the aerospace detection device 500 includes a computer device. The computer device includes a processor, a memory and a network interface connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with other external terminals or servers through a network connection. In some embodiments, the network interface can be a wired network interface, and in some embodiments, the network interface can also be a wireless network interface. When the computer program is executed by the processor, the method in the embodiment of the present application is implemented.
[0276] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0277] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0278] As used in the above embodiments, the term "when..." may be interpreted to mean "if..." or "after..." or "in response to determining..." or "in response to detecting...", depending on the context. Similarly, the phrases "upon determining..." or "if (the stated condition or event) is detected" may be interpreted to mean "if determining..." or "in response to determining..." or "upon detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)", depending on the context.
[0279] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk), etc.
[0280] Those skilled in the art can understand that to implement all or part of the processes in the above-mentioned embodiments, the processes can be completed by computer programs to instruct related hardware, and the programs can be stored in computer-readable storage media. When the programs are executed, they can include the processes of the above-mentioned method embodiments. The aforementioned storage media include: ROM or random access memory RAM, magnetic disk or optical disk and other media that can store program codes.
Claims
1. A target recognition, tracking and guidance method for an optoelectronic pod, characterized in that: include: Perform super-pixel segmentation on the acquired scanned image information and remove the background to obtain the super-pixel area; Performing target recognition on the superpixel area to identify the target; Get the original position and attitude information of the sensor; Calculating the three-dimensional coordinates of the target in the three-dimensional space based on the position information of the target in the scanned image information, using the distance information from the target to the sensor, the original position and the posture information; Obtain feature information for each target; Constructing a plurality of different motion trajectories for each target according to the feature information and the dynamic change data of the three-dimensional coordinates over time; Using the new three-dimensional coordinates received subsequently, the movement hypothesis is verified for relevance; Select the motion trajectory with the highest correlation as the trajectory data; Inputting the feature information and the trajectory data into a target recognition model to obtain the identity of the target; Determine whether the target meets the tracking conditions based on the preset target list and the target's identity; If the target meets the tracking conditions, the target is tracked; The target's position information and angular velocity information are output to guide the execution of the strike process.
2. The method according to claim 1, characterized in that After the step of selecting the motion trajectory with the highest correlation as the trajectory data, the method further includes: When the three-dimensional coordinates of the current target are not updated within a preset time, feature comparison is performed on feature information of subsequent targets in subsequent scanned image information according to feature information of the current target; the current target is any target, and the subsequent scanned image information is scanned image information obtained subsequently to the scanned image information of the current target; The subsequent target whose feature comparison result is greater than the similarity threshold and has the highest similarity is determined as the lost target of the current target; the trajectory data of the current target and the trajectory data of the lost target are spliced, wherein the missing trajectory data is supplemented by the difference.
3. The method according to claim 1, characterized in that The step of constructing a plurality of different motion trajectories for each target according to the feature information and the dynamic change data of the three-dimensional coordinates over time specifically includes: Calculate the target's velocity components in each coordinate axis direction based on the target's three-dimensional coordinate changes at consecutive time points; Calculating the acceleration component of the target in each coordinate axis direction according to the change of the velocity component at different times; Construct motion hypothesis space by changing acceleration, velocity and direction angle; The next motion state parameters of the target are extracted from the motion hypothesis space to generate multiple motion hypotheses that conform to the probability distribution. The motion state parameters include acceleration, speed, and direction angle. When extracting the speed, the current speed of the target is used as the midpoint, and the speed change amount generated by the current acceleration within a preset time step is used as the radius to determine the extraction range. When extracting the acceleration, the current acceleration of the target is used as the midpoint, and the acceleration change amplitude is used as the radius to determine the extraction range. When extracting the direction angle, the current direction angle of the target is used as the midpoint, and the maximum steering angle is used as the radius to determine the extraction range.
4. The method according to claim 1, characterized in that: Before the step of inputting the feature information and the trajectory data into a target recognition model to obtain the identity of the target, the method further includes: The historical feature information, historical trajectory data, and corresponding historical identities are used as data sets; Clustering is used to cluster historical feature information and historical trajectory data into several data subsets according to historical identities; The target recognition model is trained sequentially according to the data subsets, so that the target recognition model can recognize the historical identity according to the historical feature information and the historical trajectory data; The target recognition model is trained according to the data set, so that the target recognition model can recognize multiple historical identities according to multiple historical feature information and multiple historical trajectory data.
5. The method according to claim 1, characterized in that The step of performing target recognition on the superpixel area to identify the target specifically includes: Calculate the gradient magnitude and direction of each pixel in the superpixel area, wherein the gradient magnitude is the intensity of grayscale change of the pixel, and the direction is the direction in which the grayscale changes fastest; Perform non-maximum suppression on the gradient amplitude, retaining only the point with the maximum amplitude in the direction, and obtaining the edge; Determining a first threshold and a second threshold according to the statistical distribution of the gradient amplitude, wherein the first threshold is greater than the second threshold; determining an edge whose gradient amplitude is higher than the first threshold as a strong edge; Determine an edge whose gradient magnitude is lower than the second threshold as a weak edge; Connecting the strong edge with the corresponding weak edge to obtain a time-sensitive target; The target is detected by performing local image feature detection on the time-sensitive target to obtain the target.
6. The method according to claim 1, characterized in that After the step of obtaining the original position and posture information of the sensor, the method further includes: Fusing the thermal imaging image and the scanning image information to obtain a fused image; The step of calculating the three-dimensional coordinates of the target in the three-dimensional space based on the position information of the target in the scanned image information, using the distance information from the target to the sensor, the original position and the posture information, specifically includes: The three-dimensional coordinates of the target in the three-dimensional space are calculated based on the position information of the target in the fused image, using the distance information from the target to the sensor, the original position and the posture information.
7. An aerospace detection device, characterized in that: The aerospace detection device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the aerospace detection device to perform the method described in any one of claims 1-6.
8. A computer program product comprising instructions, characterized in that When the computer program product is run on an aerospace detection device, the aerospace detection device is caused to perform the method according to any one of claims 1 to 6.
9. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on an aerospace detection device, the aerospace detection device is caused to perform the method according to any one of claims 1 to 6.