A target positioning method and system based on aerial survey images

By combining OSD technology and deep learning algorithms with structured light projection, the problems of recognition accuracy and adaptability of flight target localization methods in diverse and complex environments have been solved, achieving high-precision and low-cost target localization.

CN120356119BActive Publication Date: 2025-11-28GUANGDONG RES INST OF WATER RESOURCES & HYDROPOWER +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510398069.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-11-28
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

Existing aerial imagery-based methods for locating flying targets suffer from unstable recognition performance, low accuracy, and insufficient adaptability when faced with diverse and complex problems.

Method used

By combining OSD technology with deep learning algorithms and structured light technology, high-precision positioning of the target object is achieved through image acquisition, recognition, structured light projection, and 3D position data integration.

Benefits of technology

It improves the accuracy and adaptability of target location identification, reduces equipment cost and power consumption, and ensures the stability and real-time performance of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356119B_ABST
    Figure CN120356119B_ABST
Patent Text Reader

Abstract

The application discloses a target positioning method and system based on aerial survey images, and relates to the technical field of image recognition. The method comprises the following steps: acquiring aerial video through an image acquisition device of an aircraft, and acquiring a plurality of aerial image sequences from the aerial video; acquiring target information in the plurality of aerial image sequences through an image recognition model; setting a structured light projection area according to a target pixel position set in the target information, and acquiring a structured light image of the structured light projection area; determining three-dimensional position data of the target in a world coordinate system according to the structured light image and the target pixel position set; and integrating the three-dimensional position data in the world coordinate system to the aerial video and displaying the same. The three-dimensional position data of the target in the world coordinate system is determined through the structured light image and the target pixel position set, so that the recognition precision of the flight target positioning is improved, and the adaptability of the application is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image recognition, and particularly relates to a target positioning method based on aerial survey images and a system thereof. BACKGROUND

[0002] Flight target positioning refers to a method of detecting, tracking, identifying and positioning targets (such as aircraft, drones, missiles, etc.) flying in the air through various sensors and algorithms. These methods usually combine signal processing, computer vision, machine learning and data fusion technologies to achieve efficient and accurate positioning of targets.

[0003] With the continuous progress of artificial intelligence and machine learning technologies, flight target positioning methods have been widely applied in many fields, including in civil aviation to ensure flight safety, prevent air collisions, and monitor drones, identify and track the flight trajectory of drones, prevent illegal flight and intrusion; in environmental monitoring, using aircraft (such as weather balloons, drones) for meteorological monitoring, monitoring wild animals, plant growth, etc. for ecological environment change monitoring; in natural disaster rescue, using unmanned aerial vehicles to quickly locate trapped personnel or disaster areas and missing ships or personnel at sea, providing rescue and search support. In addition, it is also widely used in scientific research, including topographic survey, resource investigation, and animal migration patterns, habitats, etc.

[0004] OSD (Object Spatial Distribution) technology is a method for analyzing and processing the distribution characteristics of targets in space. This technology models and analyzes the position, shape, motion trajectory, etc. of targets in a specific space to achieve target recognition, tracking and positioning. OSD technology has wide applications in computer vision, robotics, unmanned driving, etc., especially in dynamic environments for target detection and tracking. OSD technology has broad application prospects in flight target positioning methods, and through in-depth analysis of target spatial distribution, it can improve the accuracy and efficiency of target detection, tracking and positioning. With the continuous progress of technology, OSD technology will play an important role in more fields and promote the development of flight target recognition technology.

[0005] Although the flight target positioning method based on the OSD (Object Space Distribution) technology has many advantages in practical application, it also faces some technical problems. Due to the diversity of flight targets, such as multiple aircrafts operating simultaneously, multi-target tracking, and frequent changes in the relative positions of targets, and the complexity of the environment, such as complex terrain environment (such as cities, mountains, oceans, etc.), different environmental characteristics (such as light, weather changes, etc.), and other dynamic changes in the environment. The target positioning method based on aerial images faces the problems of unstable recognition effect, low recognition accuracy, and insufficient adaptability.

[0006] Therefore, it is urgent to propose a flight target positioning method with high recognition accuracy, stable effect, and strong adaptability. SUMMARY

[0007] The present application provides a target positioning method based on aerial images and a system thereof to solve at least one problem mentioned in the background art.

[0008] The specific technical solutions provided by the present application are as follows:

[0009] A target positioning method based on aerial images, comprising the steps of:

[0010] S1, acquiring aerial video through the image acquisition device of the aircraft, and acquiring a plurality of aerial image sequences from the aerial video;

[0011] S2, obtaining target information in the plurality of aerial image sequences through an image recognition model;

[0012] S3, setting a structured light projection area according to the target pixel position set in the target information, obtaining a structured light image of the structured light projection area, and determining the three-dimensional position data of the target in the world coordinate system according to the structured light image and the target pixel position set;

[0013] S4, integrating the three-dimensional position data in the world coordinate system into the aerial video and displaying;

[0014] Wherein, step S2 specifically comprises the following steps:

[0015] Each frame of image in the aerial image sequence is input into the image recognition model, and the image is preprocessed;

[0016] Each image is detected using a deep learning algorithm to identify all target objects in the image;

[0017] According to the identified results, the target information of each target object is output, including: target object identification, target object type, target object pixel position and confidence score; wherein the confidence score represents the confidence degree of the model to the recognition result;

[0018] The target object information identified in each image is stored.

[0019] As a preferred solution, step S2 further comprises comparing the target object information, and specifically comprises the following steps:

[0020] Extracting all target object information of each image and comparing with the target object information in subsequent images;

[0021] Once it is confirmed that the target object is the same object, the motion trajectory of the target object is calculated by recording the pixel position and shooting time of the target object in different images.

[0022] As a preferred solution, the calculation of the motion trajectory of the target object by recording the pixel position and shooting time of the target object in different images specifically comprises the following steps:

[0023] Recording the position of the target object in consecutive frames to obtain position change data;

[0024] Calculating the motion of the target object between different time points according to the shooting time of the image to obtain time difference data;

[0025] Using the formula: speed = position change data / time difference data, to calculate the motion speed of the target object.

[0026] As a preferred solution, step S1 specifically comprises the following steps:

[0027] S11, obtaining the flight path of the aircraft, including the size of the flight target area, the flight height and the shooting waypoint;

[0028] S12, setting the collection parameters according to the obtained flight path, the collection parameters including the image resolution and shooting frequency of the image collection device;

[0029] S13, performing image collection to obtain aerial video, and obtaining a plurality of aerial image sequences from the aerial video;

[0030] S14, preliminary checking and evaluation of the obtained aerial image sequences and classification and naming.

[0031] As a preferred solution, in step S13, the state of the aircraft and the image collection are further monitored in real time by a ground control station, and the collected aerial video and aerial image sequences are saved in real time to a local storage device; the collected image sequences are also backed up to other storage devices.

[0032] As a preferred solution, step S3 specifically comprises the following steps:

[0033] S31, using an image processing algorithm to extract the feature points of the target object, and saving the two-dimensional feature pixel coordinates in the image;

[0034] S32, setting a target region according to the two-dimensional feature pixel coordinates, and projecting a structured light pattern to the target region;

[0035] S33, locating and aligning with the target object pixel position set according to the captured structured light image; obtaining structured light deformation information by analyzing the structured light image, and calculating the depth value of each feature point according to the structured light deformation information;

[0036] S34, obtaining three-dimensional coordinates in the camera coordinate system according to the two-dimensional feature pixel coordinates and the depth information by using the camera calibration parameters;

[0037] S35, converting the three-dimensional coordinates from the camera coordinate system to the world coordinate system by using the camera extrinsic parameters;

[0038] S36, performing Bezier interpolation between the feature points in the three-dimensional space, and updating the three-dimensional coordinates of the target object in the world coordinate system.

[0039] As a preferred solution, step S31 specifically comprises:

[0040] Integral image and Haar wavelet response: generating an integral image for the image, and using Haar wavelet response to detect the fast-changing area of the image;

[0041] Interest point detection: detecting key interest points in the scale space using the Hessian matrix determinant;

[0042] Direction assignment: calculating the Haar wavelet response around the interest point, and assigning the main direction;

[0043] Feature description: extracting features around the interest point, and constructing a feature vector using the Haar wavelet response;

[0044] Feature matching: matching the feature vectors in the scene image and the target object image, identifying the most similar feature point pair, and determining the position of the target object through feature matching.

[0045] As a preferred solution, the Bezier interpolation between the feature points in the three-dimensional space comprises the steps of:

[0046] determining the known points, control points and Bezier curves; the known points are represented as and , the control points are represented as , and the Bezier curves are represented as:

[0047] ;

[0048] wherein t is a control parameter, and the value range is [0, 1];

[0049] A plurality of intermediate interpolation points are generated by setting control parameters of the Bezier curve; the coordinates of the intermediate interpolation points are represented as ;

[0050] The coordinates of the intermediate interpolation points are obtained by calculating the coordinate components of the Bezier curve in the three-dimensional space, and are represented as:

[0051]

[0052]

[0053] ;

[0054] wherein, represents the i-th control parameter, 1≤i≤N, N represents the number of intermediate interpolation points.

[0055] As a preferred solution, step S4 includes matching the three-dimensional position data with the aerial video frame, ensuring that the position of the target object in the video is consistent with the actual position; using the OSD algorithm, visualizing the three-dimensional information of the target object and superimposing it into the aerial video; further including superimposing the three-dimensional position data in the world coordinate system with the standard video information OSD data when the camera fails.

[0056] A target object positioning system based on aerial survey images, applying the aforementioned target object positioning method based on aerial survey images, comprising: an image acquisition module, an image recognition module, a data processing module and a data generation module connected in communication; wherein,

[0057] The image acquisition module is configured to obtain aerial video through the image acquisition device of the aircraft, and obtain a plurality of aerial image sequences from the aerial video;

[0058] The image recognition module is configured to obtain target object information in the plurality of aerial image sequences through the image recognition model;

[0059] The data processing module is configured to set a structured light projection area according to the target object pixel position set in the target object information, obtain a structured light image of the structured light projection area, and determine three-dimensional position data of the target object in the world coordinate system according to the structured light image and the target object pixel position set;

[0060] The data generation module is configured to integrate and display the three-dimensional position data in the world coordinate system into the aerial video.

[0061] Compared with the prior art, the present application has the following advantages:

[0062] The target positioning method based on aerial survey images provided by the present application replaces the laser ranging technology with the OSD technology to position and identify the target, thereby saving the cost, power consumption and transmission time by eliminating other superimposed devices; meanwhile, the image recognition model can accurately extract the detailed information of the target from the aerial image sequence through the similarity comparison, thereby improving the identification accuracy and adaptability of the flight target positioning and identification. In addition, the three-dimensional position data of the target in the world coordinate system is determined through the structured light image and the target pixel position set, and the three-dimensional position data is integrated into the aerial video for display, which can further improve the identification accuracy and adaptability of the flight target positioning and identification. BRIEF DESCRIPTION OF DRAWINGS

[0063] The drawings incorporated into the specification and forming a part of the specification, illustrate embodiments consistent with the present application and, together with the specification, serve to explain the principles of the present application.

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced as follows, and obviously, other drawings can be obtained by those skilled in the art without creative labor.

[0065] Figure 1 The flowchart of the target positioning method based on aerial survey images provided by the present application embodiment 1 is shown in the figure.

[0066] Figure 2 The flowchart of the target information acquisition in step S2 of the present application embodiment 1 is shown in the figure.

[0067] Figure 3 The flowchart of step S3 of the present application embodiment 1 is shown in the figure.

[0068] Figure 4 The structural diagram of the target positioning system based on aerial survey images provided by the present application embodiment 2 is shown in the figure. DETAILED DESCRIPTION

[0069] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0070] It should be noted that all the direction indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present application are only used to explain the relative position relationship, movement condition, etc. between components in a certain posture (as shown in the drawings), and if the certain posture changes, the direction indications will also change accordingly.

[0071] In addition, the descriptions involving "first", "second", etc. in the present application are only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" can explicitly or implicitly include at least one of the features. In addition, the technical solutions of various embodiments can be combined with each other, but it must be based on the fact that a person skilled in the art can realize it, and when the combination of technical solutions contradicts each other or cannot be realized, it should be considered that the combination of technical solutions does not exist and is not within the protection scope required by the present application.

[0072] The present application provides a target positioning method and system based on aerial survey images, and the specific embodiments of the present application will be described in detail below.

[0073] Embodiment 1

[0074] Please refer to Figure 1 The present application provides a target positioning method based on aerial survey images, comprising the steps of:

[0075] S1, acquiring aerial video through the image acquisition device of the aircraft, and acquiring a plurality of aerial image sequences from the aerial video; wherein the image acquisition device of the aircraft includes a high-definition camera, an infrared camera or a multispectral camera carried by a drone, etc.

[0076] Step S1 specifically includes the following steps:

[0077] S11, acquiring the flight path of the aircraft, including the size of the flight target area, the flight height and the shooting waypoint. Specifically, according to the size and complexity of the acquired flight target area, ensure that the flight route covers all the areas that need to be collected; dynamically adjust the flight height according to the size and type of the flight target area to ensure the best image quality; in combination with the shooting waypoint, ensure that the image is shot at each waypoint for a certain period of time.

[0078] S12, setting the collection parameters according to the acquired flight path, the collection parameters including the image resolution and the shooting frequency of the image acquisition device. The shooting frequency can usually be set to one image per second, or adjusted according to the flight speed and the dynamic characteristics of the target object, which is not limited in the present application.

[0079] S13, image acquisition is performed to obtain aerial video, and a plurality of aerial image sequences are obtained according to the aerial video. Specifically, the aircraft performs image acquisition according to the obtained flight path and the set acquisition parameters.

[0080] In step S13, the state of the aircraft and the image acquisition are also monitored in real time by the ground control station to ensure the integrity and effectiveness of the data; the acquired aerial video and aerial image sequences are saved in real time to a local storage device (such as an SD card, an internal storage, etc.); further, the acquired image sequences can also be backed up to other storage devices (such as an external hard disk or cloud storage) to prevent data loss.

[0081] S14, the plurality of aerial image sequences obtained are preliminarily checked and evaluated, and are classified and named, the completeness of the image sequences is checked to ensure that there is no missing frame; the image quality is evaluated to check whether there are problems such as blurring, overexposure or underexposure, etc.

[0082] Through the above steps, the image acquisition device of the aircraft can effectively obtain a plurality of aerial image sequences, providing basic data for subsequent target object recognition and positioning, and ensuring the accuracy and reliability of the data.

[0083] S2, target object information in the plurality of aerial image sequences is obtained through an image recognition model;

[0084] The image recognition model can be an existing pre-trained model, or an existing image recognition model selected according to task requirements, such as a convolutional neural network (CNN), YOLO (You Only Look Once), Faster R-CNN, etc., which is not limited in the present application.

[0085] The target object information includes target object identification, target object type, target object pixel position set and target object speed. Further,

[0086] The target object identification refers to assigning a unique identifier (ID) to each recognized target object, which can be generated according to the category and detection order of the target object;

[0087] The target object type is a static target object or a dynamic target object, the static target object can be a building, a geographical landmark (such as a mountain peak, a tree, a river, etc.), etc., and the dynamic target object can be a vehicle, a ship, a person, another aircraft, etc.;

[0088] The target object pixel position set records the pixel coordinates of the target object in the image;

[0089] Target speed: the movement speed of the target is calculated by the change of pixel position of the target in consecutive frames, for example, the pixel displacement can be obtained by comparing the pixel position of the same target in the current frame and the previous frame; the pixel displacement is converted into actual distance according to the flight speed and height of the aircraft, and the speed of the target is calculated.

[0090] Please refer to Figure 2 , step S2 specifically includes the following steps:

[0091] S21, obtaining target information, specifically including:

[0092] Step 211, inputting each frame image in the aerial image sequence into the image recognition model;

[0093] Step 212, using a deep learning algorithm (such as convolutional neural network) to detect targets in each image, and identifying all targets in the image.

[0094] Step 213, outputting target information of each target according to the identified result, the target information including: target identification, target type, target pixel position and confidence score; wherein the confidence score represents the confidence degree of the model to the identification result, and is usually a value between 0 and 1.

[0095] Step 214, storing the target information identified in each image (such as dictionary or database) for subsequent retrieval and comparison.

[0096] S22, comparing target information, specifically including:

[0097] Step 221, extracting all target information of each image and comparing with the target information in the subsequent image; wherein the comparison includes comparing target type consistency, target pixel position similarity and confidence score. Among them, the target of the same category is more likely to be the same object; the position similarity of the targets in different images is evaluated by calculating the intersection over union (IoU) of the bounding boxes of the targets, and the higher the position similarity is, the more likely it is considered to be the same object; the target with high confidence is more likely to be considered as the same object. Set a threshold (confidence), if the similarity exceeds the threshold, it is considered that the two targets are the same object.

[0098] Step 222, once it is confirmed that the target is the same object, its movement trajectory can be calculated by recording the pixel position and shooting time of the target in different images. The calculation process includes:

[0099] Recording the position of the target in consecutive frames to obtain position change data;

[0100] According to the shooting time of the image, the movement of the target between different time points is calculated to obtain time difference data;

[0101] The movement speed of the target object is calculated using the formula: speed = position change data / time difference data.

[0102] The present application ensures that the same target object is consistently identified in different images through similarity comparison, reduces misidentification and missed identification, and improves identification accuracy. It also has high real-time tracking capability, can calculate the movement trajectory of the target object in real time, and provides support for dynamic monitoring, suitable for traffic management, public security and other scenarios. In addition, the target object information in multiple images is integrated to form complete target object movement data, which is convenient for subsequent analysis and decision support.

[0103] Further, in some embodiments of the present application, step S211 further comprises image preprocessing on the plurality of aerial image sequences. The image preprocessing includes:

[0104] Step 2111, denoising, using image processing algorithms (such as Gaussian blur or filtering algorithms) to remove noise in the plurality of aerial image sequences to improve the accuracy of subsequent identification;

[0105] Step 2112, normalization, adjusting the brightness and contrast of the plurality of aerial image sequences after denoising to make them suitable for input into the image recognition model;

[0106] Step 2113, size adjustment, adjusting the image to a uniform size (such as 224x224 pixels) according to the requirements of the image recognition model.

[0107] Through the above step S2, the image recognition model can accurately extract detailed information of the target object from the aerial image sequence, including identification, type, pixel position set and speed, which provides important data support for subsequent target object positioning and analysis, ensuring the accuracy and real-time of identification. At the same time, step S2 significantly improves the identification accuracy and tracking ability of the target object by combining image recognition and target object information comparison.

[0108] S3, setting a structured light projection area according to the target object pixel position set in the target object information, obtaining a structured light image of the structured light projection area; determining the three-dimensional position data of the target object in the world coordinate system according to the structured light image and the target object pixel position set.

[0109] Among them, structured light technology is a technology that projects a specific light pattern (such as stripes, dot matrix, etc.) onto the surface of a target object to obtain its three-dimensional shape and surface characteristics. Its core principle is to use the reflection and deformation of light, by analyzing the deformation of the projected pattern on the surface of the object, to calculate the three-dimensional structure of the object.

[0110] The setting of the structured light projection region includes calculating a suitable structured light projection region according to the pixel position of the target object. The structured light projection region should cover the entire range of the target object to ensure that the structured light can be effectively projected onto the target object. For example, a certain margin can be added on the basis of the bounding box of the target object to ensure that the projection region is large enough.

[0111] The setting of the structured light projection region also includes selecting a suitable projection region geometry (such as a rectangle, a circle, etc.) according to the shape and characteristics of the target object. Generally, a rectangular region is the most commonly used choice.

[0112] Further, referring to Figure 3 , the step S3 specifically includes the steps of:

[0113] S31, extracting feature points of the target object using an image processing algorithm and saving two-dimensional feature pixel coordinates in the image. Specifically, it includes:

[0114] Step 311, integral image and Haar wavelet response: generating an integral image to improve the subsequent calculation efficiency. Using Haar wavelet response to detect the fast changing area of the image. For example, significant changing areas are detected in some corners of a scene image.

[0115] Step 312, interest point detection: using Hessian matrix determinant to detect key interest points in the scale space. For example, by calculating the Hessian matrix determinant, the edges and corner points of the objects in the image are identified.

[0116] Step 313, direction assignment: calculating the Haar wavelet response around the interest points and assigning the main direction. An interest point may be assigned a main direction of 30 degrees to ensure its rotation invariance.

[0117] Step 314, feature description: extracting features around the interest points and using Haar wavelet response to construct a feature vector (such as 64 dimensions). The feature vector is used to describe the local features of these interest points, such as the feature vector of a certain interest point may contain values such as [0.05, 0.1, 0.15,..., 0.2].

[0118] Step 315, feature matching: matching the feature vectors in the scene image and the target object image to identify the most similar feature point pairs. Through feature matching, it is assumed that the feature points of the target object are successfully matched to the corresponding positions in the scene image, thereby determining the position of the target object.

[0119] S32, setting a target region according to the two-dimensional feature pixel coordinates and projecting a structured light pattern to the target region; the structured light pattern can be a stripe pattern or a dot array pattern.

[0120] S33, positioning the target object pixel position set according to the captured structured light image; obtaining structured light deformation information by analyzing the structured light image, and calculating the depth value of each feature point according to the structured light deformation information;

[0121] S34, obtaining three-dimensional coordinates in the camera coordinate system according to the two-dimensional feature pixel coordinates and the depth information by using the camera calibration parameters; the three-dimensional coordinates in the camera coordinate system are represented as:

[0122] ;

[0123] wherein, is the three-dimensional coordinates in the camera coordinate system, is the camera intrinsic matrix, (x, y) is the two-dimensional feature pixel coordinates, and Z is the depth value.

[0124] S35, converting the three-dimensional coordinates from the camera coordinate system to the world coordinate system by using the camera extrinsic parameters (including the rotation matrix R and the translation vector T); the three-dimensional coordinates in the world coordinate system are represented as:

[0125] ;

[0126] wherein, is the three-dimensional coordinates in the world coordinate system, R is the rotation matrix of the camera extrinsic parameters, is the translation vector of the camera extrinsic parameters.

[0127] S36, performing Bezier interpolation between the feature points in the three-dimensional space to update the three-dimensional coordinates of the target object in the world coordinate system.

[0128] wherein, the Bezier interpolation uses a Bezier curve to construct an interpolation point. For example, the known points obtained through the foregoing steps are represented as and , and the control points are generated to construct a second-order Bezier curve. It can be understood that a higher-order Bezier curve can also be constructed by using more control points. The Bezier interpolation between the feature points in the three-dimensional space includes the following steps:

[0129] S361, determining the known points, the control points, and the Bezier curve; in an embodiment, the known points are represented as and , the control points are represented as ; the Bezier curve is a second-order Bezier curve, and is represented as:

[0130] ;

[0131] wherein, t is a control parameter, and the value range is [0, 1];

[0132] S362, generating a plurality of intermediate interpolation points by setting control parameters of the Bezier curve; the coordinates of the intermediate interpolation points are represented as .

[0133] wherein the coordinates of the intermediate interpolation points are The coordinates of the Bezier curve in three-dimensional space are obtained by calculation, and are represented as:

[0134]

[0135]

[0136] ,

[0137] wherein, represents the i-th control parameter, 1≤i≤N, N represents the number of intermediate interpolation points.

[0138] The curve generated by the Bezier interpolation has natural smoothness and continuity, and can better describe the shape in space. At the same time, due to the high-order property of the Bezier curve, the interpolation points can more accurately describe the small curve changes, and improve the accuracy of the feature point coordinates. In addition, by adjusting different control points and (t) values, various interpolation points can be generated flexibly to meet the needs of different applications.

[0139] S4, integrating the three-dimensional position data in the world coordinate system into the aerial video and displaying. Specifically, it includes matching the three-dimensional position data with the aerial video frame, ensuring that the position of the target object in the video is consistent with the actual position. Using the OSD algorithm, the three-dimensional information of the target object is visualized and superimposed into the aerial video for observation and analysis.

[0140] wherein the three-dimensional position data is visualized in the aerial video. It can be selected to add marks (such as points, arrows or frames) in the video frame to represent the position of the target object.

[0141] In step S4, when the FPV (First Person View) camera fails, we need to superimpose the three-dimensional position data in the world coordinate system with the standard video information OSD (Object Spatial Distribution) data to ensure the safety and effective operation of the aircraft. When the FPV (First Person View) camera fails, the aerial video cannot be displayed, at this time, by superimposing the three-dimensional position data in the world coordinate system with the standard video information OSD data, the relevant data of the target object can also be obtained and displayed, and the positioning and identification of the flight target object are performed. Through this step, even in the case of FPV camera failure, we can effectively use the standard video information to superimpose the target spatial distribution (OSD) data. This method ensures the safety and effective operation of the aircraft in the case of failure, and is suitable for various application scenarios such as unmanned aerial vehicle monitoring and flight safety guarantee.

[0142] Further, the standard video information superimposed OSD data adopts H264 encoding mode, including motion compensation and intra prediction. The specific motion compensation and intra prediction techniques can be realized by conventional techniques in the art, which are not limited by the present application.

[0143] The step S4 of the present application realizes the effective integration of the three-dimensional position data of the target object into the aerial video by using the OSD technology. The finally generated video not only shows the aerial image, but also clearly marks the position and related information of the target object, which is convenient for subsequent analysis and decision-making. This method has wide application potential in the fields of unmanned aerial vehicle monitoring, environmental monitoring, traffic management, etc.

[0144] The target object positioning method based on aerial survey image of the present application replaces the laser ranging technology with the OSD technology for positioning and identification of the target object, which saves other superimposed equipment, saves cost, power consumption and transmission time; at the same time, the image recognition model can accurately extract the detailed information of the target object from the aerial image sequence through similarity comparison, which improves the identification accuracy and adaptability of the flight target object positioning and identification. In addition, by determining the three-dimensional position data of the target object in the world coordinate system through the structured light image and the target object pixel position set, and integrating the three-dimensional position data into the aerial video for display, the identification accuracy and adaptability of the flight target object positioning and identification can be further improved.

[0145] Embodiment 2

[0146] Please refer to Figure 4The application further provides a target positioning system based on aerial survey images, and the target positioning method based on aerial survey images is applied to the system, and the system comprises an image acquisition module, an image recognition module, a data processing module and a data generation module which are communicatively connected.

[0147] The image acquisition module is configured to acquire aerial video through an image acquisition device of an aircraft, and acquire a plurality of aerial image sequences from the aerial video, wherein the image acquisition device of the aircraft comprises a high-definition camera, an infrared camera or a multispectral camera carried by a drone.

[0148] The image recognition module is configured to acquire target information in the plurality of aerial image sequences through an image recognition model.

[0149] The data processing module is configured to set a structured light projection area according to a target pixel position set in the target information, acquire a structured light image of the structured light projection area, and determine three-dimensional position data of the target in a world coordinate system according to the structured light image and the target pixel position set.

[0150] The data generation module is configured to integrate the three-dimensional position data in the world coordinate system into the aerial video and display the three-dimensional position data.

[0151] In the embodiments provided in the present application, it should be understood that the disclosed system and method can be implemented in other manners. For example, the system embodiments described above are only schematic; the division of the modules is only a logical function division; there can be another division manner for the actual implementation, for example, a plurality of modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different modules can be indirect couplings or communication connections through some interfaces, and electrical, mechanical or other forms.

[0152] The modules illustrated as separate components can or can not be physically separate, and the components illustrated as modules can or can not be physical modules, i.e., can be located in one place or distributed on a plurality of network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments.

[0153] In addition, each functional module in each embodiment of the present application can be integrated into one processing module, or each module can exist physically, or two or more modules can be integrated into one module. The integrated module can be realized in the form of hardware or in the form of a software functional module.

[0154] The integrated module, if implemented in the form of a software function module and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a dynamic hard disk, a read-only memory (ROM, read-only memory), a random access memory (RAM, random access memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0155] The above description is merely a specific implementation of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features applied herein.

Claims

1. A target location method based on aerial survey imagery, characterized in that, Including the following steps: S1. Acquire aerial video using the aircraft's image acquisition equipment, and extract several aerial image sequences from the aerial video; S2. Obtain target information in the plurality of aerial image sequences through an image recognition model; S3. Set the structured light projection area based on the target object pixel position set in the target object information, and obtain the structured light image of the structured light projection area; determine the three-dimensional position data of the target object in the world coordinate system based on the structured light image and the target object pixel position set. S4. Integrate and display the three-dimensional position data in the world coordinate system into the aerial video. Step S2 specifically includes the following steps: Each frame of the aerial image sequence is input into the image recognition model for image preprocessing; Deep learning algorithms are used to perform object detection on each image, identifying all objects in the image; Based on the identified results, the system outputs target information for each target, including: target identifier, target type, target pixel location, and confidence score; where the confidence score represents the model's confidence in the identification results. Store the target object information identified in each image.

2. The target location method based on aerial survey imagery according to claim 1, characterized in that: Step S2 also includes comparing target object information, specifically including the following steps: Extract all object information from each image and compare it with object information in subsequent images; Once the target object is confirmed to be the same object, its motion trajectory is calculated by recording the pixel position of the target object in different images and the shooting time.

3. The target location method based on aerial survey imagery according to claim 2, characterized in that: The process of calculating the motion trajectory of the target object by recording the pixel position and shooting time in different images specifically includes the following steps: Record the position of the target object in consecutive frames to obtain position change data; The motion of the target object between different time points is calculated based on the image capture time to obtain time difference data; Use the formula: Velocity = Position change data / Time difference data to calculate the velocity of the target object.

4. The target location method based on aerial survey imagery according to claim 1, characterized in that: Step S1 specifically includes the following steps: S11. Obtain the flight path of the aircraft, including the size of the target area, flight altitude, and shooting waypoints; S12. Based on the acquired flight path, set the acquisition parameters, including the image resolution and shooting frequency of the image acquisition device; S13. Perform image acquisition to obtain aerial video, and obtain several aerial image sequences based on the aerial video; S14. Conduct preliminary examination and evaluation of the acquired aerial image sequences, and classify and name them.

5. The target location method based on aerial survey imagery according to claim 4, characterized in that: In step S13, the system also includes real-time monitoring of the aircraft's status and image acquisition through a ground control station, and real-time saving of the acquired aerial videos and aerial image sequences to a local storage device; and backing up the acquired image sequences to other storage devices.

6. The target location method based on aerial survey imagery according to claim 1, characterized in that: Step S3 specifically includes the following steps: S31. Use image processing algorithms to extract feature points of the target object and store the two-dimensional feature pixel coordinates in the image; S32. Set the target area according to the coordinates of two-dimensional feature pixels, and project a structured light pattern onto the target area; S33. Based on the captured structured light image, locate and align with the set of pixel positions of the target object; Structured light deformation information is obtained by analyzing structured light images, and the depth value of each feature point is calculated based on the structured light deformation information. S34. Using camera calibration parameters, obtain the three-dimensional coordinates in the camera coordinate system based on the two-dimensional feature pixel coordinates and depth information; S35. Use camera extrinsic parameters to transform the 3D coordinates from the camera coordinate system to the world coordinate system; S36. Perform Bezier interpolation between feature points in three-dimensional space to update the three-dimensional coordinates of the target object in the world coordinate system.

7. The target location method based on aerial survey imagery according to claim 6, characterized in that: Step S31 specifically includes: Integral Image and Haar Wavelet Response: Generate an integral image of the image and use the Haar wavelet response to detect rapidly changing regions of the image; Interest point detection: Key interest points are detected in scale space using the Hessian matrix determinant; Direction assignment: Calculate the Haar wavelet response around the point of interest and assign the main direction; Feature description: Extract features around the point of interest and construct feature vectors using Haar wavelet responses; Feature matching: Matching feature vectors in scene images and target object images to identify the most similar feature point pairs, and determining the location of the target object through feature matching.

8. The target location method based on aerial survey imagery according to claim 6, characterized in that: The method of performing Bezier interpolation between feature points in three-dimensional space includes the following steps: Determine the known points, control points, and Bézier curve; the known points are represented as... and Control points are represented as The Bézier curve is represented as: , Where t is a control parameter, and its value ranges from [0, 1]. Several intermediate interpolation points are generated by setting the control parameters of the Bézier curve; the coordinates of the intermediate interpolation points are represented as follows: ; The coordinates of the intermediate interpolation points are obtained by calculating the coordinate components of the Bézier curve in three-dimensional space, and are expressed as: in, Let i represent the i-th control parameter, 1≤i≤N, where N represents the number of intermediate interpolation points.

9. The target location method based on aerial survey imagery according to claim 1, characterized in that: Step S4 includes matching the three-dimensional position data with aerial video frames to ensure that the position of the target object in the video is consistent with its actual position; The OSD algorithm is used to visualize the 3D information of the target object and overlay it onto the aerial video. It also includes overlaying OSD data with the 3D position data in the world coordinate system and standard video information when the camera malfunctions.

10. A target location system based on aerial survey imagery, employing a target location method based on aerial survey imagery as described in any one of claims 1-9, characterized in that, include: The system includes a communication-connected image acquisition module, an image recognition module, a data processing module, and a data generation module; among which, The image acquisition module is configured to acquire aerial video through the aircraft's image acquisition device and obtain a series of aerial images from the aerial video; The image recognition module is configured to obtain target information in the plurality of aerial image sequences through an image recognition model; The data processing module is configured to set the structured light projection area based on the target object pixel position set in the target object information, acquire the structured light image of the structured light projection area, and determine the three-dimensional position data of the target object in the world coordinate system based on the structured light image and the target object pixel position set. The data generation module is configured to integrate and display 3D position data in the world coordinate system into aerial video.

Citation Information

Patent Citations

  • Method, system and device for acquiring three-dimensional space data and storage medium

    CN110599546A

  • Method and device for estimating GPS coordinates of multiple target objects and tracking target objects on basis of camera image information about unmanned aerial vehicle

    WO2024096691A1