Method for precise detection and tracking of low-altitude targets by unmanned aerial vehicle based on machine vision

CN122473694BActive Publication Date: 2026-09-11ZHEJIANG STAR GENERAL AVIATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610953401.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-11
Estimated Expiration
2046-06-30

AI Technical Summary

Technical Problem

[0004]本发明的目的是为了解决现有技术中存在的难以区分真实目标的持续响应与背景中偶发出现的离散响应的缺点,而提出的基于机器视觉的无人机低空目标精准检测与跟踪方法

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122473694B_ABST
    Figure CN122473694B_ABST
Patent Text Reader

Abstract

This invention discloses a machine vision-based method for accurate detection and tracking of low-altitude targets by unmanned aerial vehicles (UAVs), relating to the field of machine vision technology. The method includes: setting morphological structural elements based on input images acquired by the UAV, and determining leaf gap aperture images based on these morphological structural elements; performing average filtering on the dilated edge images corresponding to the input images to obtain beaded edge response images; determining the leaf gap beaded edges defined by the leaf gap aperture images and beaded edge response images, and constraining the beaded edge response images within the aperture-beaded co-occurrence region to obtain aperture-constrained beaded evidence images; determining historical deposition images based on the aperture-constrained beaded evidence images, and performing recursive deposition processing on the historical deposition images to obtain the beaded trajectory deposition image at the current moment. This invention improves the accuracy of locating suspected targets in complex occlusion scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of machine vision technology, and in particular to a machine vision-based method for accurate detection and tracking of low-altitude targets by unmanned aerial vehicles (UAVs). Background Technology

[0002] In applications such as search and rescue in mountainous and forested areas, post-disaster life sign investigation, and monitoring of illegal intruders, drones typically fly at low altitudes along forest edges, forest clearings, shrub belts, or forest roads. They continuously acquire video images of the ground surface and vegetation areas using airborne visible light cameras, and perform visual detection and tracking of suspected personnel or abandoned objects. In real-world environments, targets are often partially obscured by foliage, branches, and other structures. Continuous edges can only be observed intermittently through local gaps, thus appearing as multiple separate short edges in the image. These short edges repeatedly appear at similar spatial positions in different time frames, forming a discrete appearance morphology caused by gap perspective, namely the leaf gap discrete edge appearance morphology. Under this morphology, the visible information of the target appears intermittently, spatially dispersed, and discontinuous, and its position shifts and combinations change continuously with the movement of the drone and changes in obstruction.

[0003] Under the condition of discrete edge manifestation in leaf gaps, existing technologies typically perform target tracking based on single-frame detection results combined with temporal correlation strategies. However, since discrete edge segments in different time frames originate from different occlusion gaps, existing methods are prone to misjudging responses that appear sequentially in time but are not the same continuous edge as the continuous manifestation of the same target. This leads to mistaking the temporal superposition of discrete edge segments as real continuous edges, resulting in false targets or the continuation of erroneous trajectories. Existing methods lack effective constraint mechanisms for this discrete manifestation pattern, making it difficult to distinguish between the continuous response of the real target and the occasional discrete response in the background, resulting in insufficient stability of target detection and tracking results. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies that make it difficult to distinguish between the continuous response of a real target and the discrete response that occasionally appears in the background, and to propose a machine vision-based method for accurate detection and tracking of low-altitude targets by unmanned aerial vehicles.

[0005] To address the problems existing in the prior art, the present invention adopts the following technical solution: Machine vision-based methods for accurate detection and tracking of low-altitude targets by drones include: S1. Based on the input image collected by the UAV, set the morphological structure elements, and determine the leaf gap aperture image based on the morphological structure elements. S2. Perform average filtering on the dilated edge image corresponding to the input image to obtain the beaded edge response image; S3. Based on the leaf gap aperture image and the beaded edge response image, the leaf gap beaded edge is defined, the aperture-beaded co-occurrence region is determined, and the beaded edge response image in the aperture-beaded co-occurrence region is constrained to obtain the aperture-constrained beaded evidence image. S4. Based on the aperture-constrained beaded evidence image, determine the historical deposition image, perform recursive deposition processing on the historical deposition image, and obtain the beaded trajectory deposition image at the current moment. S5. Extract connected component attributes from the current bead trajectory deposition image to obtain target tracking information, and generate the UAV target tracking result based on the target tracking information.

[0006] Preferably, morphological structural elements are defined, including: Acquire input images captured by the drone; Convert the input image to a grayscale image; Morphological structural elements are defined based on the pixel resolution of the grayscale image.

[0007] Preferably, determining the leaf slit aperture image includes: Dilation operations are performed on grayscale images and morphological structural elements to obtain dilated images; Erosion operations are performed on grayscale images and morphological structuring elements to obtain eroded images; The dilation and erosion images are differentially processed to obtain the morphological gradient image; The morphological gradient image is normalized to obtain a normalized gradient image; Based on the normalized gradient image, the leaf gap aperture image is determined.

[0008] Preferably, obtaining the beaded edge response image includes: Edge detection is performed on a grayscale image to obtain an edge response image; Perform an erosion operation on the edge response image to obtain the eroded edge image; Perform a dilation operation on the eroded edge image to obtain the dilated edge image; Local window averaging filtering is performed on the dilated edge image to obtain the beaded edge response image.

[0009] Preferably, determining the aperture-bead co-occurrence region includes: In the leaf gap aperture image and the beaded edge response image, identify the discrete short edge segments that simultaneously satisfy the aperture co-occurrence relationship; The edge feature in which discrete short side segments are arranged in a beaded pattern is defined as the leaf gap beaded edge; Based on the co-occurrence characteristics of apertures along the leaf gap beaded edge, the pixel positions of the leaf gap aperture image and the beaded edge response image are divided into aperture-dominant region, beaded-dominant region, and aperture-beaded co-occurrence region.

[0010] Preferably, obtaining the aperture-constrained beaded evidence image includes: Within the aperture-bead co-occurrence region, the bead edge response image is labeled and extracted based on the bead morphology of the leaf gap bead edge to obtain bead point labels, bead string labels, and solitary bead labels. Statistical processing was performed on the bead dot markings, bead string markings, and solitary bead markings to obtain the bead dot density, bead string length ratio, and solitary bead ratio. Based on the bead density, bead string length ratio, and solitary bead ratio, the bead edge response image in the aperture-bead co-occurrence region is constrained to obtain the aperture-constrained bead evidence image.

[0011] Preferably, determining historical sedimentary images includes: Use the aperture-constrained beaded evidence image as the evidence image at the current moment; Based on the image size of the aperture-constrained beaded evidence image, construct an all-zero image; Obtain the deposition image of the bead trajectory from the previous time step; The bead trajectory deposition image from the previous moment is recorded as the historical deposition image; when there is no bead trajectory deposition image from the previous moment, the image with all zeros is recorded as the historical deposition image.

[0012] Preferably, obtaining the bead trajectory deposition image at the current moment includes: Set the deposition recursion coefficient; Based on the depositional recursion coefficient, recursive depositional processing is performed on historical depositional images and evidence images at the current moment to obtain the beaded trajectory depositional image at the current moment.

[0013] Preferably, obtaining target tracking information includes: Connected components are extracted from the current moment's bead trajectory deposition image to obtain connected regions; Target tracking information is obtained based on connected regions.

[0014] Preferably, generating target tracking results for the UAV includes: Extract the non-zero response from the evidence image at the current moment to obtain the support domain; Determine the overlap area between the supporting domain and the connected region; When the overlap area is greater than 0, the target tracking result of the UAV is generated using the target tracking information; When the overlap area is equal to 0, the output of target tracking information is terminated.

[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention constructs leaf gap aperture images and beaded edge response images to characterize the transparent area of ​​vegetation gaps and the discrete edge display area, and further determines the co-occurrence area of ​​aperture and bead. From the source, it limits the subsequent processing to only areas that simultaneously satisfy the aperture display relationship and the edge response relationship, thereby reducing the interference of occasional background responses on target detection and improving the accuracy of suspected target localization in complex occlusion scenes.

[0016] 2. This invention extracts bead point markers, bead string markers, and isolated bead markers from the bead edge response image within the co-occurrence region of the aperture beads, and performs constraint processing by combining bead point density, bead string length ratio, and isolated bead ratio. This can distinguish between effective responses formed by discontinuous appearance of continuous edges and random discrete responses, and avoid mistaking irrelevant edge segments generated under different occlusion gaps as continuous appearance of the same target, thereby reducing the probability of false targets and erroneous trajectory continuation.

[0017] 3. This invention constructs historical deposition images based on aperture-constrained beaded evidence images and performs recursive deposition processing to obtain the beaded trajectory deposition image at the current moment. Then, it combines the overlapping relationship between the support domain and the connected region to output the target tracking result. This not only preserves the continuous display information of the real target in the time dimension, but also terminates invalid output in time when the current response is missing, thereby improving the stability and reliability of the UAV low-altitude target tracking process. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating a machine vision-based method for accurate detection and tracking of low-altitude targets by unmanned aerial vehicles (UAVs) according to an embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0020] This embodiment provides a machine vision-based method for accurate detection and tracking of low-altitude targets by UAVs. (See also...) Figure 1 Specifically, including: S1. Based on the input image collected by the UAV, set the morphological structure elements, and determine the leaf gap aperture image based on the morphological structure elements. In embodiments of the present invention, morphological structural elements are defined, including: Acquire input images captured by the drone; A drone refers to a flight platform that can fly in the air and carry an imaging device to perform image acquisition tasks. In this application, it is used to move along a predetermined area and acquire scene information of the ground surface, vegetation and target objects from the air. The drone is controlled to fly at low altitude along the search area. The onboard camera installed on the drone continuously captures scene images of the ground and vegetation areas. The raw image data output by the onboard camera at the current acquisition time is transmitted to the processing unit. The raw image data is read in complete frames. After confirming that the current frame image data is complete, the frame image data is written to the buffer area according to the image storage format. The current frame image data written to the buffer area is used as the input image for subsequent processing.

[0021] Convert the input image to a grayscale image; The input image refers to the two-dimensional optical imaging result of the scene recorded by the airborne camera on the drone at a certain acquisition time. The imaging result reflects the light intensity and surface appearance distribution at each location in the corresponding spatial area. The grayscale image refers to the image data after converting the color information in the input image into a single brightness distribution, which represents the degree of brightness response of different locations in the scene to incident light.

[0022] Read the color component data corresponding to each pixel position in the input image, extract the brightness contribution value of each color component at each pixel position, and sum the brightness contribution values ​​of each color component at the same pixel position in a weighted manner to form the single-channel brightness value of that pixel position. Rearrange the single-channel brightness values ​​of all pixel positions in the same spatial order as the input image to generate a single-channel grayscale image with the same number of rows and columns as the input image.

[0023] Morphological structural elements are defined based on the pixel resolution of the grayscale image.

[0024] Pixel resolution refers to the composition of pixels in the horizontal and vertical directions of an input image and the fineness of the spatial representation corresponding to a single pixel. This quantity is used to reflect the degree to which small structures in the image can be distinguished. Morphological structural elements refer to the basic operational units used to scan and compare local regions of an image in image morphological processing. They correspond to a local neighborhood range in the image plane and are used to characterize the spatial scale of action used when processing target boundaries, pore regions, and small occlusion structures.

[0025] The total number of pixels in the horizontal and vertical directions of the grayscale image is read. Based on the spatial coverage scale corresponding to a single pixel in the grayscale image, the neighborhood range for local scanning is determined. The center pixel position is selected in the pixel grid of the grayscale image. The neighborhood pixels are expanded in the horizontal and vertical directions around the center pixel position at the same pixel interval to form a local pixel template with a clear row and column distribution. The pixel positions in the local pixel template that are included in the neighborhood range are determined as effective positions, and the pixel positions that are not included in the neighborhood range are determined as ineffective positions. This forms the morphological structural elements used to perform dilation and erosion operations on the grayscale image.

[0026] In an embodiment of the present invention, determining the leaf slit aperture image includes: Dilation operations are performed on grayscale images and morphological structural elements to obtain dilated images; Read the brightness value of each pixel in the grayscale image, align the center of the morphological structuring element with each pixel in the grayscale image in turn, determine the grayscale image neighboring pixels covered by all effective positions in the morphological structuring element at each alignment, extract all brightness values ​​in the neighboring pixels, compare all extracted brightness values ​​one by one and determine the maximum brightness value, write the maximum brightness value to the output image pixel position corresponding to the center of the morphological structuring element, repeat the above neighbor coverage, brightness extraction, size comparison and pixel assignment process for all pixel positions in the grayscale image in turn, to form an inflated image with the same number of rows and columns as the grayscale image.

[0027] Erosion operations are performed on grayscale images and morphological structuring elements to obtain eroded images; Dilation refers to the process of selecting the larger value of pixel brightness within the coverage area of ​​the neighborhood template and assigning it to the center position. The result is a dilated image, which represents the spatial expansion of the original bright area. Erosion refers to the process of selecting the smaller value of pixel brightness within the coverage area of ​​the neighborhood template and assigning it to the center position. The result is an eroded image, which represents the spatial contraction of the original bright area.

[0028] Read the brightness value of each pixel in the grayscale image, align the center of the morphological structuring element with each pixel in the grayscale image in turn, and determine the grayscale image neighboring pixels covered by all effective positions of the morphological structuring element at each alignment. Extract all brightness values ​​of the neighboring pixels, compare all extracted brightness values ​​one by one and determine the minimum brightness value, and write the minimum brightness value to the output image pixel position corresponding to the center of the morphological structuring element. Repeat the above neighbor coverage, brightness extraction, size comparison and pixel assignment process for all pixel positions in the grayscale image in turn to form an erosion image with the same number of rows and columns as the grayscale image.

[0029] The dilation and erosion images are differentially processed to obtain the morphological gradient image; Read the pixel values ​​corresponding to the positions in the dilated and eroded images, establish a one-to-one pixel pairing relationship according to the same row and column coordinates, subtract the pixel value of the eroded image from the pixel value of the dilated image in each pair of corresponding pixels, and write the difference into the corresponding pixel position in the output image. Repeat the pixel reading, pixel pairing, value subtraction and result writing process for all corresponding pixel positions in sequence to form a morphological gradient image with the same number of rows and columns as the dilated and eroded images.

[0030] The morphological gradient image is normalized to obtain a normalized gradient image; Differential processing refers to the process of subtracting the values ​​at corresponding pixel positions in the dilated image and the eroded image, and the result is a morphological gradient image. The morphological gradient image is used to represent the distribution of the most obvious brightness changes in the image. Normalization processing refers to the process of mapping the pixel values ​​in the morphological gradient image to a uniform numerical range, and the result is a normalized gradient image.

[0031] Read all pixel values ​​in the morphological gradient image, determine the minimum and maximum values ​​among all pixel values, subtract the minimum value from each pixel value in the morphological gradient image, and divide the result of the subtraction by the difference between the maximum and minimum values ​​to obtain the normalized pixel value of the corresponding pixel position. Write all the normalized pixel values ​​into the output image in the original pixel position order to form a normalized gradient image with the same number of rows and columns as the morphological gradient image and the pixel values ​​are within a uniform numerical range.

[0032] Based on the normalized gradient image, the leaf gap aperture image is determined.

[0033] Leaf gap aperture image refers to image data obtained by distinguishing the boundary region from the gap region based on the distribution of occlusion boundaries and gaps reflected in the brightness changes in the normalized gradient image. It is used to represent the location distribution of the transparent area in the gap of vegetation.

[0034] Read the normalized pixel values ​​corresponding to each pixel position in the normalized gradient image. Perform a region continuity check on all pixels according to the adjacency relationship of the pixel positions in the image plane. Continuous regions with larger normalized pixel values ​​are identified as occlusion boundary regions. Continuous regions located between adjacent occlusion boundary regions and with smaller normalized pixel values ​​are identified as aperture candidate regions. Check each aperture candidate region one by one whether it is surrounded by occlusion boundary regions and whether it maintains a continuous distribution in the image plane. Aperture candidate regions that satisfy the surrounding relationship and continuous distribution relationship are retained as leaf gap aperture regions. Assign aperture label values ​​to the pixel positions in the normalized gradient image that belong to the leaf gap aperture region, and assign non-aperture label values ​​to the pixel positions that do not belong to the leaf gap aperture region. Generate a leaf gap aperture image with the same number of rows and columns as the normalized gradient image according to the label results of each pixel position.

[0035] S2. Perform average filtering on the dilated edge image corresponding to the input image to obtain the beaded edge response image; In an embodiment of the present invention, obtaining the beaded edge response image includes: Edge detection is performed on a grayscale image to obtain an edge response image; Edge detection refers to the process of determining the location of abrupt changes in brightness by calculating the degree of brightness change between adjacent pixels. The result is an edge response image, which represents the distribution of areas in the image where the brightness changes are more obvious.

[0036] Read the brightness value corresponding to each pixel position in the grayscale image. According to the brightness change relationship between adjacent pixels in the grayscale image, calculate the brightness change of each pixel position in the horizontal and vertical directions respectively. Combine the brightness changes of the same pixel position in the horizontal and vertical directions to obtain the edge response value corresponding to the pixel position. Write the edge response values ​​corresponding to all pixel positions into the output image in the original pixel position order to form an edge response image with the same number of rows and columns as the grayscale image.

[0037] Perform an erosion operation on the edge response image to obtain the eroded edge image; Erosion is a process that selects a smaller value of pixel brightness within a local neighborhood and assigns it to the center position. The result is the eroded edge image, which represents the spatial shrinkage of the original edge.

[0038] The edge response values ​​corresponding to each pixel position in the edge response image are read. The center position of the morphological structuring element is sequentially aligned with each pixel position in the edge response image. During each alignment, the neighboring pixels of the edge response image covered by all effective positions in the morphological structuring element are determined. All edge response values ​​in the neighboring pixels are extracted. All extracted edge response values ​​are compared one by one and the minimum edge response value is determined. The minimum edge response value is written to the pixel position of the output image corresponding to the center position of the morphological structuring element. The process of neighbor coverage, response value extraction, size comparison and pixel assignment is repeated sequentially for all pixel positions in the edge response image to form an eroded edge image with the same number of rows and columns as the edge response image.

[0039] Perform a dilation operation on the eroded edge image to obtain the dilated edge image; Dilation is a process that selects the larger value of pixel brightness within a local neighborhood and assigns it to the center position. The result is a dilated edge image, which represents the spatial expansion of the edge.

[0040] Specifically, the edge response values ​​corresponding to each pixel position in the eroded edge image are read. The center position of the morphological structuring element is sequentially aligned with each pixel position in the eroded edge image. During each alignment, the neighboring pixels of the eroded edge image covered by all effective positions in the morphological structuring element are determined. All edge response values ​​in these neighboring pixels are extracted. All extracted edge response values ​​are compared one by one, and the maximum edge response value is determined. The maximum edge response value is written to the output image pixel position corresponding to the center position of the morphological structuring element. The process of neighbor coverage, response value extraction, size comparison, and pixel assignment is repeated sequentially for all pixel positions in the eroded edge image to form an expanded edge image with the same number of rows and columns as the eroded edge image.

[0041] Local window averaging filtering is performed on the dilated edge image to obtain the beaded edge response image.

[0042] Local window averaging filtering refers to the process of summing and averaging the brightness values ​​of all pixels within a neighborhood centered on a certain pixel, resulting in a beaded edge response image.

[0043] A beaded edge response image refers to a discrete brightness response distribution image formed on the imaging plane by continuous edges in a scene after passing through gaps in the perspective in a complex occluded environment. Due to the obstruction of occlusion structures such as leaves and branches, the continuous edges are only partially transmitted to the imaging device through limited spatial gaps, thus appearing in the image as multiple local brightness enhancement regions that are separate from each other but arranged along the same spatial direction. These local brightness enhancement regions present a similar serial distribution pattern in space, which is used to reflect the intermittent appearance of the occluded edges in the image and the spatial distribution of this appearance.

[0044] Read the edge response values ​​corresponding to each pixel position in the dilated edge image. Determine the local window coverage area with any pixel position in the dilated edge image as the center. Extract the edge response values ​​corresponding to all pixel positions within the coverage area of ​​the local window. Summate all the extracted edge response values ​​and divide the summation result by the total number of pixels in the local window to obtain the average response value corresponding to the current center pixel position. Write the average response value into the pixel position corresponding to the current center pixel position in the output image. Repeat the window determination, response value extraction, summation, averaging, and pixel assignment process for all pixel positions in the dilated edge image to form a beaded edge response image with the same number of rows and columns as the dilated edge image.

[0045] S3. Based on the leaf gap aperture image and the beaded edge response image, the leaf gap beaded edge is defined, the aperture-beaded co-occurrence region is determined, and the beaded edge response image in the aperture-beaded co-occurrence region is constrained to obtain the aperture-constrained beaded evidence image. In embodiments of the present invention, determining the aperture-bead co-occurrence region includes: In the leaf gap aperture image and the beaded edge response image, identify the discrete short edge segments that simultaneously satisfy the aperture co-occurrence relationship; Specifically, the aperture identification results corresponding to each pixel position in the leaf aperture image are read, and the response values ​​corresponding to each pixel position in the beaded edge response image are read. The pixel correspondence between the two images is established according to the same row coordinates and column coordinates. The pixel positions that the aperture identification results indicate as aperture regions and whose corresponding response values ​​are not zero are determined as co-occurring pixel positions. Spatially adjacent and continuously distributed co-occurring pixel positions are connected to form multiple distinct edge regions. The continuous pixel positions contained in each edge region are counted one by one, and each edge region formed by connecting co-occurring pixel positions is determined as discrete short edge segments.

[0046] The edge feature in which discrete short side segments are arranged in a beaded pattern is defined as the leaf gap beaded edge; Discrete short edge segments refer to local edge fragments that are retained only after continuous edges are occluded in an image; beaded leaf gap edges refer to the edge representation formed by these discrete short edge segments arranged sequentially in the same direction in space; aperture co-occurrence relationship refers to the spatial correspondence where a certain pixel position is simultaneously located in a transparent gap area and there is an edge response.

[0047] Specifically, the positional distribution of each discrete short side segment in the image is read, the extension direction and endpoint position of each discrete short side segment are determined, the arrangement order, interval position and direction consistency between adjacent discrete short side segments are compared, multiple discrete short side segments that are distributed sequentially along the same extension direction and are spatially separated are merged into the same beaded edge, the beaded edge is identified as the leaf gap beaded edge, and all discrete short side segments that constitute the leaf gap beaded edge are marked in the image.

[0048] Based on the co-occurrence characteristics of apertures along the leaf gap beaded edge, the pixel positions of the leaf gap aperture image and the beaded edge response image are divided into aperture-dominant region, beaded-dominant region, and aperture-beaded co-occurrence region.

[0049] The aperture-dominant region refers to the region where the gap perspective area is the main manifestation; the bead-dominant region refers to the region where the discrete edge fragment response is the main manifestation; the aperture-bead co-occurrence region refers to the region where the gap perspective area and the discrete edge fragment appear simultaneously in space, used to characterize the positional distribution of the occluded edge that is revealed through the gap.

[0050] The aperture identification results corresponding to each pixel position in the leaf gap aperture image are read, and the response values ​​corresponding to each pixel position in the beaded edge response image are read. The pixel correspondence between the two images is established according to the same row and column coordinates. The location of the discrete short edge segment that has been identified as the leaf gap beaded edge is found in all pixel positions. The pixel position located at the leaf gap beaded edge position with the corresponding aperture identification result indicating an aperture region and the corresponding response value is not zero is determined as the aperture beaded co-occurrence region. The pixel position located at the leaf gap beaded edge position with only aperture region identification and no beaded edge response is determined as the aperture dominant region. The pixel position located at the leaf gap beaded edge position with only beaded edge response and no aperture region identification is determined as the beaded dominant region. The region division results corresponding to all pixel positions are written into the output image in the original pixel position order to represent the spatial distribution of the aperture dominant region, beaded dominant region, and aperture beaded co-occurrence region.

[0051] In an embodiment of the present invention, obtaining an aperture-constrained beaded evidence image includes: Within the aperture-bead co-occurrence region, the bead edge response image is labeled and extracted based on the bead morphology of the leaf gap bead edge to obtain bead point labels, bead string labels, and solitary bead labels. Label extraction processing refers to the process of distinguishing brightness responses with different spatial distribution states in an image and assigning them corresponding labels; bead labeling refers to labeling brightness responses that appear as a single or a small number of pixel clusters in a local area; bead string labeling refers to labeling areas where multiple brightness responses are arranged in the same direction to form a continuous distribution relationship; and solitary bead labeling refers to labeling brightness responses that exist independently without forming a continuous arrangement relationship with other brightness responses.

[0052] Read the bead edge response values ​​corresponding to each pixel position in the aperture bead co-occurrence area, establish the spatial adjacency relationship between adjacent pixels according to the row coordinates and column coordinates of the pixels in the image plane, merge the spatially connected pixel positions with continuously distributed response values ​​into the same response unit, and determine the pixel range, geometric center position, boundary range and extension direction covered by each response unit. Response units with a small coverage area and a concentrated distribution in a local area are identified as bead candidate units. Multiple bead candidate units are compared one by one according to the arrangement order of their geometric center positions. It is determined whether adjacent bead candidate units are distributed sequentially along the same extension direction and whether there is a front-to-back interval separation relationship. Multiple bead candidate units that satisfy the sequential distribution relationship and the interval separation relationship are merged into the same bead string region, and all pixel positions in the bead string region are marked as bead string markers. Bead candidate units that do not form a sequential distribution relationship with other bead candidate units and maintain an independent distribution state are marked as isolated bead markers. The pixel positions corresponding to each bead candidate unit belonging to the bead string region are marked as bead markers. Bead marker images, bead string marker images, and isolated bead marker images are generated according to the marking results corresponding to each pixel position.

[0053] Statistical processing was performed on the bead dot markings, bead string markings, and solitary bead markings to obtain the bead dot density, bead string length ratio, and solitary bead ratio. Statistical processing refers to the process of quantitatively analyzing the quantity distribution and occupancy ratio of the above-mentioned types of marks within a spatial range; bead density refers to the quantity distribution of bead marks within a unit area; bead string length ratio refers to the proportion of the spatial range covered by bead string marks within the overall mark range; and solitary bead ratio refers to the proportion of solitary bead marks among all marks.

[0054] Read the marking results corresponding to each pixel position in the bead mark image, bead string mark image and solitary bead mark image, establish the spatial correspondence of all marked pixel positions according to the row coordinates and column coordinates of the image, count the total number of pixels marked as bead marks in the co-occurrence area of ​​aperture bead, and count the total number of pixels covered by the co-occurrence area of ​​aperture bead, calculate the ratio of the total number of bead marks to the total number of pixels in the co-occurrence area of ​​aperture bead to obtain the bead density; For each separated bead string region in the bead string marking image, the positions of all marked pixels in each bead string region are extracted one by one to determine the start and end positions of each bead string region. The spatial distance between adjacent marked pixels is accumulated along the continuous arrangement direction of adjacent marked pixels in the bead string region to obtain the length of the corresponding bead string region. The lengths of all bead string regions are accumulated to obtain the total length of the bead string. The ratio of the total length of the bead string to the total length formed by all marked pixels in the aperture bead co-occurrence area along the corresponding arrangement direction is calculated to obtain the bead string length ratio. In the isolated bead marking image, the total number of pixels marked as isolated bead marks is counted, and the total number of marked pixels corresponding to bead dot marks, bead string marks, and isolated bead marks is counted. The ratio of the total number of isolated bead mark pixels to the total number of marked pixels is calculated to obtain the isolated bead ratio. The bead dot density, bead string length ratio, and isolated bead ratio are written into the statistical results data area as input data for subsequent constraint processing.

[0055] Based on the bead density, bead string length ratio, and solitary bead ratio, the bead edge response image in the aperture-bead co-occurrence region is constrained to obtain the aperture-constrained bead evidence image.

[0056] Constraint processing refers to the process of filtering and retaining the brightness response at the corresponding position in the image based on the above statistical results. Aperture-constrained beaded evidence image refers to the image result retained under the condition of simultaneously satisfying the aperture perspective relationship and the beaded edge response distribution, which is used to represent the effective area distribution of the occluded edge that is revealed through the gap.

[0057] The bead edge response values ​​corresponding to each pixel position within the co-occurrence area of ​​the aperture beaded string are read. The bead density, bead length ratio, and isolated bead ratio corresponding to each pixel position are also read. A correspondence between the statistical results and pixel positions is established according to the row and column coordinates in the image. For pixels belonging to bead string markers within the co-occurrence area, the original bead edge response values ​​are maintained. For pixels belonging to bead dot markers within the co-occurrence area, the bead edge response values ​​are adjusted based on the combination of bead density and bead length ratio, ensuring that bead dot marker pixels consistent with the continuous distribution of beads maintain a high response. For pixels belonging to isolated bead markers within the co-occurrence area, the bead edge response values ​​are suppressed based on the isolated bead ratio, weakening isolated responses that do not form a continuous bead arrangement. All pixel position response values ​​after preservation, adjustment, and suppression are rewritten to form an output image with the same number of rows and columns as the bead edge response image. This output image is then identified as the aperture constraint beaded string evidence image.

[0058] S4. Based on the aperture-constrained beaded evidence image, determine the historical deposition image, perform recursive deposition processing on the historical deposition image, and obtain the beaded trajectory deposition image at the current moment. In an embodiment of the present invention, determining historical deposition images includes: Use the aperture-constrained beaded evidence image as the evidence image at the current moment; Read the aperture constraint beaded evidence image corresponding to the current acquisition time, confirm the response value distribution of each pixel position in the aperture constraint beaded evidence image as well as the number of rows and columns of the image, write the aperture constraint beaded evidence image into the evidence data area corresponding to the current time, and determine the image data stored in the evidence data area as the evidence image at the current time, which is used to represent the response distribution of the beaded edges that satisfy the aperture constraint relationship at the current acquisition time.

[0059] Based on the image size of the aperture-constrained beaded evidence image, construct an all-zero image; The evidence image at the current moment refers to the image data used to describe the effective response information in the scene at the current acquisition moment; the image size refers to the distribution of the number of pixels in the horizontal and vertical directions; the all-zero image refers to the image data in which all pixel positions are assigned zero values ​​under the premise of having the same number of rows and columns as the target image, and is used to represent the initial state without any response.

[0060] Read the row and column information of the aperture constraint beaded evidence image, and build a pixel array of the same size as the aperture constraint beaded evidence image according to the read row and column information. Assign zero value to each pixel position in the pixel array one by one to form an all-zero image with the same row and column number as the aperture constraint beaded evidence image and all pixel values ​​are zero.

[0061] Obtain the deposition image of the bead trajectory from the previous time step; The bead trajectory deposition image of the previous moment refers to the image data recorded at a time point before the current acquisition time. This image data reflects the spatial distribution formed by the continuous appearance and gradual accumulation of discrete edge responses in the image plane over a previous period. It is used to represent the position superposition result formed by the multiple intermittent appearances of occluded edges in the time dimension.

[0062] Read the image storage area corresponding to the previous acquisition time adjacent to the current acquisition time, and search whether the bead trajectory deposition image has been stored at the previous acquisition time. If there is a stored bead trajectory deposition image, read the pixel value corresponding to each pixel position in the bead trajectory deposition image, as well as the row and column number information of the image, and output the read image data as the bead trajectory deposition image of the previous time, which is used to represent the spatial distribution of the bead edge response that has been accumulated at the time point before the current acquisition time.

[0063] The bead trajectory deposition image from the previous moment is recorded as the historical deposition image; when there is no bead trajectory deposition image from the previous moment, the image with all zeros is recorded as the historical deposition image.

[0064] Historical sedimentary images refer to image data used to carry and perpetuate the aforementioned time-accumulated results. They spatially correspond to the pixel positions of the same observation area and maintain the overall record of the distribution of past discrete edge responses, thus characterizing the temporal continuity of the distribution of beaded edges within the observation area.

[0065] When there is no bead trajectory deposition image from the previous moment, the all-zero image is recorded as the historical deposition image. This includes determining whether the bead trajectory deposition image from the previous moment has been successfully acquired. If it has been successfully acquired, the acquired bead trajectory deposition image from the previous moment is written into the historical deposition data area, and the image data in the historical deposition data area is identified as the historical deposition image. If it has not been successfully acquired, the zero values ​​corresponding to each pixel position in the all-zero image and the image size information are read, the all-zero image is written into the historical deposition data area, and the image data in the historical deposition data area is identified as the historical deposition image.

[0066] In an embodiment of the present invention, obtaining the bead trajectory deposition image at the current moment includes: Set the deposition recursion coefficient; The depositional recursion coefficient refers to a quantity used to describe the proportion of historical image information and current image information in the spatial superposition process, and is used to reflect the degree of contribution between the cumulative response over a past time and the response at the current moment.

[0067] Read the historical deposition values ​​corresponding to all pixel positions in the historical deposition image, and read the evidence response values ​​corresponding to all pixel positions in the evidence image at the current moment. Accumulate all historical deposition values ​​in the historical deposition image to obtain the total historical deposition value. Accumulate all evidence response values ​​in the evidence image at the current moment to obtain the total current evidence value. Add the total historical deposition value and the total current evidence value to obtain the total response value. Divide the total historical deposition value by the total response value to obtain the deposition recursive coefficient corresponding to the historical deposition image. Divide the total current evidence value by the total response value to obtain the current evidence weight value corresponding to the evidence image at the current moment. Write the deposition recursive coefficient and the current evidence weight value into the recursive calculation data area.

[0068] Based on the depositional recursion coefficient, recursive depositional processing is performed on historical depositional images and evidence images at the current moment to obtain the beaded trajectory depositional image at the current moment.

[0069] Recursive deposition processing refers to the process of superimposing historical deposition images with current evidence images in a certain proportion, pixel by pixel, to form a new cumulative result; the beaded trajectory deposition image at the current moment refers to the image data obtained after recursive deposition processing, which reflects the continuous superposition distribution of beaded edges in the time series.

[0070] Read the historical deposition values ​​corresponding to each pixel position in the historical deposition image, and read the evidence response values ​​corresponding to each pixel position in the evidence image at the current moment. Establish a one-to-one pixel correspondence between the historical deposition image and the evidence image at the current moment according to the same row coordinates and column coordinates. For each corresponding pixel position, multiply the historical deposition value by the deposition recursion coefficient, multiply the evidence response value by the current evidence weight value, and then sum the two product results to obtain the recursive deposition value corresponding to the pixel position. Write the recursive deposition values ​​corresponding to all pixel positions into the output image in the original pixel position order to form a beaded trajectory deposition image at the current moment with the same number of rows and columns as the historical deposition image and the evidence image at the current moment.

[0071] It should be noted that historical sedimentary images and current evidence images represent the cumulative and instantaneous response states of the same spatial location at different times, respectively. Essentially, they correspond to a temporal superposition of energy or intensity distribution within the same observation area. During continuous observation, the total historical sedimentary amount reflects the overall spatial accumulation of responses at previous times, while the total current evidence amount reflects the overall contribution of newly added responses at the current time. Summing the two yields the total response amount within the observation area. Calculating the ratio of the total historical sedimentary amount to the total response amount yields the proportion of the historical portion in the whole, and calculating the ratio of the total current evidence amount to the total response amount yields the proportion of the current portion. The proportion of the total value in the whole satisfies the proportional distribution principle under the total conservation, thus reflecting the relative contribution of historical accumulation and current input to the whole. At each pixel position, the historical deposition value and the current evidence response value are weighted and superimposed according to the above proportion. This is equivalent to continuously integrating and updating the response at the same spatial position in the time series, so that the output image retains both historical accumulation information and current response information, and forms a spatial distribution result that evolves continuously over time. The resulting image can stably reflect the superposition trajectory of the beaded edge in the time dimension. Therefore, the result of this calculation process is the deposition recursion coefficient and the beaded trajectory deposition image at the current moment.

[0072] It should be noted that under the condition of discrete edge manifestation in leaf gaps, the target edge is not continuously and stably displayed in the image. Instead, due to the occlusion and perspective effect of the leaf gaps, it appears intermittently in the form of discrete short edge segments in different time frames. These discrete short edge segments are difficult to fully reflect the real edge structure at a single moment and are easily confused with other sporadic responses in the background, resulting in unstable detection results and false connections during tracking. By constructing a beaded trajectory deposition image at the current moment, discrete responses that repeatedly appear in spatial locations in multiple time frames can be gradually accumulated, so that the responses generated by the same target edge form a continuously enhanced distribution in space. Meanwhile, random background responses, due to the lack of temporal repetition, are difficult to form a stable accumulation, thereby strengthening the existence of the real edge in the temporal dimension and suppressing false responses, so that the subsequent target generation and tracking processes can be judged based on more stable spatial distribution results.

[0073] S5. Extract connected component attributes from the current bead trajectory deposition image to obtain target tracking information, and generate the UAV target tracking result based on the target tracking information.

[0074] In embodiments of the present invention, target tracking information is obtained, including: Connected components are extracted from the current moment's bead trajectory deposition image to obtain connected regions; Connected component extraction refers to the process of merging spatially connected and continuously distributed response pixels in an image, the result of which is a connected region. A connected region is a continuous region composed of adjacent response pixels, used to represent candidate target regions with complete spatial distribution in an image.

[0075] Read the deposition response values ​​corresponding to each pixel position in the current bead trajectory deposition image. Establish the spatial adjacency relationship between pixels according to the row and column coordinates of each pixel position in the image. Merge the spatially adjacent pixel positions with continuous deposition response one by one. Check the spatial connectivity of each merged pixel set. Determine the pixel set that can be continuously reached through adjacent pixel paths as the same connected region. Record the pixel position range, boundary range and region center position of each connected region. Write all connected regions into the region data region according to the original image coordinate relationship to obtain the connected regions at the current time.

[0076] Target tracking information is obtained based on connected regions.

[0077] Target tracking information refers to the spatial location data of a target determined based on a connected region, used to indicate the location of a target that the UAV needs to continuously monitor.

[0078] The system reads the pixel location range and region center location covered by each connected region, determines the center coordinate position and outer boundary range of each connected region in the image plane, uses the center coordinate position as the target center location information, and uses the outer boundary range as the target location range information. The target center location information and target location range information corresponding to all connected regions are combined to form target tracking information that corresponds one-to-one with each connected region at the current time, and writes the target tracking information into the tracking result data area to represent the target position output of the UAV at the current time.

[0079] In an embodiment of the present invention, generating the target tracking result of the UAV includes: Extract the non-zero response from the evidence image at the current moment to obtain the support domain; Non-zero response extraction refers to the process of extracting the location distribution of non-zero pixel values ​​from the evidence image, and the result is the support domain. The support domain refers to the spatial region in the evidence image where the response actually exists, and is used to represent the real and valid edge display position at the current moment.

[0080] Read the response values ​​corresponding to each pixel position in the evidence image at the current moment, traverse all pixel positions according to the row and column coordinates of each pixel position in the image, determine the pixel positions with response values ​​that are not equal to zero as valid response pixel positions, collect all valid response pixel positions according to the original image coordinate relationship, and determine all collected valid response pixel positions as the support domain, which is used to represent the spatial region in the evidence image at the current moment where the response actually exists.

[0081] Determine the overlap area between the supporting domain and the connected region; The overlap area refers to the area covered by the support domain and the connected region in the image plane, and is used to represent the degree of overlap between the historical cumulative position and the current effective response position.

[0082] Read the coordinate information of each effective response pixel position in the supporting domain, read the pixel position coordinate information covered by each region in the connected region, establish the spatial correspondence between the supporting domain and the connected region according to the same row coordinates and column coordinates, determine whether each effective response pixel position in the supporting domain is simultaneously located within the pixel position range covered by the connected region, determine the pixel positions simultaneously located in the supporting domain and the connected region as overlapping pixel positions, count the number of all overlapping pixel positions, and determine the total number of overlapping pixels as the overlapping area between the supporting domain and the connected region.

[0083] When the overlap area is greater than 0, the target tracking result of the UAV is generated using the target tracking information; When the overlap area is equal to 0, the output of target tracking information is terminated.

[0084] The target tracking result refers to the target position result currently output by the UAV based on the target tracking information, which is used to indicate the target object that the UAV should continue to track at the current moment.

[0085] Read the numerical result corresponding to the overlapping area and determine whether the numerical result is greater than zero. If the overlapping area is greater than zero, read the target tracking information corresponding to the current connected region, write the target center position information and target position range information in the target tracking information into the target result data area at the current time, and determine the content in the target result data area as the target tracking result of the UAV. If the overlapping area is equal to zero, stop writing the target tracking information corresponding to the current connected region into the target result data area, and clear the target tracking information that is not supported by the supported domain at the current time, so that the UAV does not output the corresponding target tracking result at the current time.

[0086] It should be noted that the overlap area reflects the spatial consistency between the connected regions formed by historical accumulation and the actual effective response region at the current moment. When the overlap area is greater than zero, it means that the candidate region formed in historical deposition still has corresponding actual response support at the current moment, indicating that the region has a continuous spatial basis and can be considered as a continuation of the same target in the time series. Therefore, the target tracking result can be output based on the target tracking information. However, when the overlap area is equal to zero, it means that the candidate region formed in historical deposition lacks any actual response support at the current moment. It means that the region is only a residual distribution caused by historical accumulation or a false region formed by the superposition of intermittent manifestations. In this case, continuing to output the target tracking information will lead to misjudgment. Therefore, by terminating the output of the target tracking information, it is possible to avoid mistaking the region that does not have current response basis as the real target, thereby improving the reliability of the tracking result.

[0087] This invention achieves visual target detection by performing pixel-by-pixel processing and spatial relationship analysis on images acquired by UAVs. Specifically, it extracts brightness variation regions from the image through grayscale transformation and morphological operations to form a leaf aperture image to represent the transparent area. Simultaneously, it obtains a beaded edge response image through edge detection and filtering to represent discontinuous edge information. Based on this, it combines the spatial correspondence between the two types of images to filter effective response regions and obtains candidate target regions with spatial continuous distribution through connected component extraction. Finally, it confirms and outputs the target based on the spatial overlap between the effective response regions in the current image and the historical accumulated regions, thereby completing the visual detection and tracking process based on image brightness distribution and spatial structure information.

[0088] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for precise detection and tracking of low-altitude targets by unmanned aerial vehicles based on machine vision, characterized in that, The steps include: S1, setting morphological structural elements based on the input image collected by the UAV, and determining the leaf gap aperture image based on the morphological structural elements; S2. Perform average filtering on the dilated edge image corresponding to the input image to obtain the beaded edge response image; S3. Based on the leaf gap aperture image and the beaded edge response image, the leaf gap beaded edge is defined, the aperture-beaded co-occurrence region is determined, and the beaded edge response image in the aperture-beaded co-occurrence region is constrained to obtain the aperture-constrained beaded evidence image. The specific steps to obtain the aperture-constrained bead evidence image are as follows: Within the aperture-bead co-occurrence region, the bead edge response image is labeled and extracted based on the bead morphology of the leaf gap bead edge to obtain bead point labels, bead string labels, and solitary bead labels. Statistical processing was performed on the bead dot markings, bead string markings, and solitary bead markings to obtain the bead dot density, bead string length ratio, and solitary bead ratio. Based on the bead density, bead string length ratio, and solitary bead ratio, the bead edge response image in the aperture-bead co-occurrence region is constrained to obtain the aperture-constrained bead evidence image. S4. Based on the aperture-constrained beaded evidence image, determine the historical deposition image, perform recursive deposition processing on the historical deposition image, and obtain the beaded trajectory deposition image at the current moment. S5. Extract connected component attributes from the current bead trajectory deposition image to obtain target tracking information, and generate the UAV target tracking result based on the target tracking information. 2.The method of claim 1, wherein, Define morphological structural elements, including: Acquire input images captured by the drone; Convert the input image to a grayscale image; Morphological structural elements are defined based on the pixel resolution of the grayscale image. 3.The method of claim 2, wherein, Determine the leaf slit aperture image, including: Dilation operations are performed on grayscale images and morphological structural elements to obtain dilated images; Erosion operations are performed on grayscale images and morphological structuring elements to obtain eroded images; The dilation and erosion images are differentially processed to obtain the morphological gradient image; The morphological gradient image is normalized to obtain a normalized gradient image; Based on the normalized gradient image, the leaf gap aperture image is determined.

4. The machine vision-based method for accurate detection and tracking of low-altitude targets by unmanned aerial vehicles according to claim 3, characterized in that, The beaded edge response image is obtained, including: Edge detection is performed on a grayscale image to obtain an edge response image; Perform an erosion operation on the edge response image to obtain the eroded edge image; Perform a dilation operation on the eroded edge image to obtain the dilated edge image; Local window averaging filtering is performed on the dilated edge image to obtain the beaded edge response image.

5. The method for accurate detection and tracking of low-altitude targets by unmanned aerial vehicles based on machine vision according to claim 1, characterized in that, Determine the aperture-bead co-occurrence region, including: In the leaf gap aperture image and the beaded edge response image, identify the discrete short edge segments that simultaneously satisfy the aperture co-occurrence relationship; The edge feature in which discrete short side segments are arranged in a beaded pattern is defined as the leaf gap beaded edge; Based on the co-occurrence characteristics of apertures along the leaf gap beaded edge, the pixel positions of the leaf gap aperture image and the beaded edge response image are divided into aperture-dominant region, beaded-dominant region, and aperture-beaded co-occurrence region.

6. The method for accurate detection and tracking of low-altitude targets by unmanned aerial vehicles based on machine vision according to claim 1, characterized in that, Identify historical sedimentary images, including: Use the aperture-constrained beaded evidence image as the evidence image at the current moment; Based on the image size of the aperture-constrained beaded evidence image, construct an all-zero image; Obtain the deposition image of the bead trajectory from the previous time step; The bead trajectory deposition image from the previous moment is recorded as the historical deposition image; when there is no bead trajectory deposition image from the previous moment, the image with all zeros is recorded as the historical deposition image.

7. The machine vision-based method for accurate detection and tracking of low-altitude targets by unmanned aerial vehicles according to claim 6, characterized in that, Obtain the current moment's bead trajectory deposition image, including: Set the deposition recursion coefficient; Based on the depositional recursion coefficient, recursive depositional processing is performed on historical depositional images and evidence images at the current moment to obtain the beaded trajectory depositional image at the current moment.

8. The machine vision-based method for accurate detection and tracking of low-altitude targets by unmanned aerial vehicles according to claim 7, characterized in that, Obtain target tracking information, including: Connected components are extracted from the current moment's bead trajectory deposition image to obtain connected regions; Target tracking information is obtained based on connected regions.

9. The machine vision-based method for accurate detection and tracking of low-altitude targets by unmanned aerial vehicles according to claim 8, characterized in that, Generate target tracking results for the drone, including: Extract the non-zero response from the evidence image at the current moment to obtain the support domain; Determine the overlap area between the supporting domain and the connected region; When the overlap area is greater than 0, the target tracking result of the UAV is generated using the target tracking information; When the overlap area is equal to 0, the output of target tracking information is terminated.

Citation Information

Patent Citations

  • Crop row detection method based minimum tangent circle and morphological principle

    CN105021196A

  • Low-altitude unmanned aerial vehicle multi-target identification tracking method based on deep learning

    CN122200011A