Intelligent identification method and system for UAVs targeting occluded targets

By segmenting the occlusion area and extracting dynamic feature of the multi-frame continuous monitoring images acquired by the drone, the problem of low accuracy of occlusion target recognition in traditional methods is solved, and the accurate identification and robustness of occlusion targets are achieved.

CN120259926BActive Publication Date: 2025-08-08DEYANG JINGKAI ZHIHANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510732832.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-08-08
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

When facing the occluded target, traditional drone target recognition methods are difficult to make full use of the temporal correlation and dynamic information between multiple frames of continuous monitoring images, resulting in low recognition accuracy and poor robustness, which cannot meet the precise identification requirements for occluded targets in practical applications.

Method used

Through the acquisition module equipped by the drone, the multi-frame continuous monitoring image of the target area is obtained in real time, the occlusion area segmentation process is performed, the segmentation results of the target and occlusion area are obtained, dynamic feature extraction is performed, and dynamic recognition results are generated. Based on this, the acquisition parameters are optimized and the identification strategy is adjusted adaptively.

Benefits of technology

Accurate identification of occlusion targets is achieved, the accuracy and robustness of identification are improved, the acquisition parameters can be adjusted adaptively, and the real-timeness of identification is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259926B_ABST
    Figure CN120259926B_ABST
Patent Text Reader

Abstract

The present invention provides a drone intelligent identification method and system for occluded targets. The method comprises the following steps: a collection module carried by the drone obtains, in real time, multiple frames of continuous monitoring images of a target area containing at least a partially occluded target to be identified; an occlusion region segmentation process is performed on the multiple frames of continuous monitoring images to obtain target region segmentation results and occlusion region segmentation results; a dynamic feature extraction process is performed based on the target region segmentation results and the occlusion region segmentation results to obtain the visible feature distribution and occlusion feature distribution of the target to be identified in each frame of the image; a dynamic identification result characterizing the state migration trajectory of the target to be identified is generated based on the visible feature distribution and the occlusion feature distribution; finally, based on the dynamic identification result, the acquisition parameters of the acquisition module are optimized in real time to generate a parameter adjustment strategy for the next monitoring cycle. In this way, the multi-frame image information can be fully utilized to accurately identify the occluded target, thereby improving the recognition accuracy and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of drone technology, and in particular to a drone intelligent recognition method and system for occluded targets. Background Art

[0002] In the field of drone applications, monitoring and identifying target areas is a key task, especially when identifying occluded targets, which faces many challenges. In actual scenarios, the target to be identified may be partially or completely occluded due to various reasons, such as buildings, trees, and other objects blocking the target, or the target's own posture changes causing part of the area to be invisible. Traditional target recognition methods are often based on single-frame image analysis. When faced with occluded targets, it is difficult to accurately identify the target due to the limited effective information available. This method usually cannot fully utilize the temporal correlation and dynamic information between multiple frames of continuous monitoring images, and cannot effectively distinguish between the visible and occluded parts of the target, resulting in low recognition accuracy and poor robustness, making it difficult to meet the needs of accurate recognition of occluded targets in practical applications. Summary of the Invention

[0003] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a method for intelligently identifying a drone with an obstructed target, the method comprising:

[0004] Acquire multiple frames of continuous monitoring images of the target area in real time through an acquisition module carried by the drone, wherein the multiple frames of continuous monitoring images include at least a portion of the target to be identified that is obscured;

[0005] Performing occlusion region segmentation processing on the multiple frames of continuous monitoring images to obtain a target region segmentation result of the target to be identified in each frame of monitoring image and a corresponding occlusion region segmentation result;

[0006] Performing dynamic feature extraction based on the target region segmentation result and the occlusion region segmentation result to obtain the visible feature distribution and occlusion feature distribution of the target to be identified in each frame of the monitoring image;

[0007] Generating a dynamic recognition result of the target to be identified according to the visible feature distribution and the occlusion feature distribution, wherein the dynamic recognition result is used to characterize a state transition trajectory of the target to be identified in multiple frames of continuous monitoring images;

[0008] Based on the dynamic recognition result, the acquisition parameters of the acquisition module are optimized in real time to generate a parameter adjustment strategy for the next monitoring cycle.

[0009] On the other hand, an embodiment of the present invention also provides a drone intelligent identification system for obscured targets, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0010] Based on the above aspects, the embodiment of the present invention obtains multiple frames of continuous monitoring images of the target area in real time through the acquisition module carried by the drone, and for the at least partially obscured target to be identified contained therein, first performs occlusion area segmentation processing on the multiple frames of continuous monitoring images, accurately obtains the target area segmentation result and the corresponding occlusion area segmentation result of the target to be identified in each frame of the monitoring image, and then performs dynamic feature extraction processing based on the segmentation result to obtain the visible feature distribution and occlusion feature distribution of the target to be identified in each frame of the monitoring image, comprehensively and meticulously characterizes the feature information of the target in different states, and generates dynamic recognition results that characterize the state transition trajectory of the target to be identified in the multiple frames of continuous monitoring images based on the visible feature distribution and the occlusion feature distribution, thereby achieving accurate grasp of the dynamic changes of the occluded target, and finally, based on the dynamic recognition result, performs real-time optimization processing on the acquisition parameters of the acquisition module, generates a parameter adjustment strategy for the next monitoring cycle, and can adaptively adjust the acquisition parameters, thereby improving the accuracy, robustness and real-time performance of identifying the occluded target. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 The figure is a schematic diagram of the execution flow of the method for intelligently identifying a UAV with an obscured target provided by an embodiment of the present invention.

[0012] Figure 2 Schematic diagram of exemplary hardware and software components of a drone intelligent identification system for occluded targets provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0013] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 FIG1 is a flow chart of a method for intelligently identifying a drone with an obscured target provided by an embodiment of the present invention. The method for intelligently identifying a drone with an obscured target is described in detail below.

[0014] Step S110: acquiring multiple frames of continuous monitoring images of the target area in real time through a collection module carried by the UAV, wherein the multiple frames of continuous monitoring images include at least a partially obscured target to be identified.

[0015] In this embodiment, the scenario is a large industrial park. The drone is equipped with an image acquisition module capable of collecting image data at a specific frequency. The acquisition frequency of the image acquisition module is set to f. Within the industrial park, there are many large pieces of equipment, some of which may be obscured by surrounding buildings, transport vehicles, etc. These obscured pieces of equipment are the targets to be identified. The acquisition module continuously captures images of a specific area of the industrial park at a frequency of f, generating an image each time. These images are sequentially sequenced to form multiple frames of continuous monitoring images, many of which contain partially obscured pieces of equipment to be identified. For example, at a certain moment, the acquisition module captures an image in which a piece of large processing equipment is partially obscured by a nearby transport vehicle.

[0016] Step S120: performing occlusion region segmentation processing on the multiple frames of continuous monitoring images to obtain a target region segmentation result of the target to be identified in each frame of monitoring image and a corresponding occlusion region segmentation result.

[0017] After acquiring multiple frames of continuous monitoring images, in order to clearly define the target to be identified and its obscured part, it is necessary to perform occluded area segmentation processing.

[0018] Step S121: performing time alignment processing on the multiple frames of continuous monitoring images to generate multiple frames of aligned monitoring images with a continuous time dimension.

[0019] In an industrial park environment, the drone will be affected by factors such as airflow and its own flight attitude adjustment during flight, resulting in the multi-frame continuous monitoring images collected not being strictly continuous in the time dimension. Suppose the multi-frame continuous monitoring images are I1, I2, I3...I n , the corresponding acquisition times are t1, t2, t3...t n To ensure the accuracy of subsequent processing, these images need to be time-series aligned.

[0020] Specifically, areas with significant and stable features in the image are selected as reference points. For example, the corner points of large landmark buildings at fixed locations in industrial parks. With the help of a feature matching algorithm, the algorithm calculates the deviation of each frame image relative to the standard time series by comparing the position changes, angle changes and other information of the reference points in different frame images. For image I1, the algorithm calculates its deviation from the standard time series as d1. The calculation of the deviation takes into account the coordinate changes of the reference point in the image. For image I2, the deviation d2 is calculated, and so on, to obtain d3...d n Then, according to these deviations, each frame image is subjected to corresponding geometric transformations such as translation and rotation, so that the adjusted images I1', I2', I3'...I n'It forms continuous and precisely aligned multi-frame aligned monitoring images in the time dimension, as if these images are closely arranged in accurate time sequence, laying a good foundation for subsequent processing.

[0021] Step S122: calling a pre-trained segmentation network model to perform frame-by-frame semantic segmentation processing on the multiple aligned monitoring images to obtain an initial segmentation mask in each frame of the monitoring image.

[0022] After acquiring multiple aligned monitoring images in a continuous temporal dimension, they are processed using a pre-trained segmentation network model. Taking the common Mask-RCNN segmentation network model as an example, this model mainly consists of a backbone network, a region proposal network (RPN), a region of interest (ROI) alignment layer, and a mask prediction branch.

[0023] In an industrial park scenario, one frame from a multi-frame aligned surveillance image, such as I1', is input into the backbone network of the Mask-RCNN model. This backbone network typically consists of multiple convolutional and pooling layers, responsible for extracting features from the input image. The convolutional layers perform convolution operations on the image using different convolution kernels, K1, K2, K3, etc., sliding across the image with a set step size. Each kernel extracts specific features, such as edge features or texture features. After multiple convolution and pooling operations, feature maps F1, F2, F3, etc., with different characteristics, are generated.

[0024] These feature maps then enter the Region Proposal Network (RPN). The RPN generates a series of anchor boxes on the feature maps using a sliding window. Each anchor box corresponds to a potential target region in the image. Based on the feature map information, the RPN classifies each anchor box (determines whether it is an target) and regresses it (predicts the anchor box's position and resize amount), generating a series of candidate regions that may contain the target.

[0025] The candidate regions are then processed by the Region of Interest (ROI) Align layer. The ROIAlign layer maps candidate regions of different sizes to a fixed-size feature map for subsequent unified feature extraction and prediction.

[0026] Finally, the features processed by the ROIAlign layer enter the mask prediction branch. This branch performs mask prediction on each candidate region through a series of convolutional and fully connected layers, outputting mask information corresponding to the candidate region. This mask is processed to generate an initial segmentation mask of the same size as the input image. This mask preliminarily demarcates the approximate regions of different objects in the image. In an industrial park scenario, it can initially distinguish the regions where different objects, such as equipment, buildings, and vehicles, reside.

[0027] Step S123: extracting the contour area of the target to be identified as a candidate target area based on the initial segmentation mask, and performing edge integrity verification processing on the candidate target area to obtain a verification result.

[0028] The contour region of the target to be identified is extracted from the initial segmentation mask of each monitored image frame. In an industrial park scenario, assume that the region corresponding to a large piece of equipment in the initial segmentation mask is M1. A contour tracking algorithm is employed. Starting from a certain edge pixel in the mask image, the algorithm checks adjacent pixels in a predetermined direction (e.g., clockwise or counterclockwise) to see if they belong to the contour. If adjacent pixels meet the predetermined contour conditions (e.g., the grayscale value difference exceeds a predetermined threshold T1), they are added to the contour pixel set. Through continuous tracking, the contour region of the target to be identified (the large piece of equipment) is formed as the candidate target region C1.

[0029] Next, edge integrity verification is performed on the candidate target area C1.

[0030] Step S1231: extracting the edge pixel point set of the candidate target area, and calculating the curvature distribution characteristics of the edge pixel point set.

[0031] Extract the edge pixel point set E1 of the candidate target area C1. For each edge pixel point P(x0, y0) in the set E1, select its two adjacent pixel points P1(x1, y1) and P2(x2, y2). Calculate the curvature in the following way: Based on the coordinates of these three points, use vector operations and trigonometric functions to calculate the angle θ between the vectors P1P0 and P0P2, and then use the distance formula between the two points to calculate the lengths L1 and L2 of P1P0 and P0P2. Calculate the curvature value K(x0, y0) of the point using the conventional curvature calculation formula in the relevant technology (using θ, L1 and L2), and perform this operation on all pixel points in the set E1 to obtain the curvature distribution characteristics of the edge pixel point set E1.

[0032] Step S1232: Detecting whether there is a curvature mutation point in the candidate target area based on the curvature distribution characteristics, where the curvature mutation point is used to characterize the edge defect position of the candidate target area.

[0033] A curvature mutation threshold T2 is set. When the curvature value K(x0, y0) of a certain pixel point differs from the curvature value of the adjacent pixel point by more than T2, the pixel point is identified as a curvature mutation point. These curvature mutation points are used to characterize the edge defect position of the candidate target area C1.

[0034] Step S1233: When a curvature mutation point is detected, the local edge segment corresponding to the curvature mutation point is extracted, and the motion-compensated similarity score between the local edge segment and the edge segment at the corresponding position in the adjacent frame monitoring image is calculated in combination with the target motion estimation result.

[0035] When curvature mutation points are detected, the local edge segments L1 corresponding to these curvature mutation points are extracted. Combined with the target motion estimation results, it is assumed that the target motion estimation obtains the target's motion vector V by analyzing the displacement of the same feature points in adjacent frame images. Based on the motion vector V, the edge segment L2 at the corresponding position in the adjacent frame monitoring image is found. The motion-compensated similarity score S is calculated between the local edge segment L1 and the edge segment L2 at the corresponding position in the adjacent frame monitoring image. The calculation method is as follows: L1 and L2 are sampled at equal intervals to obtain sampling point sets S1 and S2. For each sampling point in S1, the nearest point is found in S2. The sum of the distances D between all corresponding points is calculated. Then, D is subtracted from a fixed length value L (such as the length of L1 or L2). The ratio of the obtained difference to L is the motion-compensated similarity score S.

[0036] Step S1234: determining the degree of edge defect of the candidate target region according to the motion compensation similarity score; if the degree of edge defect exceeds a preset threshold, determining that the candidate target region requires cross-frame contour compensation processing.

[0037] An edge defect threshold T3 is set. If S is less than T3, it is determined that the edge defect degree of the candidate target area exceeds the preset threshold, that is, it is determined that the candidate target area needs to be subjected to cross-frame contour compensation processing.

[0038] Step S124: performing cross-frame contour compensation processing on the candidate target region with defects according to the verification result to obtain a compensated target region segmentation result.

[0039] If the verification result of step S123 indicates that the candidate target area has edge defects and needs to be compensated, in the industrial park scenario, assume that the candidate target area C1 (such as the outline area of large equipment) has edge defects. Compensation is performed using information from the corresponding areas in the adjacent frame monitoring images.

[0040] Assume that the candidate target area C1 with defects in the current frame image is I curr , the adjacent frame image is I prev and I nxt First, the motion vector V obtained based on the target motion estimation result is iprev and inxt Find the approximate area C corresponding to C1 1prev and C 1nxtFor each pixel point P(x, y) at the edge defect in C1, according to the motion vector V, 1prev and C 1nxt Find the corresponding pixel point P prev (x prev ,y prev ) and P nxt (x nxt ,y nxt ).

[0041] Use interpolation algorithm to perform contour compensation. For example, for the gray value of point P, by calculating P prev and P nxt The grayscale values of point P and their relative position relationship are used to determine the compensation grayscale value of point P using the weighted average method. Assume that P prev The gray value is G prev , P nxt The gray value is G nxt , P and P prev 、P nxt The distances are d1 and d2 respectively, and the compensation gray value of point P is G=(G prev *d2+G nxt *d1) / (d1+d2). Similar operations are performed on all pixel points at the edge defects in C1 to obtain the compensated target area segmentation result C1'.

[0042] Step S125: determining the occluded region segmentation result based on the difference area between the compensated target region segmentation result and the initial segmentation mask.

[0043] In the industrial park scenario, the compensated target area segmentation result C1' is compared with the initial segmentation mask M1.

[0044] Step S1251: performing pixel-by-pixel comparison on the compensated target region segmentation result and the initial segmentation mask to obtain a set of difference pixels.

[0045] Through pixel-by-pixel comparison, for each pixel P(x, y), the grayscale value or category information (if the mask is based on category annotation) of C1'(x, y) and M1(x, y) are compared. If the two are different, the pixel is marked as a difference pixel, and all difference pixels constitute the difference pixel set D1.

[0046] Step S1252: performing spatial clustering processing on the difference pixel point set to generate multiple candidate occlusion regions.

[0047] Perform spatial clustering on the difference pixel set D1. The DBSCAN clustering algorithm is used, which assumes two parameters: the neighborhood radius ε and the minimum number of points, MinPts. For each pixel P(x, y) in set D1, calculate the distance (e.g., Euclidean distance) to all other pixels in its neighborhood. If the number of pixels within the radius ε is greater than or equal to MinPts, these pixels are grouped into a cluster. By processing all pixels, multiple candidate occlusion regions, O1, O2, O3, and so on, are generated.

[0048] Step S1253: performing region screening processing based on the spatiotemporal continuity characteristics of the candidate occlusion region in multiple frames of continuous monitoring images to obtain a stable occlusion region segmentation result.

[0049] For each candidate occlusion region, such as O1, its position and shape changes are tracked in multiple frames of continuous monitoring images. Its position change Δx, Δy and shape change (such as area change, aspect ratio change, etc.) in different frames are calculated. If the position and shape changes of the candidate occlusion region in multiple frames are within the set range, that is, the position change Δx, Δy is less than the set position change threshold T x 、T y , the shape change is less than the set shape change threshold T s , then the candidate occlusion region is considered to have spatiotemporal continuity and is retained as the stable occlusion region segmentation result O1'.

[0050] Step S1254: performing a single iterative optimization on the contour information of the target area segmentation result according to the stable occlusion area segmentation result. If the overlapping area of the optimized target area and the occlusion area is less than a preset threshold, the iteration is stopped to generate an optimized target area segmentation result.

[0051] Calculate the overlapping area A1 between C1' and O1'. Set an overlapping area threshold T4. If A1 is greater than T4, adjust the outline of C1'. The adjustment method is as follows: for the pixels on the C1' outline that overlap with O1', reposition these pixels based on the edge information of O1' and the feature information inside C1' to reduce the overlapping area. After adjustment, calculate the overlapping area again. If the overlapping area is less than T4, stop the iteration and generate the optimized target area segmentation result C1''.

[0052] Step S130: performing dynamic feature extraction processing based on the target region segmentation result and the occlusion region segmentation result to obtain the visible feature distribution and occlusion feature distribution of the target to be identified in each frame of the monitoring image.

[0053] After obtaining the target region segmentation result C1'' and the occlusion region segmentation result O1', perform dynamic feature extraction processing on them to obtain the visible feature distribution and occlusion feature distribution of the target to be recognized in each frame of the monitoring image.

[0054] Step S131: Divide multiple analysis windows with different scales in the target region segmentation result, and calculate the gradient magnitude distribution within each analysis window.

[0055] In the target region segmentation result C1'', divide multiple analysis windows according to different scale requirements. For example, set the small-scale window size to s1×s1, the medium-scale window size to s2×s2, and the large-scale window size to s3×s3 (where s1 < s2 < s3). For each scale of window, slide the window on C1'' with a set step size. For each pixel point within the window, use a gradient-based algorithm to calculate the gradient magnitude. For pixel point P(x, y), calculate the gray-level change in its horizontal and vertical directions, and use the gradient calculation formula to obtain the gradient magnitude G(x, y) of this point. After calculating the gradient magnitudes for all pixel points within the window, obtain the gradient magnitude distribution within each analysis window.

[0056] Step S132: Perform histogram of oriented gradients (HOG) statistical processing on the gradient magnitude distribution to obtain the texture orientation histogram feature of each analysis window.

[0057] For the gradient magnitude distribution within each analysis window, perform histogram of oriented gradients (HOG) statistics. Divide the gradient directions into several intervals, for example, divide them into n intervals, and each interval corresponds to a direction range. For the gradient direction of each pixel point within the window, according to the direction range it belongs to, accumulate the gradient magnitude of this pixel point into the corresponding interval. After the statistics are completed, obtain the texture orientation histogram feature of each analysis window, which reflects the direction distribution of the texture within the window.

[0058] Step S133: Normalize and splice the texture orientation histogram features of different scales to generate the local texture feature of the target to be recognized.

[0059] Normalize the texture orientation histogram features obtained from windows of different scales. For the texture orientation histogram feature of each scale, calculate the sum Sum of all interval values, and then divide each interval value by Sum to obtain the normalized texture orientation histogram feature. Splice the normalized texture orientation histogram features of different scales together in scale order to generate the local texture feature of the target to be recognized, which integrates the texture information at different scales.

[0060] Step S134: extracting the minimum bounding rectangle of the target area segmentation result, performing binning encoding processing on the aspect ratio and area change rate of the minimum bounding rectangle, and generating the global morphological features of the target to be identified.

[0061] Extract the minimum bounding rectangle of the target area segmentation result C1''. Assume that the length of the minimum bounding rectangle is l and the width is w, and calculate its aspect ratio r=l / w. At the same time, assume that the area of the minimum bounding rectangle corresponding to the target area in the previous frame image is A prev , the area of the current frame is A curr , calculate the area change rate v=(A curr -A prev ) / A prev .

[0062] Perform binning and encoding on the aspect ratio r. Set several aspect ratio intervals, for example, r1 <r2<…<r n , encoding r into the corresponding category value based on the interval it falls into. Similarly, the area change rate v is binned and encoded. Several area change rate intervals are set and encoded into the corresponding category value based on the interval v falls into. These two encoded category values are combined to generate the global morphological feature of the target to be identified. This feature reflects the overall shape and area change of the target area.

[0063] Step S135: performing region growing processing on the occlusion region segmentation result to generate an occlusion region extension boundary, and extracting geometric structural features of the occlusion region extension boundary.

[0064] Perform region growing processing on the occlusion region segmentation result O1'.

[0065] Step S1351: selecting a seed point from the occluded region segmentation result, and setting a region growing threshold based on the grayscale value distribution of the seed point.

[0066] In the occluded region segmentation result O1', randomly select several pixels as seed points. Analyze the grayscale distribution of these seed points and calculate the mean μ and standard deviation σ of the grayscale values. Set the region growing threshold T5 = μ + k * σ (where k is an adjustable coefficient).

[0067] Step S1352: growing the pixel region toward the adjacent pixel region with the seed point as the center until the grayscale difference between the adjacent pixels exceeds the region growth threshold, and then stopping the growth to obtain an initial growth region.

[0068] With each seed point as the center, grow toward its adjacent pixel region. For each adjacent pixel point P(x, y), calculate the grayscale difference between it and the seed point. If the grayscale difference is less than or equal to T5, then add the pixel to the growth region. Repeat this process until the grayscale difference of all adjacent pixels exceeds T5, at which point the initial growth region G1 is obtained.

[0069] Step S1353: performing a morphological closing operation on the edge of the initial growth area to generate a smoothly connected occlusion area extension boundary.

[0070] Perform a morphological closing operation on the edges of the initial growth region G1. Morphological closing generally consists of a dilation and an erosion operation. First, dilate the edges of the initial growth region G1 using a structuring element (such as a square). The dilation operation expands the edge pixels outward, filling any small holes and gaps. Then, perform an erosion operation on the dilated edge. The erosion operation contracts the dilated edge inward, removing any excess pixels generated by the dilation and smoothing the edge. This morphological closing operation generates a smoothly connected extended boundary B1 of the occluded region.

[0071] Step S1354: extracting a set of corner point coordinates of the extended boundary of the occluded area, and calculating the Euclidean distance between adjacent corner points and the angle parameter formed by three consecutive corner points.

[0072] For the extended boundary B1 of the occluded area, a corner detection algorithm is used, such as the Harris corner detection algorithm, which calculates the autocorrelation matrix of each pixel on the boundary and analyzes its eigenvalue to determine whether it is a corner point. After detection, the corner point coordinate set C = {(x1, y1), (x2, y2), ..., (x n ,y n )}.

[0073] For adjacent corner points in the corner point set C, such as (x i ,y i ) and (x i+1 ,y i+1 ), according to the Euclidean distance formula, that is, calculating the straight-line distance between two points in the plane rectangular coordinate system, let the Euclidean distance be d, d =[(xᵢ₊1- xᵢ)² + (yᵢ₊1-yᵢ)²] 1 / 2 , and get the Euclidean distance between adjacent corner points.

[0074] For three consecutive corner points, such as (x i ,y i )、(x i+1 ,yi+1 ) and (x i+2 ,y i+2 ), calculate the angle parameters formed by them through vector operations. First, according to the coordinates, we get vector V1=(x i+1 -x i ,y i+1 -y i ) and vector V2=(x i+2 -x i+1 ,y i+2 -y i+1 ). Then use the vector dot product formula, let the angle be θ, cosθ= (V 1x * V 2x + V 1y * V 2y ) / (|V1|* |V2|), and then the angle parameter θ is obtained.

[0075] Step S1355: normalizing the Euclidean distance and the angle parameter respectively, and constructing a multi-dimensional geometric feature vector of each corner point based on the normalized Euclidean distance and the angle parameter.

[0076] For the calculated Euclidean distance set D={d1, d2, ..., d n}, let the maximum value be d max , the minimum value is d min For each Euclidean distance d i Perform normalization and set the normalized Euclidean distance as d i ',d i '=(d i -d min ) / (d max -d min ).

[0077] For the angle parameter set Θ={θ1,θ2,…,θ n}, let its maximum value be θ max , the minimum value is θ min For each angle parameter θ i Perform normalization and set the normalized angle parameter as θ i ',θ i '=(θ i -θ min ) / (θ max -θ min ).

[0078] For each corner point, the normalized Euclidean distance d i ' and angle parameter θ i 'Combined together, construct the multidimensional geometric feature vector F = {(d_1', θ_1'), (d2', θ2'), ..., (d_ n ',θ_ n ')}.

[0079] Step S1356: Density clustering is performed on the geometric feature vector to divide it into multiple corner point clustering groups. Least squares straight line fitting is performed on each corner point clustering group to generate straight line segment primitives, and the directional difference between adjacent straight line segment primitives is calculated. If the directional difference is less than a preset angle threshold, the adjacent straight line segments are merged into a continuous broken line.

[0080] Using density clustering algorithm, such as DBSCAN algorithm, with multi-dimensional geometric feature vector F as input data, by setting two parameters, neighborhood radius ε and minimum number of points minPts. For each geometric feature vector f i , calculate its distance to other vectors in the neighborhood (such as Euclidean distance). If the number of vectors within the radius ε is greater than or equal to minPts, then these vectors are divided into a cluster group. In this way, multiple corner point cluster groups G1, G2, ..., G n .

[0081] For each corner cluster group, such as G i , using the least squares straight line fitting method. Assume that the corner cluster group G i The set of corner coordinates in is {(x1, y1), (x2, y2), ..., (x m ,y m )}, according to the least squares principle, find a straight line y = ax + b, so that the sum of the squares of the distances from all corner points to the straight line is minimized. By solving the relevant equations (constructing the equations and solving them according to the least squares principle), the parameters a and b of the straight line are obtained, thus generating the straight line segment primitive L i .

[0082] For adjacent straight line segment primitives L i and L i+1 , calculate their directional difference. Let the straight line L i The slope of the straight line L is k1. i+1 The slope is k2. According to the relationship between slope and angle, let the straight line L i The inclination angle is α1, tanα1=k1, and the straight line L i+1 The tilt angle is α2, tanα2=k2. The direction difference β=|α1-α2|. Set a preset angle threshold T angle , if β is less than T angle, then merge adjacent straight line segments L i and L i+1 It is a continuous broken line P1.

[0083] Step S1357: performing endpoint interpolation processing on the merged polyline to generate a closed polygon boundary, and optimizing the vertex coordinates of the closed polygon to generate the geometric structure features of the extended boundary of the occlusion area.

[0084] For the merged continuous polyline P1, let its endpoints be (x start ,y start ) and (x end ,y end ). Use an interpolation algorithm, such as linear interpolation, to insert several points between the endpoints so that the polyline forms a closed polygon boundary. Let the number of inserted points be n insert , for the jth insertion point, its x coordinate is x j =x start +j*(x end -x start ) / (n insert+1 ), the y coordinate is y j =y start +j*(y end -y start ) / (n insert+1 ), j = 1, 2, …, n insert .

[0085] After obtaining the closed polygon boundary, its vertex coordinates are optimized. By analyzing the pixel information inside and around the polygon (such as gray value, gradient and other features), for each vertex coordinate (x i ,y i ), adjusting the vertex coordinates based on the characteristics of the surrounding pixels. For example, if the grayscale values of the pixels on one side of the vertex vary significantly, this indicates that the vertex may need to be moved a certain distance in a direction with less grayscale value variation. Through multiple iterative adjustments, the polygon boundary is made to more closely match the actual shape of the occluded area, ultimately generating the geometric structure feature S1 of the extended boundary of the occluded area.

[0086] Step S136: constructing a joint feature representation of the target to be identified based on the normalized local texture features, the encoded global morphological feature vector, and the geometric structure features.

[0087] The normalized local texture feature is denoted as T, the encoded global morphological feature vector is denoted as M, and the geometric structure feature is denoted as S1. First, ensure that the dimensions of these three features match. If the local texture feature T is a multidimensional vector, the global morphological feature vector M is also a multidimensional vector, and the geometric structure feature S1 also has a set dimension. Assume that the dimension of T is (t1, t2, ..., tn ), the dimension of M is (m1, m2, ..., m m ), the dimension of S1 is (s1, s2, ..., s k ).

[0088] In order to construct the joint feature representation, the three features are spliced together in a set order. The local texture feature T can be put in the front first, followed by the global morphological feature vector M, and finally the geometric structure feature S1 to obtain the joint feature representation F. joint = [T, M, S1], whose dimensions are (t1, t2, ..., t n , m1, m2, …, m m ,s1,s2,…,s k ). This joint feature representation integrates multiple feature information of the target to be identified, from local texture to global morphology and geometric structure of the occluded area.

[0089] Step S137: performing time series modeling processing on the feature change trend in multiple frames of continuous monitoring images based on the joint feature representation to generate the visible feature distribution and the occlusion feature distribution.

[0090] The joint feature representation in multiple frames of continuous monitoring images is temporally modeled to generate visible feature distribution and occluded feature distribution.

[0091] Step S1371: performing time series alignment processing on the joint feature representation in multiple frames of continuous monitoring images to generate a time series feature vector.

[0092] Assume that the joint feature representations corresponding to multiple frames of continuous monitoring images are F joint1 , F joint2 ,…,F jointn Since different frames of images are acquired in a certain time sequence, and the joint feature representation may have slight differences in dimension or feature order due to slight differences in image acquisition, time series alignment processing is required.

[0093] First, analyze each joint feature representation F jointi The structure and dimension information of jointi The dimensions are (d1, d2, ..., d k ). For each dimension d j , find the corresponding dimension components in the joint feature representations of different frames. By comparing the changes in the eigenvalues of the corresponding dimensional components in different frames, a conventional feature matching algorithm in related technologies is used to adjust each joint feature representation so that they are consistent in the time dimension.

[0094] After adjustment, these joint feature representations are arranged in chronological order to generate the time series feature vector Ftime =[F joint 1', F joint 2',…,F jointn '], where F jointi ' is the joint feature representation after alignment, so the time series feature vector F time It is continuous in the time dimension and consistent in the feature dimension, providing a good foundation for subsequent time series analysis.

[0095] Step S1372: Input the time series feature vector into the input layer of the pre-trained time series analysis model, wherein each time step of the time series feature vector contains multiple feature channels, and each feature channel is subjected to mean-variance normalization processing respectively.

[0096] The time series feature vector F time Input into the pre-trained time series analysis model. Taking the common time series analysis model based on the recurrent neural network (RNN) architecture as an example, the time series analysis model includes an input layer, a hidden layer, and an output layer.

[0097] Time series feature vector F time Each time step in F jointi ', contains multiple feature channels, which correspond to different dimensions of information in the joint feature representation. For example, different dimensions of local texture features, different dimensions of global morphological feature vectors, and different dimensions of geometric structure features.

[0098] For each feature channel, perform mean-variance normalization. Suppose the value of a feature channel in the time series feature vector is {v1, v2, ..., v n}, first calculate the mean value of the feature channel value μ = (v1+v2+…+v n ) / n, and then calculate the variance σ² = [(v1- μ)² +(v2- μ)² + … +(v n - μ)²] / n. Then for each value v i Normalize and set the normalized value to v i ', vᵢ' = (vᵢ - μ) / (σ²) 1 / 2Through this mean-variance normalization process, the data of each feature channel has similar scale and distribution, which helps the model better learn feature changes.

[0099] The time series feature vector that has been normalized by mean and variance is input into the input layer of the time series analysis model. The input layer receives this data and passes it to the subsequent hidden layer for further processing.

[0100] Step S1373: extracting the forward time series features and the backward time series features of the time series feature vector through the bidirectional long short-term memory network layer of the time series analysis model, wherein the forward time series features and the backward time series features have the same dimension and are time-step aligned.

[0101] In the time series analysis model, the data passed from the input layer enters the bidirectional long short-term memory (Bi-LSTM) layer. The Bi-LSTM layer consists of a forward LSTM and a backward LSTM.

[0102] The forward LSTM is fed into the time series feature vector F time The data is processed from the starting time step of , and as the time step progresses, the forward time series features in the data are gradually learned. Let the hidden state of the forward LSTM at the i-th time step be h fi , which is obtained by the hidden state h of the previous time step f(i-1) and the input data x at the current time step i (i.e., time series feature vector F time Middle i The calculation is performed using the normalized data for each time step.

[0103] Backward LSTM from the time series feature vector F time The data is processed backwards from the last time step, and the backward time series features in the data are learned as the time steps are traced back. Let the hidden state of the backward LSTM at the i-th time step be h bi , which passes the hidden state h of the next time step b(i+1) and the input data x at the current time step i Perform calculations.

[0104] Through such bidirectional processing, the forward LSTM and the backward LSTM extract the forward time series features H of the time series feature vector respectively. f ={h f1 , h f2 ,…,h fn} and backward temporal features H b ={h b1 , h b2 ,…,h bnDue to the design of the Bi-LSTM layer, the forward and backward time series features have the same dimension and are time-aligned. This fully utilizes the forward and backward information of the time series data and provides rich time series features for subsequent feature fusion.

[0105] Step S1374: splice the forward time series features and the backward time series features according to the time step and input them into the multi-head self-attention layer, calculate the attention weight matrix of different feature channels in each time step, and the dimension of the attention weight matrix matches the dimension of the spliced time series features.

[0106] The forward time series feature H f and the backward temporal feature H b Splicing is performed by time step. i time steps, h fi and h bi Splice together to get the spliced time series feature h combinedi The spliced time series features of all time steps form the spliced time series feature vector H combined ={h combined1 , h combined2 ,…,h combinedn}.

[0107] H combined Input to the multi-head self-attention layer. The multi-head self-attention layer passes through multiple heads (assuming n heads ) to calculate different attention weights in parallel.

[0108] For each head, at the i-th time step, calculate the attention weights between different feature channels. Let the spliced temporal feature h combinedi The dimensions are (d1, d2, ..., d k ), h is transformed by linear transformation combinedi Projected into three different vector spaces, we get the query vector Q i , key vector ki Sum value vector V i , their dimensions are also (d1, d2, ..., d k ).

[0109] Then, calculate the attention score matrix A i , A i Element a ᵢj = (Q ᵢj × K ᵢj ) / (d k ) 1 / 2 , where j represents the index of the feature channel. The attention score matrix A is processed by the softmax functioni Perform normalization to obtain the attention weight matrix W i , W i The element w ij represents the attention weight of the jth feature channel at the i-th time step.

[0110] This calculation is performed for each time step to obtain the attention weight matrix set W={W1, W2, ..., W n}, the dimension of the attention weight matrix matches the dimension of the spliced temporal feature, that is, combinedi The dimensions of are the same, so that attention weights can be assigned according to the importance of different time steps and feature channels to highlight key features.

[0111] Step S1375: Perform weighted fusion on the spliced temporal features based on the attention weight matrix to generate a fused temporal feature vector.

[0112] For the spliced time series feature vector H combined ={h combined1 , h combined2 ,…,h combinedn} and the attention weight matrix set W = {W1, W2, ..., W n}, perform weighted fusion.

[0113] In the i time steps, and the concatenated temporal features h combinedi The feature channel values are {v1, v2, ..., v k}, attention weight matrix W i The elements are {w1, w2, ..., w k}. Weight each feature channel value to obtain the weighted feature channel value v i '=w1*v1+w2*v2+…+w k *v k .

[0114] Perform such weighted operations on all time steps to obtain the fused time series feature vector H fused ={h fused1 , h fused2 ,…,h fusedn}, where h fusedi is the feature vector after weighted fusion of the i-th time step. Through this weighted fusion, the information of different time steps and feature channels is integrated, highlighting the important temporal features and providing more valuable feature representation for subsequent predictions.

[0115] Step S1376: Input the fused time series feature vector into the fully connected regression layer for multi-step prediction, and output the visible feature prediction vector and the occlusion feature prediction vector for the next N time steps. The number of channels of the visible feature prediction vector and the occlusion feature prediction vector is consistent with the input feature vector.

[0116] The fused time series feature vector H fused Input to the fully connected regression layer. The fully connected regression layer consists of multiple neurons, each of which is connected to all neurons in the previous layer (that is, the layer where the fused time series feature vector is located).

[0117] The fully connected regression layer predicts the visible and occluded features of the next N time steps by learning and fusing the feature information in the time series feature vector. Let the weight matrix of the fully connected regression layer be W f c, bias vector is b f c.

[0118] For the fused time series feature vector H fused Each time step feature vector h fusedi , through matrix multiplication and addition operations, that is, h fusedi '=W fc *h fusedi +b fc , and get the prediction result. The prediction result contains the visible feature prediction vector V for the next N time steps pred ={v pred1 , v pred2 ,…,v predn} and occlusion feature prediction vector O pred ={o pred1 , o pred2 ,…,o predn}.

[0119] The number of channels of the visible feature prediction vector and the occlusion feature prediction vector is the same as the input feature vector (i.e., the fused temporal feature vector H fused The eigenvector h in fusedi ), which ensures that the prediction results are consistent with the original features in the feature dimension, so that the visible feature distribution and occlusion feature distribution can be accurately analyzed and processed later.

[0120] Step S1377: Perform first-order difference operations on the visible feature prediction vector and the occlusion feature prediction vector respectively to generate the feature change amount of each time step relative to the previous moment, and construct the future change trend of the visible feature distribution and the future change trend of the occlusion feature distribution based on the change amount sequence of consecutive time steps.

[0121] For the visible feature prediction vector V pred ={v pred1, v pred2 ,…,v predn}, perform first-order difference operation. Let the feature vector of the t-th time step in the visible feature prediction vector be v predt , the eigenvector of the t-1th time step is v pred(t-1) For v predt Each channel value vt in i (i represents the channel index) and v pred(t-1) The corresponding channel value v in (t-1)i , calculate the first-order difference, that is, the characteristic change Δv ti =vt i -v (t-1)i In this way, for each time step t (from 2 to N), a set of feature changes can be obtained to form a feature change sequence ΔV = {Δv2, Δv3, ..., Δv_ n}, where Δv t Is a v predt A vector with the same number of channels contains the feature change of each channel at this time step relative to the previous moment.

[0122] Similarly, for the occlusion feature prediction vector O pred ={o pred1 , o pred2 ,…,o predn}, perform first-order difference operation. Let the feature vector of the t-th time step in the occlusion feature prediction vector be o predt , the eigenvector of the t-1th time step is o pred(t-1) For o predt Each channel value in ot i (i represents the channel index) and o pred(t-1) The corresponding channel value o in (t-1)i , calculate the first-order difference, that is, the characteristic change Δo ti =ot i -o (t-1)i Thus, the feature variation sequence of the occlusion feature prediction vector ΔO={Δo2, Δo3, ..., Δo_ n}, where Δo t Is a with o predt A vector with the same number of channels contains the feature change of each channel at this time step relative to the previous moment.

[0123] The future change trend of the visible feature distribution is constructed based on the feature change sequence ΔV of the visible feature prediction vector. The change of each channel change in the feature change sequence with the time step is analyzed. For example, observe whether the feature change of a certain channel shows an increasing, decreasing or periodic change trend at different time steps. Suppose there is a channel whose feature change gradually increases in several consecutive time steps, which may indicate that the visible feature corresponding to the channel has an increasing trend in the future. By analyzing the feature change sequence of all channels, the future change trend of the visible feature distribution is comprehensively obtained. This change trend can be represented as a multidimensional vector sequence T V , where each element corresponds to a description of the change trend of a time step. The change trend description can be a certain feature representation obtained based on the analysis of the change amount of each channel feature, such as a comprehensive value obtained by performing some weighted combination or cluster analysis on the change amount of each channel feature to reflect the change direction and degree of the entire visible feature distribution at that time step.

[0124] Similarly, the future change trend of the occlusion feature distribution is constructed based on the feature change sequence ΔO of the occlusion feature prediction vector. A similar analysis is performed on the change of each channel change in the occlusion feature change sequence over time. For example, if the feature change of a channel first decreases and then increases within several time steps, this reflects a specific change pattern of the occlusion feature corresponding to the channel. By analyzing all channels, the future change trend of the occlusion feature distribution is expressed as a multidimensional vector sequence T O ,Each element is also a comprehensive value obtained based on the analysis of the change in the characteristics of each channel, which is used to describe the change in the distribution of occlusion features at the corresponding time step.

[0125] Step S1378: performing compensation and correction processing on the visible feature distribution and the occlusion feature distribution in the current frame monitoring image according to the future change trend, and generating updated visible feature distribution and occlusion feature distribution.

[0126] For the visible feature distribution in the current frame monitoring image, let it be V current , which is a multidimensional vector, each dimension corresponds to a different visible feature channel. According to the future change trend of the visible feature distribution T V , for V current Make compensation correction. Assume T V The trend vector of the t-th time step is t Vt , which contains the description information of the future changes of each channel of the visible feature distribution.

[0127] For V current Each channel value v in i (i represents the channel index), according to t VtFor example, if t Vt The information of a certain channel in the equation indicates that the channel feature has an increasing trend in the future, so V is increased accordingly. current The specific adjustment method can be through some weighted relationship, assuming the weight is w i , then the adjusted channel value v i '=v i +w i *t V t i , where t V t i It is t Vt The value of the corresponding channel in V current All channels of the tfl are adjusted in this way to obtain the updated visible feature distribution V updated .

[0128] Similarly, for the occlusion feature distribution in the current frame monitoring image, let it be O current , which is also a multidimensional vector. According to the future change trend of the occlusion feature distribution T O Perform compensation correction. Assume T O The trend vector of the t-th time step is t ot , for O current Each channel value o in i ( i represents the channel index), in a similar way to the visible feature distribution, according to t ot The information of the corresponding channel and the corresponding weight w i 'Adjust, that is, the adjusted channel value o i '=o i +w i '*t oti , where t oti It is t ot By adjusting all channels, the updated occlusion feature distribution O is generated. updated Through this compensation and correction process, the visible feature distribution and occlusion feature distribution in the current frame monitoring image can better reflect the future changes predicted based on multi-frame analysis, providing more accurate feature information for subsequent target recognition and parameter optimization.

[0129] Step S140: generating a dynamic recognition result of the target to be identified according to the visible feature distribution and the occlusion feature distribution, wherein the dynamic recognition result is used to characterize a state transition trajectory of the target to be identified in multiple frames of continuous monitoring images.

[0130] After obtaining the updated visible feature distribution V updated and occlusion feature distribution O updatedFinally, these feature information are used to generate dynamic recognition results of the target to be identified.

[0131] First, V updated and O updated Since they are both multidimensional vectors, and the dimensions correspond to the feature channels processed previously, the two vectors are concatenated together in the set order to form a fused feature vector F fusion =[V updated , O updated ]. The fused feature vector combines the information of visible features and occlusion features.

[0132] Then, based on the fusion feature vector F fusion To generate dynamic recognition results, conventional pattern recognition algorithms in related technologies are used to analyze the changes in the fused feature vector at different time steps. For example, the state changes of the target to be identified can be determined by observing the change patterns of the values of each dimension in the fused feature vector over time.

[0133] Assuming that at a certain time step, certain dimensional values of the fused feature vector undergo a specific combination of changes, based on pre-set rules or a model trained through machine learning, it can be determined that the target to be identified has entered a new state. By analyzing the fused feature vectors in multiple frames of continuous monitoring images frame by frame, the state of the target to be identified at different time steps is recorded, thus generating a state transition trajectory of the target to be identified in the multiple frames of continuous monitoring images.

[0134] The state transition trajectory is the dynamic recognition result, which can be expressed as a state sequence S={s1, s2, ..., s n}, where s n Indicates the n The state of the target to be identified at each time step. Each state can be a feature description derived from the analysis of the fused feature vector. For example, by performing cluster analysis on the fused feature vector, different clustering results can be defined as different states. Each state represents a comprehensive feature expression of the target to be identified at that moment, which may include a comprehensive reflection of information such as the target's position, posture, and degree of occlusion. In this way, a dynamic recognition result is generated that can characterize the state transition of the target to be identified in multiple frames of continuous monitoring images.

[0135] Step S150: performing real-time optimization processing on the acquisition parameters of the acquisition module based on the dynamic recognition result, and generating a parameter adjustment strategy for the next monitoring cycle.

[0136] After obtaining the dynamic recognition result of the target to be recognized, the acquisition parameters of the acquisition module are optimized in real time according to the result to generate the parameter adjustment strategy for the next monitoring cycle.

[0137] Step S151: Calculating a motion speed estimation value and a motion direction estimation value of the target to be identified according to the state transition trajectory in the dynamic identification result.

[0138] Analyze the state transition trajectory S={s1, s2, ..., s n Assume that each state s contains the position information of the target to be identified at the corresponding time step, and let the position of the i-th time step be (x i ,y i ), the position of the i+1th time step is (x( i+1 ), y( i+1 )).

[0139] Calculate the estimated value of the motion speed. In the plane coordinate system, the speed is calculated based on the position change and the time interval. Let the time interval be Δt (assuming that the time interval of each time step is the same), then the velocity component in the x direction vx = (x( i+1 )-x i ) / Δt, velocity component v in the y direction y =(y( i+1 )-y i ) / Δt. Using these two velocity components, the Pythagorean theorem is used to calculate the total velocity v=(v x ²+v y ²) 1 / 2 , the combined velocity v is the estimated value of the moving speed of the target to be identified.

[0140] Calculate the estimated value of the direction of movement. Determine the direction of movement based on the velocity component. Let the angle between the direction of movement and the positive direction of the x-axis be θ, then tanθ=v y / v x , through the inverse tangent function arctan(v y / v x ) to obtain the angle θ, where θ is the estimated value of the moving direction of the target to be identified. Through such calculations, the estimated value of the moving speed and moving direction of the target to be identified are obtained from the state transition trajectory.

[0141] Step S152: predicting the expected location area of the target to be identified in the next monitoring period based on the estimated value of the motion speed and the estimated value of the motion direction.

[0142] According to the calculated motion speed estimation value v and motion direction estimation value θ, the expected position area of the target to be identified in the next monitoring cycle is predicted.

[0143] Assume that the current position of the target to be identified is (x0, y0) and the duration of the next monitoring cycle is T. In the x-direction, the position change Δx = v*cosθ*T is calculated based on the speed and time, and the position change Δy = v*sinθ*T in the y-direction.

[0144] Then the expected position of the target to be identified in the next monitoring cycle in the x direction is x1=x0+Δx, and the expected position in the y direction is y1=y0+Δy.

[0145] With (x1, y1) as the center, set the expected position area according to the set range. For example, set a circular area with a radius of r as the expected position area, and all points (x, y) in the expected position area meet (x - x1)² + (y- y1)² ≤ The radius r² can be set based on actual conditions, such as the target's motion stability and the acquisition module's accuracy. This calculation and setting yields the expected location of the target in the next monitoring cycle, providing a basis for adjusting acquisition module parameters.

[0146] Step S153: adjusting the focal length parameter and the shooting angle parameter of the acquisition module according to the expected position area, and generating a first parameter adjustment instruction.

[0147] According to the expected location area of the target to be identified in the next monitoring cycle, the focal length parameters and shooting angle parameters of the acquisition module are adjusted to ensure that the target can be clearly photographed.

[0148] To adjust the focal length parameter, analyze the imaging relationship between the target location and the current acquisition module. Assume that the acquisition module's imaging principle follows certain optical imaging laws, and the distance d from the target location to the acquisition module is the target location. If d is large, the focal length needs to be increased to clearly image the target within the target location. Conversely, if d is small, the focal length needs to be decreased. Let the current focal length be f0. Based on the relationship between distance d and focal length (for example, using a pre-established distance-focal length mapping table or calculation logic based on optical imaging principles), the adjusted focal length f1 is calculated.

[0149] To adjust the shooting angle parameters, use the acquisition module's current shooting direction as a reference and calculate the angle adjustment required to center the target location area within the acquisition module's field of view. Let α be the angle between the current shooting direction and the positive x-axis, and β be the angle between the direction of the target location area's center relative to the acquisition module's current position and the positive x-axis. The shooting angle adjustment Δα is then calculated as β - α. Based on this angle adjustment, determine the direction and angle in which the acquisition module needs to rotate to position the target location area within the acquisition module's effective field of view.

[0150] The adjusted focal length parameter f1 and the shooting angle adjustment amount Δα are combined to generate a first parameter adjustment instruction I1, which is used to guide the acquisition module to adjust the focal length and shooting angle in the next monitoring cycle to better capture the target to be identified.

[0151] Step S154: predicting a potential occlusion area in the next monitoring period based on the historical change trend of the occlusion feature distribution, and adjusting the exposure parameters and resolution parameters of the acquisition module according to the potential occlusion area to generate a second parameter adjustment instruction.

[0152] Review the historical trend of occlusion feature distribution in the previous multi-frame continuous monitoring images. Assume that the occlusion feature distribution of the previous frames is O history ={O1,O2,…,O m}, and analyze the changes in the distribution of these occlusion features.

[0153] Observe the changing patterns of each part of the occlusion feature distribution, for example, whether the occlusion features in certain areas persist over several consecutive frames or appear and disappear regularly. By analyzing this historical data, predict the potential occlusion areas that may be occluded in the next monitoring cycle.

[0154] Assuming that through analysis it is found that a certain area has been frequently occluded in the past multiple frames, and the changes in its occlusion characteristics have a certain stability, then this area will be regarded as the potential occlusion area P in the next monitoring cycle.

[0155] For the potential occlusion area P, adjust the exposure and resolution parameters of the acquisition module. If the brightness of the potential occlusion area P is low, in order to clearly display the details in this area, the acquisition module's exposure intensity in this area needs to be increased. Assume that the current acquisition module's exposure intensity for the entire image is E0. Based on the brightness of the potential occlusion area P (for example, by analyzing the grayscale value distribution of this area in historical images to determine the brightness), calculate the exposure intensity ΔE that needs to be increased in this area. The adjusted exposure intensity of the potential occlusion area P is E1=E0+ΔE. At the same time, adjust the exposure intensity of other areas accordingly to ensure exposure balance for the entire image.

[0156] Regarding resolution parameter adjustment, if the acquisition module supports regional adaptive resolution adjustment and the potential occlusion region P is of high importance (for example, it may contain critical target information), the image acquisition resolution of the potential occlusion region P is increased. Assuming the overall resolution of the current acquisition module is R0, the resolution of the potential occlusion region P is increased to R1. At the same time, the resolution of adjacent regions is adjusted through an interpolation algorithm to ensure image continuity and consistency.

[0157] The adjusted exposure parameters (including the exposure intensity of the potential occlusion area P and other areas) and resolution parameters (the resolution of the potential occlusion area P and adjacent areas) are combined together to generate a second parameter adjustment instruction I2. The second parameter adjustment instruction I2 is used to guide the acquisition module to optimize the exposure and resolution of the potential occlusion area in the next monitoring cycle.

[0158] Step S155: Fusing the first parameter adjustment instruction and the second parameter adjustment instruction to generate the parameter adjustment strategy.

[0159] The first parameter adjustment instruction I1 includes adjustment information of the focal length parameter and the shooting angle parameter of the acquisition module, and the second parameter adjustment instruction I2 includes adjustment information of the exposure parameter and the resolution parameter of the acquisition module.

[0160] These two instructions are combined to form a comprehensive parameter adjustment strategy P. The parameter information in I1 and I2 can be integrated according to a predefined format. For example, the focus parameter adjustment value and the shooting angle adjustment value are listed first, followed by the exposure parameter adjustment value (including potential occlusion areas and other areas), and the resolution parameter adjustment value (potential occlusion areas and adjacent areas).

[0161] The parameter adjustment strategy P will serve as the basis for the acquisition module to adjust parameters in the next monitoring cycle, so that the acquisition module can optimize and adjust the focal length, shooting angle, exposure and resolution according to the dynamic recognition results of the target to be identified, so as to improve the monitoring and recognition effect of the target.

[0162] Next, we will describe the model training process. When building a pre-trained segmentation network model, we can choose a network architecture suitable for semantic segmentation tasks, such as the UNet architecture. This architecture consists of an encoder, a decoder, and skip connections connecting the two.

[0163] The encoder consists of multiple downsampling blocks. Each downsampling block includes a convolutional layer and a pooling layer. The convolutional layer convolves the input image using different convolution kernels (K1, K2, K3, etc.). The convolution kernel slides across the image to extract various features, such as edges and textures. The pooling layer downsamples the convolved feature map, reducing the data size while retaining important features.

[0164] The decoder corresponds to the encoder and consists of multiple upsampling blocks. These blocks restore low-resolution feature maps to high resolution through deconvolution or interpolation. They also include convolutional layers for further feature fusion and refinement.

[0165] The skip connection connects the feature maps of different levels in the encoder with the feature maps of the corresponding levels in the decoder, so that the decoder can utilize the feature information of different levels of the encoder in the process of restoring the resolution, thereby improving the accuracy of segmentation.

[0166] Among them, a large amount of image data containing the targets to be identified and occlusion situations is collected as a training set. The image data should cover different scenes, different types of targets to be identified and various occlusion situations.

[0167] Each image is annotated to identify the target area and the occluded area. Annotation can be done using pixel-level masks, which label each pixel in the image as belonging to the target area, the occluded area, or the background.

[0168] Divide the labeled image data into a training set, a validation set, and a test set. For example, you can divide the data into a set ratio (e.g., 70% for training, 15% for validation, and 15% for testing). The training set is used to learn model parameters, the validation set is used to adjust model hyperparameters to prevent overfitting, and the test set is used to evaluate the final performance of the model.

[0169] Set the learning rate, which controls the step size of each parameter update. Let the learning rate be lr. Usually, a small value such as 0.001 is chosen to ensure that the model can converge stably during training.

[0170] Set the number of training epochs, which represents the number of times the model will complete the training set. For example, setting epoch to 100 means that the model will complete the training set 100 times.

[0171] Set batch size _size , that is, the number of images input to the model each time training. For example, setting batch _size If it is 16, it means that 16 images are selected from the training set and input into the model for training at the same time.

[0172] The image data in the training set is sequentially input into the constructed segmentation network model. The image first enters the encoder, where it undergoes convolution and pooling operations to extract features.

[0173] In the decoder, the feature map is restored to the same size as the input image through upsampling and convolution operations to generate a predicted segmentation mask. The predicted segmentation mask is compared with the annotated ground-truth mask, and the difference between the two is measured by a loss function.

[0174] Choose a suitable loss function, such as the cross entropy loss function. For each pixel in the predicted mask, the difference between its predicted category probability and the true category is quantified by the cross entropy loss function. Assume that the predicted mask is P and the true mask is G. For pixel i, its predicted category probability is P i , the true category is G i , the cross entropy loss function calculates the loss value L of the pixel i , by summing and averaging the loss values of all pixels, we get the loss value L of the entire image.

[0175] Based on the loss value L, an optimization algorithm is used to update the model parameters. For example, the stochastic gradient descent (SGD) algorithm adjusts the parameters based on the gradient of the loss function with respect to the model parameters. The gradient ∇L / ∇θ of the loss function L with respect to each parameter θ in the model is calculated. The parameter update formula is θ = θ - lr * ∇L / ∇θ, where lr is the set learning rate. Through continuous iteration of this process, the model parameters are gradually optimized and the loss value is continuously reduced.

[0176] After each epoch of training, input the validation set image data into the model and calculate the validation set loss. Observe the changes in the validation set loss. If the validation set loss stops decreasing after several consecutive epochs, or even starts to increase, the model may be overfitting. In this case, you can adjust the learning rate, reduce model complexity, and so on to avoid overfitting. For example, reduce the learning rate to one-tenth of its original value (lr = lr / 10), and then continue training.

[0177] After a set number of epochs, the model completes training on the training set. The trained model is then evaluated using the test set, calculating metrics such as segmentation accuracy and recall. For example, segmentation accuracy is calculated as the ratio of correctly segmented pixels to the total number of pixels, while recall is calculated as the ratio of correctly segmented target pixels to the actual number of target pixels. These metrics are used to determine whether the model meets the expected performance requirements. If not, further adjustments to the model structure or training parameters can be made and training can be repeated.

[0178] Furthermore, a time series analysis model was constructed using a bidirectional long short-term memory (Bi-LSTM) network combined with a multi-head self-attention mechanism. The model mainly consists of an input layer, a bidirectional long short-term memory network layer, a multi-head self-attention layer, and a fully connected regression layer.

[0179] The input layer is responsible for receiving the time series feature vector. The time series feature vector is a joint feature representation sequence after time series alignment and mean variance normalization. Assume that the time series feature vector is F time , the dimension of the feature vector of each time step is d, and the input layer will be F time Input to subsequent layers sequentially by time step.

[0180] The bidirectional long short-term memory network layer consists of a forward LSTM and a backward LSTM. The forward LSTM processes data from the beginning of the time series, and the backward LSTM processes data in reverse from the end of the time series. For the forward LSTM, at the tth time step, it receives the hidden state h of the previous time step. f(t-1) and the input feature x at the current time step t , through the internal forget gate, input gate and output gate mechanisms, update the hidden state h ft At the tth time step, the backward LSTM receives the hidden state h of the next time step. b(t+1) and the input feature x at the current time step t , also update the hidden state h through the internal mechanism b t. The forward LSTM and backward LSTM extract the forward time series features H f and the backward temporal feature H b , and the two time series features have the same dimensions and are time-step aligned.

[0181] The multi-head self-attention layer receives the forward time series feature H f and the backward temporal feature H b The concatenated feature vector H by time step combined This layer passes multiple heads (assuming n heads For each head, at the tth time step, H combined The current time step feature vector in is projected to the query vector Q through linear transformation t , key vector K t Sum value vector V t . Then calculate the attention score matrix A t , whose element a ij According to Q t and K t The dot product and feature dimension of A are calculated, and then the softmax function is used to calculate the feature dimension of A. t Normalize to get the attention weight matrix W t , W t The element w ijrepresents the attention weight of the jth feature channel at the tth time step. The multi-head self-attention layer highlights the importance of different time steps and feature channels in this way.

[0182] The fully connected regression layer receives the fused time series feature vector H after weighted fusion of the multi-head self-attention layer fused The fully connected regression layer consists of multiple neurons, each of which is connected to H fused All elements in are connected. Through the weight matrix W f c and bias vector b f c, for H fused Perform linear transformation and bias addition operation, that is, y=W f c*H fused +b f c. Output the visible feature prediction vector and occlusion feature prediction vector for the next N time steps.

[0183] Furthermore, time series data is extracted from the joint feature representation of multiple frames of continuous monitoring images as a training set. Assume that the joint feature representation sequence is F jointsequence ={F joint1 , F joint2 ,…,F jointm}, divide it into multiple time series segments in chronological order. Each time series segment contains several consecutive time steps, and each segment contains T time steps. For example, the first time series segment is {F joint1 , F joint2 ,…,F jointT}, the second time series segment is {F joint2 , F joint3 ,…,F joint(T+1)}, and so on.

[0184] For each time series segment, the corresponding true visible feature prediction vector and true occlusion feature prediction vector are determined according to the actual visible feature distribution and occlusion feature distribution, and the true vector is used as the training label data.

[0185] All time series segments and their corresponding labeled data are divided into training, validation, and test sets. Again, this division is done according to a set ratio (e.g., 70% for training, 15% for validation, and 15% for testing). The training set is used for model parameter learning, the validation set is used for adjusting model hyperparameters, and the test set is used for evaluating model performance.

[0186] Set the learning rate lr seq , which controls the step size of each parameter update in the time series analysis model. Similar to the learning rate setting of the segmentation network model, a small value such as 0.0001 is usually selected to ensure the stability and convergence of the model during training.

[0187] Set the number of training rounds epoch seq , which indicates the number of times the model performs a complete traversal of the training set. For example, setting epoch seq If set to 50, the model will learn the training set 50 times.

[0188] Fixed batch size _sizeseq , which is the number of time series segments that are input to the model each time training. For example, setting batch _sizeseq If it is 8, it means that 8 time series segments are selected from the training set and input into the model for training at the same time.

[0189] The time series segments in the training set are sequentially input into the constructed time series analysis model. The time series segments first enter the input layer and then pass to the bidirectional long short-term memory network layer.

[0190] In the bidirectional long short-term memory network layer, the forward LSTM and backward LSTM process the time series segments respectively, extracting forward time series features and backward time series features. The forward time series features and backward time series features are spliced and then enter the multi-head self-attention layer.

[0191] In the multi-head self-attention layer, the attention weights of different time steps and feature channels are calculated through multiple heads, and the concatenated time series features are weighted fused to generate a fused time series feature vector.

[0192] The fused time series feature vector is input into a fully connected regression layer, which outputs the visible feature prediction vector and the occlusion feature prediction vector for the next N time steps. The predicted vector is compared with the actual visible feature prediction vector and the actual occlusion feature prediction vector, and the difference is measured using a loss function.

[0193] Take the mean square error (MSE) loss function as an example. For the predicted visible feature prediction vector V pred and the true visible feature prediction vector V true , and the predicted occlusion feature prediction vector O pred and the true occlusion feature prediction vector O true , respectively calculate the mean square error between them. For the visible feature prediction vector, the mean square error MSE V = (1 / N) * ∑ (V predi -V truei ) 2 , where i ranges from 1 to N; for the occlusion feature prediction vector, the mean square error MSE O =(1 / N)*∑(O predi -O truei ) 2, where i ranges from 1 to N. These two mean squared errors are added and averaged to get the loss value L for the entire model seq .

[0194] According to the loss value L seq , use the optimization algorithm to update the model parameters. Also choose the stochastic gradient descent (SGD) algorithm to calculate the loss function L seq For each parameter θ in the model seq The gradient ∇L seq / ∇θ seq , the parameter update formula is θ seq =θ seq -lr seq *∇L seq / ∇θ seq .

[0195] In each round of training (epoch seq ) After the training, input the time series of the validation set into the model and calculate the loss value on the validation set. Observe the changes in the validation set loss value. If the validation set loss value stops decreasing after several rounds of training, or even starts to increase, it means that the model may be overfitting. In this case, you can adjust the learning rate, reduce the model complexity, and other methods to avoid overfitting. For example, reduce the learning rate to half of the original value, i.e., lr seq =lr seq / 2, and then continue training.

[0196] After the set number of training rounds epoch seq After that, the model is trained on the training set. The performance of the trained model is evaluated using the test set, and indicators such as the prediction accuracy of the model on the test set are calculated. For example, for the visible feature prediction vector, the error rate between the predicted value and the true value is calculated, error rate = (1 / N) * ∑|V predi -V truei | / |V truei |, where i ranges from 1 to N. Similar calculations are performed for the occlusion feature prediction vector. These metrics are used to determine whether the model meets the expected performance requirements. If not, further adjustments can be made to the model structure or training parameters and retraining can be performed.

[0197] During the data collection process, when collecting images that may contain privacy-sensitive data, differential privacy technology is used to protect privacy. Differential privacy perturbs the data by adding noise, making it difficult for an attacker to infer sensitive information from the data even if they obtain partial data.

[0198] Encryption technology is employed during data storage and transmission. Symmetric encryption algorithms, such as AES (Advanced Encryption Standard), are used for captured image data and processed feature data. An encryption key is selected and the data is encrypted according to the AES algorithm's rules. During data storage, the encrypted data is stored in secure storage media. During data transmission, the encrypted data is transmitted over a secure network channel. After receiving the encrypted data, the recipient uses the same key to decrypt it according to the AES algorithm's decryption rules, recovering the original data for subsequent processing. Encryption technology further ensures the security of privacy-sensitive data during storage and transmission, preventing data theft or tampering.

[0199] Figure 2 A schematic diagram illustrates exemplary hardware and software components of a drone intelligent identification system 100 for occluded targets, provided in some embodiments of the present application, that can implement the concepts of the present application. For example, processor 120 can be used in drone intelligent identification system 100 for occluded targets and perform the functions described in the present application.

[0200] The drone intelligent identification system 100 for obstructed targets can be a general-purpose server or a special-purpose server, both of which can be used to implement the drone intelligent identification method for obstructed targets of this application. Although this application only shows a single server, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0201] For example, the drone intelligent identification system 100 for obstructed targets may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and storage media 140 in various forms, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the drone intelligent identification system 100 for obstructed targets may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application may be implemented according to these program instructions. The drone intelligent identification system 100 for obstructed targets also includes an I / O interface 150 between the computer and other input and output devices.

[0202] For ease of explanation, only one processor is described in the drone intelligent identification system 100 for obscured targets. However, it should be noted that the drone intelligent identification system 100 for obscured targets in this application may also include multiple processors, so the steps performed by one processor described in this application may also be performed jointly or individually by multiple processors. For example, if the processor of the drone intelligent identification system 100 for obscured targets executes steps A and B, it should be understood that steps A and B may also be executed jointly by two different processors or individually in one processor. For example, the first processor executes step A and the second processor executes step B, or the first processor and the second processor execute steps A and B together.

[0203] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When a processor executes the computer-executable instructions, the above-mentioned drone intelligent identification method for obscured targets is implemented.

[0204] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.

Claims

1. A method for intelligently identifying obscured drone targets, characterized in that: The method comprises: Acquire multiple frames of continuous monitoring images of the target area in real time through an acquisition module carried by the drone, wherein the multiple frames of continuous monitoring images include at least a portion of the target to be identified that is obscured; Performing occlusion region segmentation processing on the multiple frames of continuous monitoring images to obtain a target region segmentation result of the target to be identified in each frame of monitoring image and a corresponding occlusion region segmentation result; Performing dynamic feature extraction based on the target region segmentation result and the occlusion region segmentation result to obtain the visible feature distribution and occlusion feature distribution of the target to be identified in each frame of the monitoring image; Generating a dynamic recognition result of the target to be identified according to the visible feature distribution and the occlusion feature distribution, wherein the dynamic recognition result is used to characterize a state transition trajectory of the target to be identified in multiple frames of continuous monitoring images; Based on the dynamic recognition result, the acquisition parameters of the acquisition module are optimized in real time to generate a parameter adjustment strategy for the next monitoring cycle; The performing of the occlusion region segmentation process on the multiple frames of continuous monitoring images to obtain the target region segmentation result of the target to be identified in each frame of the monitoring image and the corresponding occlusion region segmentation result includes: Performing time sequence alignment processing on the multiple frames of continuous monitoring images to generate multiple frame aligned monitoring images with continuous time dimension; Calling a pre-trained segmentation network model to perform frame-by-frame semantic segmentation processing on the multiple aligned monitoring images to obtain an initial segmentation mask in each frame of the monitoring image; Extracting the contour area of the target to be identified as a candidate target area based on the initial segmentation mask, and performing edge integrity verification processing on the candidate target area to obtain a verification result; performing cross-frame contour compensation processing on the candidate target area with defects according to the verification result to obtain a compensated target area segmentation result; The occluded region segmentation result is determined based on a difference area between the compensated target region segmentation result and the initial segmentation mask.

2. The method for intelligently identifying obscured targets by drones according to claim 1, wherein: The performing edge integrity verification on the candidate target area to obtain a verification result includes: Extracting an edge pixel set of the candidate target area and calculating a curvature distribution feature of the edge pixel set; Detecting whether there is a curvature mutation point in the candidate target area based on the curvature distribution feature, wherein the curvature mutation point is used to characterize the edge defect position of the candidate target area; When a curvature mutation point is detected, the local edge segment corresponding to the curvature mutation point is extracted, and the motion-compensated similarity score between the local edge segment and the edge segment at the corresponding position in the adjacent frame monitoring image is calculated in combination with the target motion estimation result; The degree of edge defect of the candidate target region is determined according to the motion compensation similarity score. If the degree of edge defect exceeds a preset threshold, it is determined that the candidate target region needs to be subjected to cross-frame contour compensation processing.

3. The method for intelligently identifying obstructed targets by drones according to claim 1, wherein: The determining the occluded area segmentation result based on the difference area between the compensated target area segmentation result and the initial segmentation mask includes: Comparing the compensated target area segmentation result with the initial segmentation mask pixel by pixel to obtain a set of difference pixels; Performing spatial clustering processing on the difference pixel point set to generate multiple candidate occlusion regions; Performing region screening based on the spatiotemporal continuity characteristics of the candidate occlusion region in multiple frames of continuous monitoring images to obtain a stable occlusion region segmentation result; A single iterative optimization is performed on the contour information of the target area segmentation result according to the stable occlusion area segmentation result. If the overlapping area of the optimized target area and the occlusion area is less than a preset threshold, the iteration is stopped to generate an optimized target area segmentation result.

4. The method for intelligently identifying a UAV with an obstructed target according to claim 1, wherein: The dynamic feature extraction process is performed based on the target region segmentation result and the occlusion region segmentation result to obtain the visible feature distribution and occlusion feature distribution of the target to be identified in each frame of the monitoring image, including: Dividing the target region segmentation result into a plurality of analysis windows of different scales, and calculating the gradient amplitude distribution within each analysis window; Performing directional histogram statistical processing on the gradient amplitude distribution to obtain texture directional histogram features of each analysis window; Normalizing and splicing the texture direction histogram features of different scales to generate local texture features of the target to be identified; Extracting the minimum bounding rectangle of the target area segmentation result, performing binning and encoding processing on the aspect ratio and area change rate of the minimum bounding rectangle respectively, and generating the global morphological features of the target to be identified; Performing region growing processing on the occlusion region segmentation result to generate an occlusion region extension boundary, and extracting geometric structural features of the occlusion region extension boundary; Constructing a joint feature representation of the target to be identified based on the normalized local texture features, the encoded global morphological feature vector, and the geometric structure features; Based on the joint feature representation, a temporal modeling process is performed on the feature change trend in multiple frames of continuous monitoring images to generate the visible feature distribution and the occlusion feature distribution.

5. The method for intelligently identifying a UAV with an obstructed target according to claim 4, wherein: The performing region growing processing on the occlusion region segmentation result to generate an occlusion region extension boundary, and extracting geometric structural features of the occlusion region extension boundary includes: Selecting a seed point from the occluded region segmentation result, and setting a region growing threshold based on the grayscale value distribution of the seed point; Growing the pixel region toward the adjacent pixel region with the seed point as the center until the grayscale difference between the adjacent pixels exceeds the regional growth threshold, and then stopping the growth to obtain an initial growth region; Performing morphological closing operations on the edges of the initial growing region to generate smoothly connected extended boundaries of the occluded region; Extracting a set of corner point coordinates of the extended boundary of the occluded area, and calculating the Euclidean distance between adjacent corner points and the angle parameter formed by three consecutive corner points; Normalizing the Euclidean distance and the angle parameter respectively, and constructing a multidimensional geometric feature vector of each corner point based on the normalized Euclidean distance and the angle parameter; Density clustering is performed on the geometric feature vector to divide the vector into multiple corner point cluster groups. Least squares linear fitting is performed on each corner point cluster group to generate straight line segment primitives. The directional difference between adjacent straight line segment primitives is calculated. If the directional difference is less than a preset angle threshold, the adjacent straight line segments are merged into a continuous polyline. Endpoint interpolation processing is performed on the merged polyline to generate a closed polygon boundary, and the vertex coordinates of the closed polygon are optimized to generate geometric structural features of the extended boundary of the occluded area.

6. The method for intelligently identifying obstructed targets by unmanned aerial vehicles according to claim 4, wherein: The performing time series modeling processing on the feature change trend in multiple frames of continuous monitoring images based on the joint feature representation to generate the visible feature distribution and the occlusion feature distribution includes: Performing time series alignment processing on the joint feature representation in multiple frames of continuous monitoring images to generate a time series feature vector; Inputting the time series feature vector into the input layer of a pre-trained time series analysis model, wherein each time step of the time series feature vector includes multiple feature channels, and each feature channel is subjected to mean-variance normalization processing; Extracting forward time series features and backward time series features of the time series feature vector through the bidirectional long short-term memory network layer of the time series analysis model, wherein the forward time series features and the backward time series features have the same dimension and are time-step aligned; The forward time series features and the backward time series features are spliced together according to the time step and input into the multi-head self-attention layer, and the attention weight matrix of different feature channels in each time step is calculated. The dimension of the attention weight matrix matches the dimension of the spliced time series features; Performing weighted fusion on the spliced temporal features based on the attention weight matrix to generate a fused temporal feature vector; Input the fused time series feature vector into a fully connected regression layer for multi-step prediction, and output a visible feature prediction vector and an occlusion feature prediction vector for the next N time steps, where the number of channels of the visible feature prediction vector and the occlusion feature prediction vector is consistent with that of the input feature vector; Performing first-order difference operations on the visible feature prediction vector and the occlusion feature prediction vector respectively to generate feature changes at each time step relative to the previous moment, and constructing future change trends of the visible feature distribution and the occlusion feature distribution based on a sequence of changes in consecutive time steps; Compensation and correction processing is performed on the visible feature distribution and the occlusion feature distribution in the current frame monitoring image according to the future change trend to generate updated visible feature distribution and occlusion feature distribution.

7. The method for intelligently identifying a UAV with an obstructed target according to claim 1, wherein: The real-time optimization processing of the acquisition parameters of the acquisition module based on the dynamic recognition result to generate a parameter adjustment strategy for the next monitoring cycle includes: Calculating a motion speed estimation value and a motion direction estimation value of the target to be identified based on the state transition trajectory in the dynamic identification result; Predicting an expected location area of the target to be identified in the next monitoring period based on the estimated motion speed and the estimated motion direction; Adjusting the focal length parameter and the shooting angle parameter of the acquisition module according to the expected position area to generate a first parameter adjustment instruction; Predicting a potential occlusion area in the next monitoring period based on a historical change trend of the occlusion feature distribution, and adjusting exposure parameters and resolution parameters of the acquisition module according to the potential occlusion area to generate a second parameter adjustment instruction; The first parameter adjustment instruction and the second parameter adjustment instruction are merged to generate the parameter adjustment strategy.

8. The method for intelligently identifying obstructed targets by unmanned aerial vehicles according to claim 7, wherein: The step of adjusting the exposure parameter and the resolution parameter of the acquisition module according to the potential occlusion area to generate a second parameter adjustment instruction includes: Calculating the occurrence frequency and duration of the potential occlusion area in the historical monitoring period; determining a stability score of the potential occlusion region according to the occurrence frequency and the duration; When the stability score exceeds a first threshold and the acquisition module supports regional adaptive resolution adjustment, reducing the image acquisition resolution corresponding to the potential occlusion area and improving the resolution of adjacent areas through an interpolation algorithm; When the stability score is lower than a second threshold, the exposure intensity of the acquisition module corresponding to the potential occlusion area is increased, and the exposure intensity of adjacent areas is reduced.

9. An intelligent identification system for drones targeting obscured targets, characterized in that: The method comprises a processor and a memory, wherein the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the method for intelligent identification of drones for obscured targets as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Unmanned aerial vehicle holder camera target identification tracking method and system based on artificial intelligence

    CN117197695A

  • Multi-modal remote sensing image sea target identification method based on AI

    CN119274062A