Unmanned aerial vehicle intelligent identification method and system for shielding target

By segmenting and feature extraction of the multi-frame continuous monitoring images acquired by the drone, dynamic recognition results are generated and acquisition parameters are optimized, the problem of low accuracy of occlusion target recognition in traditional methods is solved, and efficient occlusion target recognition and dynamic change capture are achieved.

CN120259926AActive Publication Date: 2025-07-04DEYANG JINGKAI ZHIHANG TECH CO LTD

Patent Information

Application Number
CN202510732832.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

When facing the occlusion target, traditional drone target recognition methods are difficult to accurately identify and process occlusion information, resulting in low recognition accuracy and poor robustness, which cannot meet the actual application needs.

Method used

The acquisition module equipped by the drone acquires multi-frame continuous monitoring images, performs occlusion area segmentation processing, extracts visible and occlusion feature distributions, generates dynamic recognition results, and optimizes acquisition parameters in real time to improve recognition accuracy and robustness.

Benefits of technology

Accurate identification of occlusion targets and capture dynamic changes, improve the accuracy and robustness of the identification, and adapt to target monitoring in different states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259926A_ABST
    Figure CN120259926A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle intelligent identification method and system for a sheltered target, and the method comprises the steps: obtaining a multi-frame continuous monitoring image of a to-be-identified target which is at least partially sheltered in a target region in real time through a collection module carried by an unmanned aerial vehicle, and carrying out the sheltered region segmentation processing of the multi-frame continuous monitoring image, performing dynamic feature extraction processing on the basis of the target region segmentation result and the shielding region segmentation result to obtain visible feature distribution and shielding feature distribution of the target to be recognized in each frame of image; generating a dynamic recognition result representing the state transition track of the target to be recognized according to the visible feature distribution and the occlusion feature distribution, and finally optimizing the collection parameters of the collection module in real time based on the dynamic recognition result, and generating a parameter adjustment strategy of the next monitoring period, so that multi-frame image information can be fully utilized, the occlusion target can be accurately recognized, and the recognition efficiency is improved. And the recognition accuracy and robustness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicles, and in particular, to an intelligent recognition method and system for unmanned aerial vehicles for occluded targets. Background Art

[0002] In the field of unmanned aerial vehicle applications, the monitoring and recognition of target areas is a key task, especially when facing challenges in recognizing occluded targets. In actual scenarios, the target to be recognized may be partially or completely occluded for various reasons, such as occlusion by buildings, trees, or other objects, or partial invisibility due to the change of the target's own posture. Traditional target recognition methods often analyze single-frame images. When facing occluded targets, due to limited available effective information, it is difficult to accurately recognize the target. This method usually cannot make full use of the temporal correlation and dynamic information between multiple consecutive monitoring images, and cannot effectively distinguish the visible part and the occluded part of the target, resulting in low recognition accuracy and poor robustness, and it is difficult to meet the requirements of accurate recognition of occluded targets in practical applications. Summary of the Invention

[0003] In view of the above-mentioned problems, in combination with the first aspect of the present invention, embodiments of the present invention provide an intelligent recognition method for unmanned aerial vehicles for occluded targets, and the method includes: Real-time acquisition of multiple consecutive monitoring images of a target area by an acquisition module carried by an unmanned aerial vehicle, where the multiple consecutive monitoring images include at least a partially occluded target to be recognized; Performing occlusion area segmentation processing on the multiple consecutive monitoring images to obtain the target area segmentation result and the corresponding occlusion area segmentation result of the target to be recognized in each frame of the monitoring image; Performing dynamic feature extraction processing based on the target area segmentation result and the occlusion area segmentation result to obtain the visible feature distribution and the occlusion feature distribution of the target to be recognized in each frame of the monitoring image; Generating a dynamic recognition result of the target to be recognized according to the visible feature distribution and the occlusion feature distribution, where the dynamic recognition result is used to characterize the state transition trajectory of the target to be recognized in multiple consecutive monitoring images; Performing real-time optimization processing on the acquisition parameters of the acquisition module based on the dynamic recognition result to generate a parameter adjustment strategy for the next monitoring period.

[0004] In another aspect, an embodiment of the present invention further provides an intelligent UAV recognition system for occluded targets, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0005] Based on the above aspects, in the embodiment of the present invention, a multi-frame continuous monitoring image of a target area is obtained in real time through an acquisition module carried by a UAV. For at least some of the to-be-recognized targets included therein that are occluded, first, an occlusion area segmentation process is performed on the multi-frame continuous monitoring images to accurately obtain the target area segmentation result and the corresponding occlusion area segmentation result of the to-be-recognized target in each frame of the monitoring image. Then, based on the segmentation results, a dynamic feature extraction process is performed to obtain the visible feature distribution and occlusion feature distribution of the to-be-recognized target in each frame of the monitoring image, which comprehensively and meticulously depicts the feature information of the target in different states. A dynamic recognition result representing the state transition trajectory of the to-be-recognized target in the multi-frame continuous monitoring images is generated according to the visible feature distribution and occlusion feature distribution, realizing an accurate grasp of the dynamic changes of the occluded target. Finally, based on the dynamic recognition result, the acquisition parameters of the acquisition module are optimized in real time to generate a parameter adjustment strategy for the next monitoring period, which can adaptively adjust the acquisition parameters, improving the accuracy, robustness and real-time performance of the recognition of occluded targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1 is a schematic flowchart of the execution of the intelligent UAV recognition method for occluded targets provided by an embodiment of the present invention.

[0007] Figure 2 is a schematic diagram of an exemplary hardware and software component of the intelligent UAV recognition system for occluded targets provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0008] The present invention will be specifically described below with reference to the accompanying drawings of the specification. Figure 1 is a schematic flowchart of the intelligent UAV recognition method for occluded targets provided by an embodiment of the present invention. The intelligent UAV recognition method for occluded targets will be introduced in detail below.

[0009] Step S110: Obtain multi-frame continuous monitoring images of the target area in real time through an acquisition module carried by the UAV. The multi-frame continuous monitoring images include at least some of the to-be-recognized targets that are occluded.

[0010] In this embodiment, the set scenario is a large industrial park. The drone is equipped with an image acquisition module, which has the ability to collect image data at a specific frequency. The acquisition frequency of the image acquisition module is set to f. In the industrial park, there are many large devices, and some of the devices may be blocked by surrounding buildings, transport vehicles, etc. These blocked devices are the targets to be recognized. The acquisition module continuously acquires images of a specific area in the industrial park at frequency f. Each acquisition generates an image, and these images form a multi-frame continuous monitoring image in chronological order. In many of these images, the devices to be recognized are partially blocked. For example, at a certain moment, the acquisition module acquires an image in which a large processing device is partially blocked by a transport vehicle beside it.

[0011] Step S120: Perform occlusion area segmentation processing on the multi-frame continuous monitoring images to obtain the target area segmentation results and the corresponding occlusion area segmentation results of the target to be recognized in each frame of monitoring image.

[0012] After obtaining the multi-frame continuous monitoring images, in order to clearly define the target to be recognized and its occluded part, occlusion area segmentation processing needs to be performed.

[0013] Step S121: Perform time series alignment processing on the multi-frame continuous monitoring images to generate multi-frame aligned monitoring images that are continuous in the time dimension.

[0014] In the industrial park environment, when the drone flies, it will be affected by factors such as air flow and its own flight attitude adjustment, resulting in the multi-frame continuous monitoring images collected not being strictly continuous in the time dimension. Let the multi-frame continuous monitoring images be I1, I2, I3... I n , and the corresponding acquisition times are t1, t2, t3... t n . To ensure the accuracy of subsequent processing, these images need to be time series aligned.

[0015] Specifically, select areas with significant and stable features in the images as reference points. For example, the corner points of large landmark buildings at fixed positions in the industrial park, etc. With the help of an algorithm based on feature matching, this algorithm calculates the deviation amount of each frame of image relative to the standard time series by comparing information such as the position change and angle change of the reference points in different frames of images. For image I1, the deviation amount d1 relative to the standard time series is calculated through this algorithm. The calculation of the deviation amount comprehensively considers the coordinate change of the reference point in the image. For image I2, the deviation amount d2 is calculated, and so on, to obtain d3... d n . Then, according to these deviation amounts, corresponding geometric transformations such as translation and rotation are performed on each frame of image, so that the adjusted images I1', I2', I3'... I n"Form continuous and precisely aligned multi-frame aligned monitoring images in the time dimension, as if arranging these images closely in the accurate chronological order, laying a good foundation for subsequent processing.

[0016] Step S122: Call the pre-trained segmentation network model to perform frame-by-frame semantic segmentation processing on the multi-frame aligned monitoring images, and obtain the initial segmentation masks in each frame of the monitoring images.

[0017] After obtaining multi-frame aligned monitoring images with continuous time dimension, use the pre-trained segmentation network model to process them. Taking the common Mask-RCNN segmentation network model as an example, this model is mainly composed of a backbone network, a region proposal network (RPN), a region of interest (ROI) Align layer, a mask prediction branch, etc.

[0018] In the industrial park scenario, input one frame of the multi-frame aligned monitoring images, such as I1', into the backbone network of the Mask-RCNN model. The backbone network is generally composed of multiple convolutional layers and pooling layers, which are responsible for extracting features from the input image. The convolutional layers use different convolutional kernels K1, K2, K3... to perform convolutional operations on the image. The convolutional kernels slide on the image with a set stride, and each convolutional kernel is responsible for extracting specific types of features, such as edge features, texture features, etc. After multiple convolutional and pooling operations, feature maps F1, F2, F3... with different features are obtained.

[0019] Next, these feature maps enter the region proposal network (RPN). The RPN generates a series of anchor boxes on the feature map by sliding a window, and each anchor box corresponds to a potential target region in the image. The RPN classifies (judges whether it is a target) and regresses (predicts the position and size adjustment amount of the anchor box) each anchor box according to the feature map information, so as to generate a series of candidate regions that may contain targets.

[0020] Then, the candidate regions are processed by the region of interest (ROI) Align layer. The ROI Align layer maps candidate regions of different sizes to a feature map of a fixed size for subsequent unified feature extraction and prediction.

[0021] Finally, the features processed by the ROI Align layer enter the mask prediction branch. The mask prediction branch performs mask prediction on each candidate region through a series of convolutional layers and fully connected layers, outputs the mask information corresponding to the candidate region, and this mask information is processed to generate an initial segmentation mask with the same size as the input image. This mask preliminarily divides the approximate regions of different objects in the image. In the industrial park scenario, it can initially distinguish the regions where different objects such as equipment, buildings, and vehicles are located.

[0022] Step S123: Extract the contour region of the target to be recognized based on the initial segmentation mask as the candidate target region, and perform edge integrity verification processing on the candidate target region to obtain a verification result.

[0023] Extract the contour region of the target to be recognized from the initial segmentation mask of each frame of the monitoring image. In the industrial park scenario, assume that the region corresponding to a large device in the initial segmentation mask is M1. An algorithm based on contour tracking is used. This algorithm starts from an edge pixel point of the mask image and sequentially checks whether adjacent pixel points belong to the contour in a set direction (such as clockwise or counterclockwise). If an adjacent pixel point meets the set contour condition (such as the gray value difference exceeds the set threshold T1), it is added to the contour pixel point set. Through continuous tracking, the contour region of the target to be recognized (large device) is formed as the candidate target region C1.

[0024] Then perform edge integrity verification processing on the candidate target region C1.

[0025] Step S1231: Extract the set of edge pixel points of the candidate target region and calculate the curvature distribution characteristics of the set of edge pixel points.

[0026] Extract the set of edge pixel points E1 of the candidate target region C1. For each edge pixel point P(x0, y0) in the set E1, select its two adjacent pixel points P1(x1, y1) and P2(x2, y2). Calculate the curvature in the following way: According to the coordinates of these three points, use vector operations and trigonometric function relationships to calculate the angle θ between the vectors P1P0 and P0P2, and then combine the distance formula between two points to calculate the lengths L1 and L2 of P1P0 and P0P2. Calculate the curvature value K(x0, y0) of this point through the conventional curvature calculation formula in the related technology (using θ, L1, and L2), and perform this operation on all pixel points in the set E1, so as to obtain the curvature distribution characteristics of the set of edge pixel points E1.

[0027] Step S1232: Detect whether there are curvature mutation points in the candidate target region based on the curvature distribution characteristics, and the curvature mutation points are used to characterize the edge defect positions of the candidate target region.

[0028] Set a curvature mutation threshold T2. When the difference between the curvature value K(x0, y0) of a certain pixel point and the curvature value of its adjacent pixel point exceeds T2, this pixel point is recognized as a curvature mutation point, and these curvature mutation points are used to characterize the edge defect positions of the candidate target region C1.

[0029] Step S1233: When a curvature mutation point is detected, extract the local edge segment corresponding to the curvature mutation point, and calculate the motion compensation similarity score between the local edge segment and the edge segment at the corresponding position in the monitoring image of the adjacent frame in combination with the target motion estimation result.

[0030] When a curvature mutation point is detected, extract the local edge segment L1 corresponding to these curvature mutation points. In combination with the target motion estimation result, assume that the target motion estimation obtains the motion vector V of the target by analyzing the displacement of the same feature points in the adjacent frame images. According to the motion vector V, find the edge segment L2 at the corresponding position in the monitoring image of the adjacent frame. Calculate the motion compensation similarity score S between the local edge segment L1 and the edge segment L2 at the corresponding position in the monitoring image of the adjacent frame. The calculation method is as follows: sample L1 and L2 at equal intervals respectively to obtain the sampling point sets S1 and S2. For each sampling point in S1, find the point with the closest distance in S2, calculate the sum D of the distances between all corresponding points, and then subtract D from a fixed length value L (such as the length of L1 or L2). The ratio of the obtained difference to L is the motion compensation similarity score S.

[0031] Step S1234: Determine the degree of edge defect of the candidate target area according to the motion compensation similarity score. If the degree of edge defect exceeds the preset threshold, it is determined that the candidate target area needs to be processed by cross-frame contour compensation.

[0032] Set an edge defect threshold T3. If S is less than T3, it is determined that the degree of edge defect of the candidate target area exceeds the preset threshold, that is, it is determined that the candidate target area needs to be processed by cross-frame contour compensation.

[0033] Step S124: Perform cross-frame contour compensation processing on the candidate target area with defects according to the verification result to obtain the compensated target area segmentation result.

[0034] If the verification result of step S123 indicates that the candidate target area has an edge defect that needs to be compensated. In the industrial park scenario, assume that the candidate target area C1 (such as the contour area of large equipment) has an edge defect. Compensate using the information in the corresponding area of the adjacent frame monitoring image.

[0035] Let the candidate target area C1 with defects be I in the current frame image curr , the adjacent frame images be I prev and I nxt . First, based on the motion vector V obtained from the target motion estimation result, find the approximate areas C iprev and inxt corresponding to C1 in 1prev and C 1nxtFor each pixel point P(x, y) at the edge defect in C1, according to the motion vector V, find the corresponding pixel points P 1prev and P 1nxt in C prev (x prev , y prev ) and P nxt (x nxt , y nxt ).

[0036] Adopt an interpolation algorithm for contour compensation. For example, for the gray value of point P, by calculating the gray values of P prev and P nxt and their relative position relationship with point P, use the weighted average method to determine the compensated gray value of point P. Assume that the gray value of P prev is G prev , the gray value of P nxt is G nxt , the distances between P and P prev , P nxt are d1 and d2 respectively, and the compensated gray value G of point P = (G prev *d2 + G nxt *d1) / (d1 + d2). Perform similar operations on all pixel points at the edge defects in C1 to obtain the compensated target region segmentation result C1'.

[0037] Step S125: Determine the occluded region segmentation result based on the difference region between the compensated target region segmentation result and the initial segmentation mask.

[0038] In the industrial park scenario, compare the compensated target region segmentation result C1' with the initial segmentation mask M1.

[0039] Step S1251: Compare the compensated target region segmentation result with the initial segmentation mask pixel by pixel to obtain a set of difference pixel points.

[0040] Through pixel-by-pixel comparison, for each pixel point P(x, y), compare the gray value or class information of C1'(x, y) and M1(x, y) (if the mask is based on class annotation). If the two are different, mark this pixel point as a difference pixel point, and all difference pixel points form the set of difference pixel points D1.

[0041] Step S1252: Perform spatial clustering processing on the set of difference pixel points to generate multiple candidate occluded regions.

[0042] Perform spatial clustering on the set of differential pixel points D1. The DBSCAN clustering algorithm is used, which sets two parameters, the neighborhood radius ε and the minimum number of points MinPts. For each pixel point P(x, y) in the set D1, calculate its distance from other pixel points within the neighborhood (such as the Euclidean distance). If the number of pixel points within the radius ε is greater than or equal to MinPts, these pixel points are divided into one cluster. By processing all pixel points, multiple candidate occlusion regions O1, O2, O3... are generated.

[0043] Step S1253: Perform region screening based on the spatio-temporal continuity characteristics of the candidate occlusion regions in multiple consecutive monitoring images to obtain a stable segmentation result of the occlusion region.

[0044] For each candidate occlusion region, such as O1, track its position and shape changes in multiple consecutive monitoring images. Calculate the position change amounts Δx, Δy and shape change amounts (such as area change, aspect ratio change, etc.) in different frame images. If in multiple frames of images, the position and shape changes of the candidate occlusion region are within the set range, that is, the position change amounts Δx, Δy are less than the set position change threshold T x 、T y , and the shape change amount is less than the set shape change threshold T s , then it is considered that the candidate occlusion region has spatio-temporal continuity and is retained as the stable segmentation result O1' of the occlusion region.

[0045] Step S1254: Perform a single iteration optimization on the contour information of the target region segmentation result according to the stable segmentation result of the occlusion region. If the overlapping area between the optimized target region and the occlusion region is less than the preset threshold, stop the iteration and generate the optimized target region segmentation result.

[0046] Calculate the overlapping area A1 between C1' and O1'. Set an overlapping area threshold T4. If A1 is greater than T4, adjust the contour of C1'. The adjustment method is: for the pixel points on the contour of C1' that overlap with O1', re-determine the positions of these pixel points according to the edge information of O1' and the feature information inside C1' to reduce the overlapping area. After adjustment, calculate the overlapping area again. If the overlapping area is less than T4, stop the iteration and generate the optimized target region segmentation result C1''.

[0047] Step S130: Perform dynamic feature extraction based on the target region segmentation result and the occlusion region segmentation result to obtain the visible feature distribution and occlusion feature distribution of the target to be recognized in each frame of the monitoring image.

[0048] After obtaining the target region segmentation result C1'' and the occlusion region segmentation result O1', perform dynamic feature extraction processing on them to obtain the visible feature distribution and occlusion feature distribution of the target to be recognized in each frame of the monitoring image.

[0049] Step S131: Divide multiple analysis windows with different scales in the target region segmentation result, and calculate the gradient magnitude distribution within each analysis window.

[0050] In the target region segmentation result C1'', divide multiple analysis windows according to different scale requirements. For example, set the small-scale window size to s1×s1, the medium-scale window size to s2×s2, and the large-scale window size to s3×s3 (where s1 < s2 < s3). For each scale of window, slide the window on C1'' with a set step size. For each pixel point within the window, use a gradient-based algorithm to calculate the gradient magnitude. For the pixel point P(x, y), by calculating the gray-level changes in its horizontal and vertical directions, use the gradient calculation formula to obtain the gradient magnitude G(x, y) of this point. After calculating the gradient magnitudes for all pixel points within the window, obtain the gradient magnitude distribution within each analysis window.

[0051] Step S132: Perform histogram of oriented gradients (HOG) statistical processing on the gradient magnitude distribution to obtain the texture orientation histogram feature of each analysis window.

[0052] For the gradient magnitude distribution within each analysis window, perform histogram of oriented gradients (HOG) statistics. Divide the gradient directions into several intervals, for example, divide them into n intervals, and each interval corresponds to a direction range. For the gradient direction of each pixel point within the window, according to the direction range it belongs to, accumulate the gradient magnitude of this pixel point into the corresponding interval. After the statistics are completed, obtain the texture orientation histogram feature of each analysis window, which reflects the direction distribution of the texture within the window.

[0053] Step S133: Normalize and splice the texture orientation histogram features of different scales to generate the local texture feature of the target to be recognized.

[0054] Normalize the texture orientation histogram features obtained from windows of different scales. For the texture orientation histogram feature of each scale, calculate the sum Sum of all interval values, and then divide each interval value by Sum to obtain the normalized texture orientation histogram feature. Splice the normalized texture orientation histogram features of different scales together in scale order to generate the local texture feature of the target to be recognized, which synthesizes the texture information at different scales.

[0055] Step S134: Extract the minimum bounding rectangle of the target region segmentation result, perform bin coding processing on the aspect ratio and area change rate of the minimum bounding rectangle respectively, and generate the global morphological features of the target to be recognized.

[0056] Extract the minimum bounding rectangle of the target region segmentation result C1''. Let the length of the minimum bounding rectangle be l and the width be w, and calculate its aspect ratio r = l / w. At the same time, assume that the area of the minimum bounding rectangle corresponding to the target region in the previous frame of the image is A prev , and the area of the current frame is A curr , calculate the area change rate v=(A curr - A prev ) / A prev .

[0057] Perform bin coding processing on the aspect ratio r. Set several aspect ratio intervals, for example, r1 < r2 < … < r n , and according to the interval where r is located, encode it into the corresponding class value. Similarly, perform bin coding processing on the area change rate v, set several area change rate intervals, and encode it into the corresponding class value according to the interval where v is located. Combine these two encoded class values together to generate the global morphological features of the target to be recognized, which reflect the overall shape and area change of the target region.

[0058] Step S135: Perform region growing processing on the occlusion region segmentation result to generate an extended boundary of the occlusion region, and extract the geometric structure features of the extended boundary of the occlusion region.

[0059] Perform region growing processing on the occlusion region segmentation result O1'.

[0060] Step S1351: Select seed points in the occlusion region segmentation result, and set a region growing threshold based on the gray value distribution of the seed points.

[0061] In the occlusion region segmentation result O1', randomly select several pixel points as seed points. Analyze the gray value distribution of these seed points, calculate the mean μ and standard deviation σ of the gray values. Set the region growing threshold T5 = μ + k * σ (where k is an adjustable coefficient).

[0062] Step S1352: Perform growing processing from the seed points to the adjacent pixel regions until the gray difference between adjacent pixels exceeds the region growing threshold and then stop growing to obtain the initial growing region.

[0063] Centered on each seed point, growth is carried out towards its adjacent pixel regions. For each adjacent pixel point P(x, y), calculate the grayscale difference between it and the seed point. If the grayscale difference is less than or equal to T5, then add this pixel point to the growth region. Continuously repeat this process until the grayscale differences of all adjacent pixels exceed T5. At this time, the initial growth region G1 is obtained.

[0064] Step S1353: Perform morphological closing operation on the edge of the initial growth region to generate a smoothly connected extended boundary of the occlusion region.

[0065] Perform morphological closing operation on the edge of the initial growth region G1. Morphological closing operation generally consists of dilation operation and erosion operation. First, use a structuring element (such as a square structuring element) to perform dilation operation on the edge of the initial growth region G1. The dilation operation will expand the edge pixel points outward to fill some small holes and gaps. Then, perform erosion operation on the dilated edge. The erosion operation will contract the dilated edge inward to remove some redundant pixels generated by dilation, so that the edge becomes smoother. Through such morphological closing operation, a smoothly connected extended boundary B1 of the occlusion region is generated.

[0066] Step S1354: Extract the set of corner point coordinates of the extended boundary of the occlusion region, and calculate the Euclidean distance between adjacent corner points and the included angle parameter formed by three consecutive corner points.

[0067] For the extended boundary B1 of the occlusion region, adopt a corner point detection algorithm, such as the principle of the Harris corner point detection algorithm. It judges whether a pixel point is a corner point by calculating the autocorrelation matrix of each pixel point on the boundary and analyzing its eigenvalues. After detection, the set of corner point coordinates C = {(x1, y1), (x2, y2), …, (x n , y n )} is obtained.

[0068] For adjacent corner points in the corner point set C, such as (x i , y i ) and (x i+1 , y i+1 ), according to the Euclidean distance formula, that is, calculate the straight-line distance between two points in the plane rectangular coordinate system. Let the Euclidean distance be d, d = [(xᵢ₊1 - xᵢ)² + (yᵢ₊1 - yᵢ)²] 1 / 2 , and obtain the Euclidean distance between adjacent corner points.

[0069] For three consecutive corner points, such as (x i , y i ), (x i+1 , yi+1 ) and (x i+2 , y i+2 ), calculate the included angle parameter formed by them through vector operations. First, obtain the vector V1 = (x i+1 - x i , y i+1 - y i ) and the vector V2 = (x i+2 - x i+1 , y i+2 - y i+1 ) according to the coordinates. Then, using the vector dot product formula, let the included angle be θ, cosθ = (V 1x * V 2x + V 1y * V 2y ) / (|V1| * |V2|), and then obtain the included angle parameter θ.

[0070] Step S1355: Normalize the Euclidean distance and the included angle parameter respectively, and construct the multi-dimensional geometric feature vector of each corner point based on the normalized Euclidean distance and the included angle parameter.

[0071] For the set of Euclidean distances D = {d1, d2,..., d n}, let the maximum value among them be d max , and the minimum value be d min . Normalize each Euclidean distance d i , and let the normalized Euclidean distance be d i ', d i ' = (d i - d min ) / (d max - d min ).

[0072] For the set of included angle parameters Θ = {θ1, θ2,..., θ n}, let its maximum value be θ max , and the minimum value be θ min . Normalize each included angle parameter θ i , and let the normalized included angle parameter be θ i ', θ i ' = (θ i - θ min ) / (θ max - θ min ).

[0073] For each corner point, combine the normalized Euclidean distance d i ' and the included angle parameter θ i ' to construct the multi-dimensional geometric feature vector F of each corner point: F = {(d_1', θ_1'), (d2', θ2'), …, (d_ n ', θ_ n )}.

[0074] Step S1356: Perform density clustering on the geometric feature vector, divide it into multiple corner point clustering groups, perform least squares line fitting on each corner point clustering group to generate line segment primitives, and calculate the direction difference degree between adjacent line segment primitives. If the direction difference degree is less than the preset angle threshold, merge the adjacent line segments into a continuous broken line.

[0075] Adopt a density clustering algorithm, such as the DBSCAN algorithm, use the multi-dimensional geometric feature vector F as the input data, and set two parameters, the neighborhood radius ε and the minimum number of points minPts. For each geometric feature vector f i , calculate its distance (such as the Euclidean distance) from other vectors within the neighborhood. If the number of vectors within the radius ε is greater than or equal to minPts, divide these vectors into a clustering group. In this way, divide it into multiple corner point clustering groups G1, G2, …, G n .

[0076] For each corner point clustering group, such as G i , adopt the least squares line fitting method. Let the set of corner point coordinates in the corner point clustering group G i be {(x1, y1), (x2, y2), …, (x m , y m ). According to the least squares principle, find a straight line y = ax + b such that the sum of the squares of the distances from all corner points to this straight line is the smallest. Solve the relevant equations (construct and solve the equations according to the least squares principle) to obtain the parameters a and b of the straight line, thereby generating the line segment primitive L i .

[0077] For adjacent line segment primitives L i and L i+1 , calculate their direction difference degree. Let the slope of the straight line L i be k1, and the slope of the straight line L i+1 be k2. According to the relationship between the slope and the angle, let the inclination angle of the straight line L i be α1, tanα1 = k1, and the inclination angle of the straight line L i+1 be α2, tanα2 = k2. The direction difference degree β = |α1 - α2|. Set a preset angle threshold T angle . If β is less than T angle, then merge adjacent straight line segments L i and L i+1 into a continuous polyline P1.

[0078] Step S1357: Perform endpoint interpolation on the merged polyline to generate a closed polygon boundary, and optimize the vertex coordinates of the closed polygon to generate the geometric structure features of the extended boundary of the occlusion area.

[0079] For the merged continuous polyline P1, let its endpoints be (x start , y start ) and (x end , y end ). Using an interpolation algorithm, such as linear interpolation, insert several points between the endpoints so that the polyline forms a closed polygon boundary. Let the number of inserted points be n insert , for the j-th inserted point, its x coordinate is x j = x start + j * (x end - x start ) / (n insert+1 ), and the y coordinate is y j = y start + j * (y end - y start ) / (n insert+1 ), j = 1, 2,..., n insert .

[0080] After obtaining the closed polygon boundary, optimize its vertex coordinates. By analyzing the pixel information (such as features like gray value, gradient, etc.) inside and around the polygon, for each vertex coordinate (x i , y i ), adjust the vertex coordinate according to the features of the surrounding pixels. For example, if the gray value of the pixels on one side around the vertex changes significantly, it indicates that the vertex may need to move a certain distance in the direction where the gray value changes less. Through multiple iterative adjustments, make the polygon boundary more conform to the actual shape of the occlusion area, and finally generate the geometric structure features S1 of the extended boundary of the occlusion area.

[0081] Step S136: Construct the joint feature representation of the target to be recognized according to the normalized local texture features, the encoded global morphological feature vector, and the geometric structure features.

[0082] Denote the normalized local texture features as T, the encoded global morphological feature vector as M, and the geometric structure features as S1. First, ensure that the dimensions of these three features can match. If the local texture feature T is a multi-dimensional vector, the global morphological feature vector M is also a multi-dimensional vector, and the geometric structure feature S1 also has a set dimension. Assume the dimension of T is (t1, t2,..., tn ) The dimension of M is (m1, m2, …, m m ), and the dimension of S1 is (s1, s2, …, s k ).

[0083] To construct the joint feature representation, these three features are concatenated together in a set order. The local texture feature T can be placed first, followed by the concatenation of the global shape feature vector M, and finally the geometric structure feature S1, to obtain the joint feature representation F joint = [T, M, S1], whose dimension is (t1, t2, …, t n , m1, m2, …, m m , s1, s2, …, s k ). This joint feature representation synthesizes the feature information of the target to be recognized from multiple aspects such as local texture to global shape and the geometric structure of the occluded area.

[0084] Step S137: Based on the joint feature representation, perform a temporal modeling process on the feature change trend in multiple consecutive monitoring images to generate the visible feature distribution and the occluded feature distribution.

[0085] Perform a temporal modeling process on the joint feature representation in multiple consecutive monitoring images to generate the visible feature distribution and the occluded feature distribution.

[0086] Step S1371: Perform a time series alignment process on the joint feature representation in multiple consecutive monitoring images to generate a time series feature vector.

[0087] Let the joint feature representations corresponding to multiple consecutive monitoring images be F joint1 , F joint2 , …, F jointn . Since there is a sequence in the acquisition time of different frame images, and the joint feature representation may have slight differences in dimension or feature order due to small differences in image acquisition, a time series alignment process is required.

[0088] First, analyze the structure and dimension information of each joint feature representation F jointi . Assume that the dimension of F jointi is (d1, d2, …, d k ). For each dimension d j , find the corresponding dimension component in the joint feature representations of different frames. By comparing the change of the feature values of the corresponding dimension components in different frames, use the conventional feature matching algorithm in related technologies to adjust each joint feature representation so that they are consistent in the time dimension.

[0089] After adjustment, arrange these joint feature representations in chronological order to generate a time series feature vector Ftime =[F joint 1', F joint 2', …, F jointn ']), where F jointi ' is the joint feature representation after alignment processing, so that the time series feature vector F time is continuous in the time dimension and has a consistent feature dimension, providing a good foundation for subsequent time series analysis.

[0090] Step S1372: Input the time series feature vector into the input layer of a pre-trained time series analysis model, where each time step of the time series feature vector contains multiple feature channels, and each feature channel is respectively subjected to mean-variance normalization processing.

[0091] Input the time series feature vector F time into a pre-trained time series analysis model. Taking a common time series analysis model based on the recurrent neural network (RNN) architecture as an example, this time series analysis model includes an input layer, a hidden layer, and an output layer.

[0092] For each time step in the time series feature vector F time , that is, each F jointi ', it contains multiple feature channels, and these feature channels correspond to different dimensional information in the joint feature representation. For example, different dimensions of local texture features, different dimensions of global morphological feature vectors, and different dimensions of geometric structure features, etc.

[0093] For each feature channel, perform mean-variance normalization processing. Suppose the values of a certain feature channel in the time series feature vector are {v1, v2, …, v n}. First, calculate the mean μ = (v1 + v2 + … + v n ) / n, and then calculate the variance σ² = [(v1 - μ)² + (v2 - μ)² + … + (v n - μ)²] / / n. Then normalize each value v i . Suppose the normalized value is v i ', vᵢ' = (vᵢ - μ) / (σ²) 1 / 2Through such mean-variance normalization processing, the data of each feature channel has similar scales and distributions, which helps the model better learn the feature changes.

[0094] Input the time series feature vector after mean-variance normalization processing into the input layer of the time series analysis model. The input layer receives this data and passes it to the subsequent hidden layer for further processing.

[0095] Step S1373: Extract the forward time series features and backward time series features of the time series feature vector through the bidirectional long short-term memory network layer of the time series analysis model. The forward time series features and the backward time series features have the same dimension and are time step-aligned.

[0096] In the time series analysis model, the data passed by the input layer enters the bidirectional long short-term memory (Bi-LSTM) layer. The Bi-LSTM layer consists of a forward LSTM and a backward LSTM.

[0097] The forward LSTM processes the data starting from the starting time step of the time series feature vector F time and gradually learns the forward time series features in the data as the time step progresses. Let the hidden state of the forward LSTM at the i-th time step be h fi , which is calculated through the hidden state h f(i-1) of the previous time step and the input data x i at the current time step (i.e., the data after normalization at the time -th time step in the time series feature vector F i ).

[0098] The backward LSTM processes the data backward starting from the last time step of the time series feature vector F time and learns the backward time series features in the data as the time step goes back. Let the hidden state of the backward LSTM at the i-th time step be h bi , which is calculated through the hidden state h b(i+1) of the previous time step and the input data x i at the current time step.

[0099] Through such bidirectional processing, the forward LSTM and the backward LSTM respectively extract the forward time series features H f = {h f1 , h f2 , …, h fn} and the backward time series features H b = {h b1 , h b2 , …, h bn}。Due to the design of the Bi-LSTM layer, the dimensions of the forward temporal features and the backward temporal features are the same and the time steps are aligned, so that the front and back information of the time series data can be fully utilized to provide rich temporal features for subsequent feature fusion.

[0100] Step S1374: Concatenate the forward temporal features and the backward temporal features by time step and input them into the multi-head self-attention layer to calculate the attention weight matrix of different feature channels within each time step, and the dimension of the attention weight matrix matches the dimension of the concatenated temporal features.

[0101] Concatenate the forward temporal feature H f and the backward temporal feature H b by time step. For the i th time step, concatenate h fi and h bi together to obtain the concatenated temporal feature h combinedi . The concatenated temporal features of all time steps form the concatenated temporal feature vector H combined ={h combined1 , h combined2 , …, h combinedn}.

[0102] Input H combined into the multi-head self-attention layer. The multi-head self-attention layer calculates different attention weights in parallel through multiple heads (assumed to be n heads heads).

[0103] For each head, at the i-th time step, calculate the attention weights between different feature channels. Assume that the dimension of the concatenated temporal feature h combinedi is (d1, d2, …, d k ). Project h combinedi into three different vector spaces through linear transformation to obtain the query vector Q i , the key vector ki , and the value vector V i , and their dimensions are also (d1, d2, …, d k ).

[0104] Then, calculate the attention score matrix A i , and the element aᵢⱼ i of A = (Qᵢⱼ × Kᵢⱼ) / (d k ), where j represents the index of the feature channel. Pass the attention score matrix A 1 / 2 through the softmax function. iPerform normalization to obtain the attention weight matrix W i , W i The element w ij in it represents the attention weight of the j-th feature channel at the i-th time step.

[0105] Perform such calculations for each time step to obtain the set of attention weight matrices for different feature channels within each time step W = {W1, W2, …, W n}}, and the dimension of the attention weight matrix matches the dimension of the concatenated temporal features, that is, the same as the dimension of h combinedi . In this way, attention weights can be assigned according to the importance of different time steps and feature channels, highlighting key features.

[0106] Step S1375: Based on the attention weight matrix, perform weighted fusion on the concatenated temporal features to generate a fused temporal feature vector.

[0107] For the concatenated temporal feature vector H combined ={h combined1 , h combined2 , …, h combinedn} and the set of attention weight matrices W = {W1, W2, …, W n}, perform weighted fusion.

[0108] At the i -th time step, assume that the feature channel values of the concatenated temporal feature h combinedi are {v1, v2, …, v k}, and the elements of the attention weight matrix W i are {w1, w2, …, w k}. Weight each feature channel value to obtain the weighted feature channel value v i ' = w1*v1 + w2*v2 + … + w k *v k .

[0109] Perform such weighted operations for all time steps to obtain the fused temporal feature vector H fused ={h fused1 , h fused2 , …, h fusedn}, where h fusedi is the feature vector after weighted fusion at the i-th time step. Through this weighted fusion, the information of different time steps and feature channels is integrated, highlighting important temporal features and providing a more valuable feature representation for subsequent predictions.

[0110] Step S1376: Input the fused temporal feature vector into a fully-connected regression layer for multi-step prediction, and output the predicted visible feature vectors and occluded feature vectors for the next N time steps. The number of channels of the predicted visible feature vectors and occluded feature vectors is the same as that of the input feature vector.

[0111] Input the fused temporal feature vector H fused into the fully-connected regression layer. The fully-connected regression layer consists of multiple neurons, and each neuron is connected to all neurons in the previous layer (i.e., the layer where the fused temporal feature vector is located).

[0112] The fully-connected regression layer predicts the visible features and occluded features for the next N time steps by learning the feature information in the fused temporal feature vector. Let the weight matrix of the fully-connected regression layer be W f c and the bias vector be b f c.

[0113] For each time-step feature vector h fused in the fused temporal feature vector H fusedi , through matrix multiplication and addition operations, i.e., h fusedi ' = W fc * h fusedi + b fc , the prediction result is obtained. The prediction result includes the predicted visible feature vectors V pred = {v pred1 , v pred2 , …, v predn} and the predicted occluded feature vectors O pred = {o pred1 , o pred2 , …, o predn}.

[0114] The number of channels of the predicted visible feature vectors and occluded feature vectors is the same as that of the input feature vector (i.e., the feature vector h fused in the fused temporal feature vector H fusedi ), which can ensure the consistency of the prediction result in the feature dimension with the original features, so as to accurately analyze and process the visible feature distribution and occluded feature distribution in the subsequent steps.

[0115] Step S1377: Perform a first-order difference operation on the predicted visible feature vectors and occluded feature vectors respectively to generate the feature change amount at each time step relative to the previous moment, and construct the future change trends of the visible feature distribution and the occluded feature distribution according to the change amount sequences of consecutive time steps.

[0116] For the predicted visible feature vectors V pred = {v pred1, v pred2 , …, v predn}, perform the first-order difference operation. Let the feature vector at the t-th time step in the visible feature prediction vector be v predt , and the feature vector at the (t - 1)-th time step be v pred(t-1) . For each channel value vt predt (i represents the channel index) in v i and the corresponding channel value v pred(t-1) in v (t-1)i , calculate the first-order difference, that is, the feature change amount Δv ti = vt i - v (t-1)i . In this way, for each time step t (from 2 to N), a set of feature change amounts can be obtained, forming a sequence of feature change amounts ΔV = {Δv2, Δv3, …, Δv_ n}, where Δv t is a vector with the same number of channels as v predt , containing the feature change amount of each channel at this time step relative to the previous moment.

[0117] Similarly, for the occluded feature prediction vector O pred = {o pred1 , o pred2 , …, o predn}, perform the first-order difference operation. Let the feature vector at the t-th time step in the occluded feature prediction vector be o predt , and the feature vector at the (t - 1)-th time step be o pred(t-1) . For each channel value ot predt (i represents the channel index) in o i and the corresponding channel value o pred(t-1) in o (t-1)i , calculate the first-order difference, that is, the feature change amount Δo ti = ot i - o (t-1)i . Thus, a sequence of feature change amounts of the occluded feature prediction vector ΔO = {Δo2, Δo3, …, Δo_ n} is obtained, where Δo t is a vector with the same number of channels as o predt , containing the feature change amount of each channel at this time step relative to the previous moment.

[0118] Construct the future change trend of the visible feature distribution based on the sequence of feature change amounts ΔV of the visible feature prediction vector. Analyze the variation of the change amounts of each channel in the sequence of feature change amounts over time steps. For example, observe whether the change amount of a certain channel shows trends such as increasing, decreasing, or periodic changes at different time steps. Suppose there is a channel whose feature change amount gradually increases in several consecutive time steps, which may indicate an increasing trend of the corresponding visible feature in the future. Through the analysis of the sequences of feature change amounts of all channels, comprehensively obtain the future change trend of the visible feature distribution. This change trend can be represented as a multi-dimensional vector sequence T V , where each element corresponds to the description of the change trend at a time step. This description of the change trend can be a certain feature representation obtained based on the analysis of the feature change amounts of each channel. For example, it can be a comprehensive value obtained through operations such as performing a certain weighted combination or clustering analysis on the feature change amounts of each channel to reflect the change direction and degree of the entire visible feature distribution at this time step.

[0119] Similarly, construct the future change trend of the occluded feature distribution based on the sequence of feature change amounts ΔO of the occluded feature prediction vector. Conduct a similar analysis on the variation of the change amounts of each channel in the sequence of occluded feature change amounts over time steps. For example, if the change amount of a certain channel first decreases and then increases within several time steps, this reflects a specific change pattern of the corresponding occluded feature. Through the analysis of all channels, form the future change trend of the occluded feature distribution, which is represented as a multi-dimensional vector sequence T O , and each element is also a comprehensive value obtained based on the analysis of the feature change amounts of each channel, used to describe the change situation of the occluded feature distribution at the corresponding time step.

[0120] Step S1378: Perform compensation and correction processing on the visible feature distribution and the occluded feature distribution in the current frame monitoring image according to the future change trend, and generate the updated visible feature distribution and occluded feature distribution.

[0121] For the visible feature distribution in the current frame monitoring image, let it be V current , which is a multi-dimensional vector, and each dimension corresponds to a different visible feature channel. According to the future change trend T V of the visible feature distribution, perform compensation and correction on V current . Suppose the change trend vector at the t-th time step in T V is t Vt , which contains the description information of the future changes of each channel of the visible feature distribution.

[0122] For each channel value v current (i represents the channel index) in V i , according to t VtAdjust the information of the corresponding channels in. For example, if t Vt The information of a certain channel in indicates that the feature of this channel has an increasing trend in the future, then correspondingly increase V current The value of this channel in. The specific adjustment method can be through a certain weighting relationship. Let the weight be w i Then the adjusted channel value v i ' = v i + w i * t V t i where t V t i is the value of the corresponding channel in t Vt in. Perform such adjustments on all channels of V current to obtain the updated visible feature distribution V updated .

[0123] Similarly, for the occluded feature distribution in the current frame monitoring image, let it be O current , which is also a multi-dimensional vector. Compensate and correct according to the future change trend T O of the occluded feature distribution. Let the change trend vector at the t-th time step in T O be t ot , for each channel value o current in O i ( i represents the channel index), in a similar way to the visible feature distribution, according to the information of the corresponding channel in t ot and the corresponding weight w i ' for adjustment, that is, the adjusted channel value o i ' = o i + w i ' * t oti , where t oti is the value of the corresponding channel in t ot . Through the adjustment of all channels, generate the updated occluded feature distribution O updated . Through this compensation and correction process, the visible feature distribution and the occluded feature distribution in the current frame monitoring image can better reflect the future change situation predicted based on multi-frame analysis, providing more accurate feature information for subsequent target recognition and parameter optimization.

[0124] Step S140: Generate a dynamic recognition result of the target to be recognized according to the visible feature distribution and the occluded feature distribution, where the dynamic recognition result is used to characterize the state transition trajectory of the target to be recognized in multi-frame consecutive monitoring images.

[0125] After obtaining the updated visible feature distribution V updated and the occluded feature distribution O updatedAfter that, use these feature information to generate the dynamic recognition result of the target to be recognized.

[0126] First, fuse V updated and O updated Since they are both multi-dimensional vectors and their dimensions correspond to the previously processed feature channels, the two vectors are concatenated together in a set order to form a fused feature vector F fusion =[V updated , O updated . This fused feature vector synthesizes the information of the visible feature and the occluded feature.

[0127] Then, based on the fused feature vector F fusion to generate the dynamic recognition result. Use the conventional pattern recognition algorithm in related technologies to analyze the changes of the fused feature vector at different time steps. For example, observe the change pattern of each dimension value in the fused feature vector over time to judge the state change of the target to be recognized.

[0128] Suppose that at a certain time step, certain dimension values of the fused feature vector have a specific combination of changes. According to the pre-set rules or the model obtained through machine learning training, it can be judged that the target to be recognized has entered a new state. By analyzing the fused feature vector frame by frame in multiple consecutive monitoring images, record the state of the target to be recognized at different time steps, so as to generate the state transition trajectory of the target to be recognized in multiple consecutive monitoring images.

[0129] This state transition trajectory is the dynamic recognition result, which can be expressed as a state sequence S={s1, s2, …, s n}, where s n represents the state of the target to be recognized at the n th time step. Each state can be a certain feature description obtained based on the analysis of the fused feature vector. For example, through clustering analysis of the fused feature vector, different clustering results are defined as different states, and each state represents a comprehensive feature manifestation of the target to be recognized at that moment, which may include the comprehensive manifestation of information such as the position, pose, and occlusion degree of the target. In this way, a dynamic recognition result that can characterize the state transition of the target to be recognized in multiple consecutive monitoring images is generated.

[0130] Step S150: Based on the dynamic recognition result, perform real-time optimization processing on the acquisition parameters of the acquisition module to generate a parameter adjustment strategy for the next monitoring cycle.

[0131] After obtaining the dynamic recognition result of the target to be recognized, perform real-time optimization on the acquisition parameters of the acquisition module according to this result to generate a parameter adjustment strategy for the next monitoring cycle.

[0132] Step S151: Calculate the estimated motion speed value and the estimated motion direction value of the target to be recognized according to the state transition trajectory in the dynamic recognition result.

[0133] Analyze the state transition trajectory S = {s1, s2, …, s n} in the dynamic recognition result. Assume that each state s contains the position information of the target to be recognized at the corresponding time step. Let the position at the i-th time step be (x i , y i ), and the position at the (i + 1)-th time step be (x( i+1 ), y( i+1 )).

[0134] Calculate the estimated motion speed value. In the plane coordinate system, calculate the speed according to the position change and the time interval. Let the time interval be Δt (assuming that the time interval of each time step is the same). Then the velocity component vx in the x direction is (x( i+1 ) - x i ) / Δt, and the velocity component v y in the y direction is (y( i+1 ) - y i ) / Δt. Through these two velocity components, use the Pythagorean theorem to calculate the resultant velocity v = (v x ² + v y ²) 1 / 2 . This resultant velocity v is the estimated motion speed value of the target to be recognized.

[0135] Calculate the estimated motion direction value. Determine the motion direction according to the velocity components. Let the angle between the motion direction and the positive x-axis direction be θ. Then tanθ = v y / v x . Obtain the angle θ through the arctangent function arctan(v y / v x ). θ is the estimated motion direction value of the target to be recognized. Through such calculations, the estimated motion speed value and the estimated motion direction value of the target to be recognized are obtained from the state transition trajectory.

[0136] Step S152: Predict the expected position area of the target to be recognized in the next monitoring period based on the estimated motion speed value and the estimated motion direction value.

[0137] According to the calculated estimated motion speed value v and the estimated motion direction value θ, predict the expected position area of the target to be recognized in the next monitoring period.

[0138] Assume that the current position of the target to be recognized is (x0, y0), and the time length of the next monitoring period is T. In the x direction, calculate the position change Δx = v * cosθ * T according to the speed and time. In the y direction, the position change is Δy = v * sinθ * T.

[0139] Then, the expected position of the target to be recognized in the x direction in the next monitoring period is x1 = x0 + Δx, and the expected position in the y direction is y1 = y0 + Δy.

[0140] Taking (x1, y1) as the center, set the expected position area according to the set range. For example, set a circular area with a radius of r as the expected position area. All points (x, y) within this expected position area satisfy (x - x1)² + (y - y1)² ≤ r². The setting of the radius r here can be determined according to actual situations, such as the motion stability of the target to be recognized, the accuracy of the acquisition module, etc. Through such calculations and settings, the expected position area of the target to be recognized in the next monitoring period is obtained, providing a basis for adjusting the parameters of the acquisition module.

[0141] Step S153: Adjust the focal length parameter and shooting angle parameter of the acquisition module according to the expected position area, and generate the first parameter adjustment instruction.

[0142] According to the expected position area of the target to be recognized in the next monitoring period, adjust the focal length parameter and shooting angle parameter of the acquisition module to ensure that the target can be clearly photographed.

[0143] For the adjustment of the focal length parameter, analyze the imaging relationship between the expected position area and the current acquisition module. Assume that the imaging principle of the acquisition module follows certain optical imaging laws, and the distance between the expected position area and the acquisition module is d. If d is large, in order to make the target within the expected position area clearly imaged in the image, the focal length needs to be increased; conversely, if d is small, the focal length needs to be decreased. Let the current focal length be f0, and according to the relationship between the distance d and the focal length (for example, through a pre-established distance-focal length mapping relationship table or a calculation logic based on optical imaging principles), obtain the adjusted focal length f1.

[0144] For the adjustment of the shooting angle parameter, based on the current shooting direction of the acquisition module, calculate the amount of angle adjustment required to make the expected position area located at the center of the field of view of the acquisition module. Let the angle between the current shooting direction and the positive direction of the x-axis be α, and the angle between the direction of the center of the expected position area relative to the current position of the acquisition module and the positive direction of the x-axis be β. Then the shooting angle adjustment amount Δα = β - α. According to this angle adjustment amount, determine the direction and angle by which the acquisition module needs to rotate, so that the expected position area can be within the effective shooting field of view of the acquisition module.

[0145] Combine the adjusted focal length parameter f1 and the shooting angle adjustment amount Δα to generate a first parameter adjustment instruction I1, which is used to guide the acquisition module to adjust the focal length and shooting angle in the next monitoring cycle to better capture the target to be recognized.

[0146] Step S154: Predict the potential occlusion area in the next monitoring cycle based on the historical change trend of the occlusion feature distribution, and adjust the exposure parameter and resolution parameter of the acquisition module according to the potential occlusion area to generate a second parameter adjustment instruction.

[0147] Review the historical change trend of the occlusion feature distribution in the previous multiple consecutive monitoring images. Let the occlusion feature distributions in the previous several frames be O history ={O1, O2,..., O m}, and analyze the changes in these occlusion feature distributions.

[0148] Observe the change patterns of each part in the occlusion feature distribution. For example, whether the occlusion features in certain areas persist in several consecutive frames or appear and disappear regularly. Through the analysis of these historical data, predict the potential occlusion area where occlusion may occur in the next monitoring cycle.

[0149] Suppose it is found through analysis that a certain area has frequently shown occlusion in the past multiple frames and the change of its occlusion features has a certain stability, then regard this area as the potential occlusion area P in the next monitoring cycle.

[0150] For the potential occlusion area P, adjust the exposure parameter and resolution parameter of the acquisition module. If the brightness of the potential occlusion area P is low, in order to clearly display the details in this area, it is necessary to increase the exposure intensity of the acquisition module in this area. Let the current exposure intensity of the acquisition module for the entire image be E0, and calculate the increased exposure intensity ΔE required in this area according to the brightness situation of the potential occlusion area P (for example, determine the brightness by analyzing the gray value distribution of this area in the historical images), then the adjusted exposure intensity of the potential occlusion area P is E1 = E0 + ΔE, and at the same time, adjust the exposure intensity of other areas accordingly to ensure the exposure balance of the entire image.

[0151] For the adjustment of the resolution parameter, if the acquisition module supports region - adaptive resolution adjustment and the potential occlusion region P has a high importance (for example, the potential occlusion region P may contain key target information), then increase the image acquisition resolution of the potential occlusion region P. Let the overall resolution of the current acquisition module be R0, increase the resolution of the potential occlusion region P to R1, and at the same time adjust the resolution of the adjacent regions through the interpolation algorithm to ensure the continuity and consistency of the image.

[0152] Combine the adjusted exposure parameters (including the exposure intensities of the potential occlusion region P and other regions) and resolution parameters (the resolutions of the potential occlusion region P and adjacent regions) to generate the second parameter adjustment instruction I2. This second parameter adjustment instruction I2 is used to guide the acquisition module to optimize the exposure and resolution adjustments for the potential occlusion region in the next monitoring cycle.

[0153] Step S155: Combine the first parameter adjustment instruction and the second parameter adjustment instruction to generate the parameter adjustment strategy.

[0154] The first parameter adjustment instruction I1 contains the adjustment information of the focal length parameter and shooting angle parameter of the acquisition module, and the second parameter adjustment instruction I2 contains the adjustment information of the exposure parameter and resolution parameter of the acquisition module.

[0155] Fuse these two instructions to form a comprehensive parameter adjustment strategy P. The parameter information in I1 and I2 can be integrated in a set format. For example, first list the adjusted values of the focal length parameter and the adjustment amount of the shooting angle, and then list the adjusted values of the exposure parameter (including the potential occlusion region and other regions) and the adjusted values of the resolution parameter (the potential occlusion region and adjacent regions).

[0156] This parameter adjustment strategy P will be used as the basis for the acquisition module to adjust parameters in the next monitoring cycle, enabling the acquisition module to optimize the adjustment in terms of focal length, shooting angle, exposure, and resolution according to the dynamic recognition results of the target to be recognized, so as to improve the monitoring and recognition effects of the target.

[0157] Next, describe the model training part. When constructing the pre - trained segmentation network model, a network architecture suitable for semantic segmentation tasks can be selected, such as the UNet architecture. This architecture consists of an encoder, a decoder, and skip connections connecting the two.

[0158] The encoder part is composed of multiple down - sampling blocks. Each down - sampling block contains a convolutional layer and a pooling layer. The convolutional layer uses different convolutional kernels K1, K2, K3... to perform convolutional operations on the input image. The convolutional kernels slide on the image to extract different features in the image, such as edge features, texture features, etc. The pooling layer downsamples the convolutional feature map to reduce the data volume while retaining important features.

[0159] The decoder part corresponds to the encoder and consists of multiple upsampling blocks. The upsampling blocks restore the low-resolution feature maps to high resolution by means of deconvolution or interpolation, etc. It also contains convolutional layers for further fusing and refining the features.

[0160] The skip connections connect the feature maps at different levels in the encoder with the corresponding feature maps in the decoder, enabling the decoder to utilize the feature information at different levels in the encoder during the process of restoring the resolution, thereby improving the accuracy of segmentation.

[0161] Among them, a large number of image data containing the target to be recognized and occlusion situations are collected as the training set. This image data should cover different scenarios, different types of targets to be recognized, and various occlusion situations.

[0162] Each image is annotated to mark the area of the target to be recognized and the occlusion area. The annotation can be in the form of a pixel-level mask, that is, for each pixel in the image, mark whether it belongs to the target to be recognized, the occlusion area, or the background.

[0163] The annotated image data is divided into a training set, a validation set, and a test set. For example, it is divided according to a set ratio (such as 70% as the training set, 15% as the validation set, and 15% as the test set). The training set is used for parameter learning of the model, the validation set is used to adjust the hyperparameters of the model to prevent overfitting, and the test set is used to evaluate the final performance of the model.

[0164] Set the learning rate, which controls the step size of the model during each parameter update. Let the learning rate be lr, and usually a relatively small value, such as 0.001, is selected to ensure that the model can converge stably during training.

[0165] Set the number of training epochs, which represents the number of times the model makes a complete pass through the training set. For example, set epoch to 100, which means the model will perform 100 complete learning processes on the training set.

[0166] Set the batch size batch _size , that is, the number of images input into the model each time during training. For example, set batch _size to 16, indicating that 16 images are selected from the training set and input into the model for training at the same time.

[0167] The image data in the training set is sequentially input into the constructed segmentation network model. The image first enters the encoder and extracts features through convolutional and pooling operations.

[0168] In the decoder part, the feature map is restored to the same size as the input image through upsampling and convolution operations to generate the predicted segmentation mask. The predicted segmentation mask is compared with the annotated ground truth mask, and the loss function is used to measure the difference between the two.

[0169] Select a suitable loss function, such as the cross-entropy loss function. For each pixel point in the predicted mask, the difference between its predicted class probability and the ground truth class is quantified by the cross-entropy loss function. Suppose the predicted mask is P, the ground truth mask is G, for pixel point i, its predicted class probability is P i , and the ground truth class is G i , and the cross-entropy loss function calculates the loss value L of this pixel point i . By summing and averaging the loss values of all pixel points, the loss value L of the entire image is obtained.

[0170] According to the loss value L, an optimization algorithm is used to update the parameters of the model. Taking the stochastic gradient descent (SGD) algorithm as an example, it adjusts the parameters according to the gradient of the loss function with respect to the model parameters. Calculate the gradient ∇L / ∇θ of the loss function L with respect to each parameter θ in the model, and the parameter update formula is θ = θ - lr * ∇L / ∇θ, where lr is the set learning rate. By continuously iterating this process, the parameters of the model are gradually optimized and the loss value continuously decreases.

[0171] After each round of training (epoch), the image data of the validation set is input into the model, and the loss value on the validation set is calculated. Observe the change of the validation set loss value. If the validation set loss value no longer decreases or even starts to increase after several consecutive rounds of training, it indicates that the model may have overfitting. At this time, the learning rate can be adjusted, the model complexity can be reduced, etc. to avoid overfitting. For example, reduce the learning rate to one-tenth of the original, that is, lr = lr / 10, and then continue training.

[0172] After the set number of training epochs, the model is trained on the training set. The trained model is evaluated using the test set, and metrics such as the segmentation accuracy and recall rate of the model on the test set are calculated. For example, the segmentation accuracy is calculated as the ratio of the number of correctly segmented pixels to the total number of pixels, and the recall rate is calculated as the ratio of the number of correctly segmented target pixels to the actual number of target pixels. These metrics are used to determine whether the model meets the expected performance requirements. If not, the model structure or training parameters can be further adjusted and retrained.

[0173] Furthermore, a time series analysis model is constructed by combining a bidirectional long short-term memory network (Bi-LSTM) with a multi-head self-attention mechanism. The model mainly includes an input layer, a bidirectional long short-term memory network layer, a multi-head self-attention layer, and a fully connected regression layer.

[0174] The input layer is responsible for receiving the time series feature vector. This time series feature vector is a joint feature representation sequence after time series alignment processing and mean-variance normalization processing. Assume the time series feature vector is F time , and the feature vector dimension of each time step is d. The input layer inputs F time to the subsequent layers step by step according to the time steps.

[0175] The bidirectional long short-term memory network layer consists of a forward LSTM and a backward LSTM. The forward LSTM processes data starting from the beginning of the time series, and the backward LSTM processes data in reverse starting from the end of the time series. For the forward LSTM, at the t-th time step, it receives the hidden state h f(t-1) from the previous time step and the input feature x t at the current time step, and updates the hidden state h ft through internal mechanisms such as the forget gate, input gate, and output gate. For the backward LSTM, at the t-th time step, it receives the hidden state h b(t+1) from the next time step and the input feature x t at the current time step, and also updates the hidden state h b t through internal mechanisms. The forward LSTM and the backward LSTM respectively extract the forward time series feature H f and the backward time series feature H b , and these two time series features have the same dimension and are aligned in time steps.

[0176] The multi-head self-attention layer receives the feature vector H f formed by concatenating the forward time series feature H b and the backward time series feature H combined step by step. This layer calculates different attention weights in parallel through multiple heads (assumed to be n heads heads). For each head, at the t-th time step, the feature vector at the current time step in H combined is projected to the query vector Q t , key vector K t , and value vector V t through linear transformation. Then calculate the attention score matrix A t , and its element a ij is calculated according to the dot product of Q t and K t and the feature dimension, and then the softmax function is used to normalize A t to obtain the attention weight matrix W t , and the element w t of W ijIt represents the attention weight of the j-th feature channel at the t-th time step. The multi-head self-attention layer highlights the importance of different time steps and feature channels in this way.

[0177] The fully connected regression layer receives the fused temporal feature vector H after weighted fusion by the multi-head self-attention layer. fused . The fully connected regression layer consists of multiple neurons, and each neuron is connected to all elements in H. fused . Through the weight matrix W f c and the bias vector b f c, a linear transformation and bias addition operation is performed on H fused , that is, y = W f c * H fused + b f c, and the predicted visible feature vector and the predicted occluded feature vector for the next N time steps are output.

[0178] Furthermore, time series data is extracted from the joint feature representation of multi-frame consecutive monitoring images as the training set. Assume that the joint feature representation sequence is F jointsequence = {F joint1 , F joint2 , …, F jointm}, and it is divided into multiple time series segments in chronological order. Each time series segment contains several consecutive time steps, and assume that each segment contains T time steps. For example, the first time series segment is {F joint1 , F joint2 , …, F jointT}, the second time series segment is {F joint2 , F joint3 , …, F joint(T+1)}, and so on.

[0179] For each time series segment, according to the actual visible feature distribution and occluded feature distribution, the corresponding true visible feature prediction vector and true occluded feature prediction vector are determined, and this true vector is used as the labeled data for training.

[0180] All time series segments and their corresponding labeled data are divided into a training set, a validation set, and a test set. The division is also carried out according to a set ratio (such as 70% as the training set, 15% as the validation set, and 15% as the test set). The training set is used for parameter learning of the model, the validation set is used to adjust the hyperparameters of the model, and the test set is used to evaluate the performance of the model.

[0181] Set the learning rate lr seq , which controls the step size of the temporal analysis model in each parameter update. Similar to the learning rate setting of the segmentation network model, a relatively small value, such as 0.0001, is usually selected to ensure the stability and convergence of the model during training.

[0182] Set the number of training epochs seq , which represents the number of times the model makes a complete pass through the training set. For example, setting the epoch seq to 50 means the model will perform 50 complete learning processes on the training set.

[0183] Set the batch size _sizeseq , which is the number of time series segments input to the model during each training. For example, setting the batch _sizeseq to 8 means that 8 time series segments are selected from the training set and input to the model for training simultaneously each time.

[0184] Input the time series segments in the training set into the constructed time series analysis model in sequence. The time series segments first enter the input layer and then are passed to the bidirectional long short-term memory network layer.

[0185] In the bidirectional long short-term memory network layer, the forward LSTM and the backward LSTM process the time series segments respectively, extract the forward time series features and the backward time series features, and after concatenation, the forward and backward time series features enter the multi-head self-attention layer.

[0186] In the multi-head self-attention layer, the attention weights for different time steps and feature channels are calculated through multiple heads, and the concatenated time series features are weighted and fused to generate a fused time series feature vector.

[0187] The fused time series feature vector is input to the fully connected regression layer, and the fully connected regression layer outputs the visible feature prediction vector and the occlusion feature prediction vector for the next N time steps. The prediction vectors are compared with the true visible feature prediction vector and the true occlusion feature prediction vector, and the difference is measured through a loss function.

[0188] Select the mean squared error (MSE) loss function as an example. For the predicted visible feature prediction vector V pred and the true visible feature prediction vector V true , as well as the predicted occlusion feature prediction vector O pred and the true occlusion feature prediction vector O true , calculate the mean squared error between them respectively. For the visible feature prediction vector, the mean squared error MSE V =(1 / N)*∑(V predi -V truei ) 2 , where i ranges from 1 to N; for the occlusion feature prediction vector, the mean squared error MSE O =(1 / N)*∑(O predi -O truei ) 2, where i ranges from 1 to N. Add and average these two mean squared errors to obtain the loss value L of the entire model. seq .

[0189] According to the loss value L seq , use an optimization algorithm to update the parameters of the model. Similarly, select the Stochastic Gradient Descent (SGD) algorithm to calculate the gradient ∇L of the loss function L seq with respect to each parameter θ seq in the model, and the parameter update formula is θ seq = θ seq - lr seq * ∇L seq / ∇θ seq * ∇L seq / ∇θ seq .

[0190] After each round of training (epoch seq ), input the time series segments of the validation set into the model and calculate the loss value on the validation set. Observe the change of the validation set loss value. If the validation set loss value no longer decreases or even starts to increase after several consecutive rounds of training, it indicates that the model may be overfitting. At this time, ways such as adjusting the learning rate and reducing the model complexity can be used to avoid overfitting. For example, reduce the learning rate to half of the original, that is, lr seq = lr seq / 2, and then continue the training.

[0191] After the set number of training rounds epoch seq , the model completes training on the training set. Use the test set to evaluate the performance of the trained model and calculate metrics such as the prediction accuracy of the model on the test set. For example, for the visible feature prediction vector, calculate the error rate between the predicted value and the true value, and the error rate = (1 / N) * ∑|V predi - V truei | / |V truei |, where i ranges from 1 to N; similar calculations are also performed for the occluded feature prediction vector. Use these metrics to determine whether the model meets the expected performance requirements. If not, the model structure or training parameters can be further adjusted and retrained.

[0192] During the data collection process, when it comes to image collection that may contain privacy-sensitive data, differential privacy technology is used for privacy protection. Differential privacy perturbs the data by adding noise to the data, so that even if an attacker obtains part of the data, it is difficult to infer sensitive information from the data.

[0193] During data storage and transmission, encryption technology is adopted. For the collected image data and processed feature data, etc., a symmetric encryption algorithm such as AES (Advanced Encryption Standard) is used. A encryption key key is selected, and the data is encrypted according to the rules of the AES algorithm. When storing data, the encrypted data is stored in a secure storage medium. When transmitting data, the encrypted data is transmitted through a secure network channel. After receiving the encrypted data, the receiving party uses the same key key to decrypt it according to the decryption rules of the AES algorithm to restore the original data for subsequent processing. Through encryption technology, the security of privacy-sensitive data during storage and transmission is further ensured, preventing data from being stolen or tampered with.

[0194] Figure 2 FIG. shows a schematic diagram of exemplary hardware and software components of an unmanned aerial vehicle intelligent recognition system 100 for occluded targets provided by some embodiments of the present application that can implement the idea of the present application. For example, the processor 120 can be used on the unmanned aerial vehicle intelligent recognition system 100 for occluded targets and is used to execute the functions in the present application.

[0195] The unmanned aerial vehicle intelligent recognition system 100 for occluded targets can be a general-purpose server or a special-purpose server, both of which can be used to implement the method for unmanned aerial vehicle intelligent recognition of occluded targets in the present application. Although only one server is shown in the present application, for convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0196] For example, the unmanned aerial vehicle intelligent recognition system 100 for occluded targets can include a network port 110 connected to the network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROM, or RAM, or any combination thereof. Exemplarily, the unmanned aerial vehicle intelligent recognition system 100 for occluded targets can also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. According to these program instructions, the method of the present application can be implemented. The unmanned aerial vehicle intelligent recognition system 100 for occluded targets also includes an I / O interface 150 between the computer and other input / output devices.

[0197] For ease of description, only one processor is described in the UAV intelligent recognition system 100 for occluded targets. However, it should be noted that the UAV intelligent recognition system 100 for occluded targets in this application may also include multiple processors. Therefore, the steps performed by one processor described in this application can also be jointly performed or separately performed by multiple processors. For example, if the processor of the UAV intelligent recognition system 100 for occluded targets performs step A and step B, it should be understood that step A and step B can also be jointly performed by two different processors or separately performed in one processor. For example, the first processor performs step A, the second processor performs step B, or the first processor and the second processor jointly perform steps A and B.

[0198] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above-mentioned UAV intelligent recognition method for occluded targets is implemented.

[0199] It should be noted that, in order to simplify the description of the present invention disclosure and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are merged into one embodiment, drawing, or description thereof.

Claims

1. An intelligent recognition method for UAVs of occluded targets, characterized in that, The method includes: Obtaining multiple consecutive monitoring images of a target area in real time through an acquisition module carried by a drone, where the multiple consecutive monitoring images include at least a partially occluded target to be recognized; Performing occluded area segmentation processing on the multiple consecutive monitoring images to obtain the target area segmentation result and the corresponding occluded area segmentation result of the target to be recognized in each frame of the monitoring image; Performing dynamic feature extraction processing based on the target area segmentation result and the occluded area segmentation result to obtain the visible feature distribution and the occluded feature distribution of the target to be recognized in each frame of the monitoring image; Generating a dynamic recognition result of the target to be recognized according to the visible feature distribution and the occluded feature distribution, where the dynamic recognition result is used to characterize the state transition trajectory of the target to be recognized in the multiple consecutive monitoring images; Performing real-time optimization processing on the acquisition parameters of the acquisition module based on the dynamic recognition result to generate a parameter adjustment strategy for the next monitoring period.

2. The method for intelligent identification of an unmanned aerial vehicle for an occluded target according to claim 1, wherein The performing occluded area segmentation processing on the multiple consecutive monitoring images to obtain the target area segmentation result and the corresponding occluded area segmentation result of the target to be recognized in each frame of the monitoring image includes: Performing temporal alignment processing on the multiple consecutive monitoring images to generate multiple aligned monitoring images that are continuous in the time dimension; Invoking a pre-trained segmentation network model to perform frame-by-frame semantic segmentation processing on the multiple aligned monitoring images to obtain an initial segmentation mask in each frame of the monitoring image; Extracting the contour area of the target to be recognized based on the initial segmentation mask as a candidate target area, and performing edge integrity verification processing on the candidate target area to obtain a verification result; Performing cross-frame contour compensation processing on the candidate target area with defects according to the verification result to obtain a compensated target area segmentation result; Determining the occluded area segmentation result based on the difference area between the compensated target area segmentation result and the initial segmentation mask.

3. The method for intelligent recognition of an unmanned aerial vehicle for an occluded target according to claim 2, wherein The performing edge integrity verification processing on the candidate target area to obtain a verification result includes: Extracting the set of edge pixel points of the candidate target area and calculating the curvature distribution feature of the set of edge pixel points; Detecting whether there are curvature mutation points in the candidate target area based on the curvature distribution feature, where the curvature mutation points are used to characterize the edge defect positions of the candidate target area; When detecting the existence of curvature mutation points, extracting the local edge segment corresponding to the curvature mutation point, and calculating the motion compensation similarity score between the local edge segment and the edge segment at the corresponding position in the adjacent frame of the monitoring image in combination with the target motion estimation result; Determining the edge defect degree of the candidate target area according to the motion compensation similarity score, and if the edge defect degree exceeds a preset threshold, determining that the candidate target area needs to be subjected to cross-frame contour compensation processing.

4. The method for intelligent recognition of an unmanned aerial vehicle for an occluded target according to claim 2, wherein The determining the occluded area segmentation result based on the difference area between the compensated target area segmentation result and the initial segmentation mask includes: Performing pixel-by-pixel comparison between the compensated target area segmentation result and the initial segmentation mask to obtain a set of difference pixel points; Perform spatial clustering on the set of differential pixel points to generate multiple candidate occlusion regions; Based on the spatio-temporal continuity characteristics of the candidate occlusion regions in multiple consecutive monitoring images, perform region screening to obtain a stable occlusion region segmentation result; According to the stable occlusion region segmentation result, perform a single iteration optimization on the contour information of the target region segmentation result. If the overlapping area between the optimized target region and the occlusion region is less than a preset threshold, stop the iteration and generate an optimized target region segmentation result.

5. The method for intelligent identification of an unmanned aerial vehicle for an occluded target according to claim 1, wherein The dynamic feature extraction process based on the target region segmentation result and the occlusion region segmentation result to obtain the visible feature distribution and occlusion feature distribution of the target to be recognized in each frame of the monitoring image includes: Divide multiple analysis windows with different scales in the target region segmentation result and calculate the gradient magnitude distribution within each analysis window; Perform histogram of oriented gradients (HOG) statistical processing on the gradient magnitude distribution to obtain the texture orientation histogram features of each analysis window; Normalize and splice the texture orientation histogram features of different scales to generate the local texture features of the target to be recognized; Extract the minimum bounding rectangle of the target region segmentation result, and perform bin coding processing on the aspect ratio and area change rate of the minimum bounding rectangle respectively to generate the global morphological features of the target to be recognized; Perform region growing on the occlusion region segmentation result to generate an extended boundary of the occlusion region, and extract the geometric structure features of the extended boundary of the occlusion region; Construct a joint feature representation of the target to be recognized based on the normalized local texture features, the encoded global morphological feature vector, and the geometric structure features; Based on the joint feature representation, perform temporal modeling on the feature change trend in multiple consecutive monitoring images to generate the visible feature distribution and the occlusion feature distribution.

6. The method for intelligent identification of an unmanned aerial vehicle for an occluded target according to claim 5, wherein The process of performing region growing on the occlusion region segmentation result to generate an extended boundary of the occlusion region and extracting the geometric structure features of the extended boundary of the occlusion region includes: Select seed points in the occlusion region segmentation result and set a region growing threshold based on the gray value distribution of the seed points; Perform growth processing from the seed points to adjacent pixel regions until the gray difference between adjacent pixels exceeds the region growing threshold and then stop growing to obtain an initial growth region; Perform morphological closing operation on the edge of the initial growth region to generate a smoothly connected extended boundary of the occlusion region; Extract the set of corner coordinates of the extended boundary of the occlusion region, and calculate the Euclidean distance between adjacent corners and the included angle parameters formed by three consecutive corners; Normalize the Euclidean distance and the included angle parameters respectively, and construct a multi-dimensional geometric feature vector for each corner based on the normalized Euclidean distance and included angle parameters; Perform density clustering on the geometric feature vectors, divide them into multiple corner clustering groups, perform least squares line fitting on each corner clustering group to generate line segment primitives, and calculate the direction difference degree between adjacent line segment primitives. If the direction difference degree is less than a preset angle threshold, merge adjacent line segments into a continuous broken line; Endpoint interpolation is performed on the merged polyline to generate a closed polygon boundary, and the vertex coordinates of the closed polygon are optimized to generate the geometric structure features of the extended boundary of the occlusion area.

7. The method for intelligent identification of an unmanned aerial vehicle for an occluded target according to claim 5, wherein Performing temporal modeling on the feature change trends in multiple consecutive monitoring images based on the joint feature representation to generate the visible feature distribution and the occlusion feature distribution, including: Performing time series alignment on the joint feature representation in multiple consecutive monitoring images to generate time series feature vectors; Inputting the time series feature vectors into the input layer of a pre-trained temporal analysis model, where each time step of the time series feature vectors contains multiple feature channels, and each feature channel is respectively normalized by mean and variance; Extracting the forward temporal features and backward temporal features of the time series feature vectors through the bidirectional long short-term memory network layer of the temporal analysis model, where the forward temporal features and the backward temporal features have the same dimension and are time step-aligned; Inputting the forward temporal features and the backward temporal features into a multi-head self-attention layer after splicing by time step, and calculating the attention weight matrix of different feature channels within each time step, where the dimension of the attention weight matrix matches the dimension of the spliced temporal features; Performing weighted fusion on the spliced temporal features based on the attention weight matrix to generate a fused temporal feature vector; Inputting the fused temporal feature vector into a fully connected regression layer for multi-step prediction, and outputting the visible feature prediction vectors and occlusion feature prediction vectors for the next N time steps, where the number of channels of the visible feature prediction vectors and the occlusion feature prediction vectors is the same as that of the input feature vectors; Performing a first-order difference operation on the visible feature prediction vectors and the occlusion feature prediction vectors respectively to generate the feature change amount at each time step relative to the previous moment, and constructing the future change trends of the visible feature distribution and the occlusion feature distribution according to the change amount sequence of consecutive time steps; Performing compensation and correction processing on the visible feature distribution and the occlusion feature distribution in the current frame monitoring image according to the future change trends to generate the updated visible feature distribution and occlusion feature distribution.

8. The method for intelligent recognition of an unmanned aerial vehicle for an occluded target according to claim 1, wherein, Performing real-time optimization processing on the acquisition parameters of the acquisition module based on the dynamic recognition result to generate the parameter adjustment strategy for the next monitoring period, including: Calculating the estimated motion speed value and the estimated motion direction value of the target to be recognized according to the state transition trajectory in the dynamic recognition result; Predicting the expected position area of the target to be recognized in the next monitoring period based on the estimated motion speed value and the estimated motion direction value; Adjusting the focal length parameter and the shooting angle parameter of the acquisition module according to the expected position area to generate the first parameter adjustment instruction; Predicting the potential occlusion area in the next monitoring period based on the historical change trend of the occlusion feature distribution, and adjusting the exposure parameter and the resolution parameter of the acquisition module according to the potential occlusion area to generate the second parameter adjustment instruction; Fusing the first parameter adjustment instruction and the second parameter adjustment instruction to generate the parameter adjustment strategy.

9. The method for intelligent recognition of an unmanned aerial vehicle for an occluded target according to claim 8, wherein Adjusting the exposure parameters and resolution parameters of the acquisition module according to the potential occlusion area to generate a second parameter adjustment instruction includes: Calculating the occurrence frequency and duration of the potential occlusion area in the historical monitoring period; Determining the stability score of the potential occlusion area according to the occurrence frequency and the duration; When the stability score exceeds the first threshold and the acquisition module supports area adaptive resolution adjustment, reducing the image acquisition resolution corresponding to the potential occlusion area and enhancing the resolution of adjacent areas through an interpolation algorithm; When the stability score is lower than the second threshold, increasing the exposure intensity of the acquisition module corresponding to the potential occlusion area and reducing the exposure intensity of adjacent areas.

10. An intelligent UAV recognition system for occluded targets, characterized in that, It includes a processor and a memory. The memory is connected to the processor. The memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the method for intelligent identification of drones for occluded targets described in any one of claims 1-9 above.

Citation Information

Patent Citations

  • Unmanned aerial vehicle holder camera target identification tracking method and system based on artificial intelligence

    CN117197695A

  • High-robustness target tracking method based on unmanned aerial vehicle

    CN117853532A

  • Multi-modal remote sensing image sea target identification method based on AI

    CN119274062A

  • Dynamic target tracking method and system based on image return

    CN119310859A

  • Target identification method and system under view angle of unmanned aerial vehicle

    CN120014495A

Cited By

  • Business district consumer population dynamic prediction method and system based on unmanned aerial vehicle identification

    CN120451841A

  • Smart city information display method and system based on digital twinning

    CN120931873A

  • Unmanned aerial vehicle landing positioning method based on visual identification

    CN120953567A

  • Unmanned aerial vehicle road property inspection image recognition method and system based on semantic segmentation

    CN122024092A