A low-altitude unmanned aerial vehicle cross-field re-identification method fusing observation space prior
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UESTC (SHENZHEN) ADVANCED RES INST
- Filing Date
- 2026-07-13
- Publication Date
- 2026-08-07
AI Technical Summary
[0005]本发明的目的在于提供一种融合观测空间先验的低空无人机跨视场重识别方法,以解决现有技术中存在的低空无人机跨视场重识别方法单纯依赖视觉外观特征,忽略固定观测设备空间约束,在复杂低空场景中的身份连续性和关联准确率均较低,容易产生误关联或漏关联的技术问题
本申请将入口区域一致性纳入空间先验权重,使无人机首次进入下一相机视场的位置与固定相机部署关系相互印证;将初始空间先验校准为可学习空间先验
,并以先验保持损失约束其不过度偏离物理几何基础,使空间先验能够适应真实场景数据,但仍受固定相机几何关系约束;将可学习空间先验作为TransReID注意力模块中两幅图像之间的跨图像注意力偏置,使模型在视觉特征提取过程中即按空间合理性调制两图的特征级交互、在两两相似度形成之前感知跨设备转移合理性,而不是在视觉匹配之后简单叠加人工规则,从而提高了复杂低空场景中的无人机身份连续性和关联准确率,有效避免了误关联或漏关联。
Smart Images

Figure CN122530883A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned aerial vehicle (UAV) technology, and in particular to a cross-field-of-view re-identification method for low-altitude UAVs that integrates observation space priors. Background Technology
[0002] With the rapid development of the low-altitude economy, drones are increasingly being used on a large scale in logistics, power line inspection, emergency rescue, and security monitoring. The number and density of drone flights in urban low-altitude airspace are growing daily. Drone targets typically exhibit characteristics such as small size, rapid speed changes, low flight altitude, easy obstruction by buildings, and frequent disappearance and reappearance between different camera fields of view. To achieve continuous perception of low-altitude targets in key areas, it is usually necessary to deploy multiple fixed electro-optical cameras at different locations. These cameras work together to detect drones, perform short-term tracking, and re-identify them across fields of view. Cross-field-of-view re-identification is crucial for maintaining continuity of identity, determining whether a drone that leaves one camera's field of view and subsequently reappears in another camera's field of view is the same target.
[0003] Due to differences in camera installation location, orientation, field of view, observation distance, and background environment, the scale, attitude, sharpness, and direction of motion of the same UAV may vary significantly across different cameras. Existing cross-field-of-view re-identification methods for low-altitude UAVs mainly rely on target appearance similarity, intra-device tracking results, or simple temporal proximity rules. UAV targets in images are typically small in size, with inconspicuous differences in appearance texture and color. Under conditions of long-distance observation, backlighting, occlusion, or complex backgrounds, relying solely on visual appearance features and ignoring the spatial constraints of fixed observation equipment can easily lead to false or missed associations, resulting in low identity continuity and association accuracy in complex low-altitude scenes.
[0004] In the process of realizing this invention, the inventors discovered at least the following problems in the prior art: The cross-field-of-view re-identification method for low-altitude UAVs relies solely on visual appearance features and ignores the spatial constraints of fixed observation equipment. In complex low-altitude scenarios, the continuity of identity and the accuracy of association are both low, and it is easy to produce false associations or missed associations. Summary of the Invention
[0005] The purpose of this invention is to provide a cross-field-of-view re-identification method for low-altitude UAVs that integrates prior observations of the observation space. This addresses the technical problems of existing cross-field-of-view re-identification methods for low-altitude UAVs that rely solely on visual appearance features, ignore the spatial constraints of fixed observation equipment, and suffer from low identity continuity and association accuracy in complex low-altitude scenarios, easily leading to false or missed associations. The various technical effects of the preferred solutions among the many technical solutions provided by this invention are detailed below.
[0006] To achieve the above objectives, the present invention provides the following technical solution: This invention provides a low-altitude UAV cross-field-of-view re-identification method that integrates observation space priors, comprising the following steps: S100: Camera ,camera After successively detecting the low-altitude drones, short trajectory segments were generated. Short trajectory segments As candidate fragments S200: Camera-based Field of view, camera Three-dimensional field-of-view models were established for each field of view. 3D field of view model S300: Based on candidate fragment pairs and 3D field of view model 3D field of view model Get the camera ,camera The initial spatial prior weights are calculated by multiplying the spatial adjacency of the field of view, the shortest reach time on one side, and the entrance region consistency. S400: Candidate fragment pairs are processed via the TransReID encoder. The two images are jointly encoded to extract drone visual re-identification features that characterize the drone's appearance outline, local morphology, and cross-view identity information. Unmanned aerial vehicle (UAV) visual re-identification features S500: Sets the initial spatial prior as a learnable spatial prior and uses it as the cross-image attention bias and fusion decision factor between two images in the attention module of the TransReID encoder, outputting candidate segment pairs. The probability of association between drones belonging to the same type.
[0007] Preferably, in step S100, the short trajectory segment Short trajectory segments The expressions are as follows: , , in, Represents short trajectory segments The camera it belongs to Camera serial number, Represents short trajectory segments First time with a camera The time during which the field of view is observed, Represents short trajectory segments Leave the camera Field of view or the time of the last observation Represents short trajectory segments The drone image patch or image patch sequence obtained by cropping from the detection frame. Represents short trajectory segments The sequence of detection bounding boxes; Represents short trajectory segments The camera it belongs to Camera serial number, Represents short trajectory segments First time with a camera The time during which the field of view is observed, Represents short trajectory segments Leave the camera Field of view or the time of the last observation Represents short trajectory segments The drone image patch or image patch sequence obtained by cropping from the detection frame. Represents short trajectory segments The sequence of detection boxes.
[0008] Preferably, in step S200, the three-dimensional field-of-view model 3D field of view model The expressions are as follows: in, Indicates camera Any point in the low-altitude space of the field of view, Indicates camera Any point in the low-altitude space of the field of view, Representing three-dimensional real space, Indicates camera Position in a unified coordinate system Indicates camera Oriented towards the unit vector, and They represent cameras respectively. The imaging effect is acceptable at both the closest and furthest working distances. Indicates camera Half of the field of view, This represents the angle between two vectors; Indicates camera Position in a unified coordinate system Indicates camera Oriented towards the unit vector, and They represent cameras respectively. The imaging effect is acceptable at both the closest and furthest working distances. Indicates camera Half of the field of view.
[0009] Preferably, in step S300, the camera ,camera spatial adjacency between fields of view The expression is: , in, ( , () indicates camera ,camera The shortest spatial distance or engineering calibration distance between two fields of view. This is a spatial scale parameter used to control the impact of distance changes on the spatial adjacency of the field of view.
[0010] Preferably, in step S300, the camera ,camera The shortest reachable time on one side between them includes: the shortest reachable time on one side of the non-overlapping field of view. Shortest reachable time on one side of overlapping field of view The expressions for the two are as follows: , , , , in, Represents short trajectory segments Enter camera Time, Represents short trajectory segments Leave the camera Time, Represents short trajectory segments Short trajectory segments The time interval between events; This indicates the maximum reasonable flight speed that the drone can be set to. Indicates that the drone is from the camera Field of view shifted to camera The shortest reasonable time required for the field of view; For time scale parameters; if Greater than or equal to If so, the weight of the shortest reach time term on one side is not reduced. Less than If the candidate target appears too early, the weight of the shortest reach time term on one side will be reduced.
[0011] Preferably, in step S300, if the camera The camera can be determined in the image plane. Transfer to camera For a typical entry region, the entry region consistency weight... The expression is: , in, Indicates from camera Transfer to camera At that time, in the camera Entry region weighting function in the image plane, and Represents short trajectory segments First time entering the camera Image coordinates at time; if camera The camera cannot be determined in the image plane. Transfer to camera For a typical entry region, the entry region consistency weight... The value is set to 1.
[0012] Preferably, in step S400, the candidate fragment pairs are processed by the TransReID encoder. When jointly encoding the two images, the initial joint token sequence is fed into the TransReID encoder. The expression is: , in, , These represent the category tokens of the two images, used to aggregate their respective global identity features; to , to These represent the data obtained by dividing two drone image blocks. One patchtoken; This represents a positional encoding used to preserve the relative position of image patches within an image; The camera encoding is used to characterize the domain differences brought about by different cameras and to identify the corresponding camera of the image to which each token belongs. After joint encoding by the TransReID encoder, the outputs of the category tokens of the two images are taken as short trajectory segments. Visual re-identification features of drones Short trajectory segments Visual re-identification features of drones .
[0013] Preferably, in step S500, the formula for calculating the cross-image attention bias between two images in the attention module of the TransReID encoder is as follows: , , , , , , , , in, Indicates the dot product similarity between tokens within a joint sequence; Indicates each attention head and Feature dimensions; Used to normalize the dot product result; This represents the query matrix, used to represent the information actively sought by the current token. This represents the key matrix, used to represent the index features of the retrieved token; This represents the Value matrix, used to represent the aggregated content features; , , These represent the learnable linear mapping parameters; Represents the spatial prior bias matrix. This represents the feature matrix of the input token; This indicates the strength of the effect of spatial prior bias. Indicates prevention The operation is performed on extremely small positive numbers with unstable values, taking values only across image patches between two image patches and 0 within image patches; This represents the spatial prior weights after training and calibration. The parameter is A lightweight multilayer sensor. Represents the Sigmoid function; This represents the spatial prior feature vector of the candidate fragment pair, containing the initial spatial prior. Spatial adjacency Reasonableness of time Consistency of entrance area and normalized cross-field time interval , Represents short trajectory segments Short trajectory segments The time interval Indicates a reference time.
[0014] Preferably, candidate fragment pairs Association probability of belonging to the same drone The calculation formula is: , , in, This represents the Sigmoid function. Indicates candidate fragment pairs The weighting coefficients of visual appearance similarity, Indicates candidate fragment pairs Visual similarity Represents cosine similarity. Represents short trajectory segments The visual re-identification features of drones Represents short trajectory segments The visual re-identification features of drones, The weight coefficients represent the spatial prior weights after training and calibration. This represents the spatial prior weights after training and calibration. This indicates the bias term.
[0015] Preferably, the low-altitude UAV cross-field-of-view re-identification method uses manually confirmed cross-field-of-view segments as supervised samples for training, and the total loss function during the training phase is... The expression is: , , , , , , , in, This represents the identity association classification loss, used to directly supervise association probabilities. Does the candidate fragment pair express itself correctly? The probability that they belong to the same drone. =1 indicates a short trajectory segment With short trajectory segments Belonging to the same drone, =0 indicates a short trajectory segment With short trajectory segments They do not belong to the same drone; This represents a spatially weighted loss, used to make features of the same UAV more similar and features of different UAVs more distinct. This represents the negative sample interval threshold. The weights represent the spatially weighted loss measure. This represents the visual feature distance; the smaller the value, the more similar the visual features. Represents cosine similarity. Indicates the weight of the loss metric for positive samples. This indicates the strength of positive samples that are difficult to control. Indicates the weight of the loss metric for negative samples. Indicates the intensity of negative samples that are difficult to control; This represents the prior preservation loss, used to constrain the prior in the learnable space. However, it does not deviate excessively from the initial space prior. This allows spatial priors to remain interpretable while adapting to the distribution of real samples. This represents the weight of the prior preservation loss.
[0016] Implementing one of the above-described technical solutions of the present invention has the following advantages or beneficial effects: This application incorporates entrance area consistency into spatial prior weights, ensuring that the location of the UAV's first entry into the next camera's field of view corroborates the deployment relationship of the fixed cameras; and incorporates initial spatial priors. Calibration as a learnable space prior Furthermore, the prior-preserving loss is used to constrain it from deviating excessively from the physical geometry, enabling the spatial prior to adapt to real-world scene data while still being constrained by the fixed camera geometry. The learnable spatial prior is used as the cross-image attention bias between two images in the TransReID attention module, allowing the model to modulate the feature-level interaction of the two images according to spatial rationality during the visual feature extraction process and perceive the rationality of cross-device transfer before pairwise similarity is formed, rather than simply superimposing manual rules after visual matching. This improves the continuity and accuracy of drone identity association in complex low-altitude scenes and effectively avoids false or missed associations. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1This is a flowchart of a low-altitude UAV cross-field-of-view re-identification method that integrates observation space priors according to an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the consistency between the camera's three-dimensional field of view and the entrance region in a low-altitude UAV cross-field-of-view re-identification method that integrates observation space priors according to an embodiment of the present invention. Figure 3 This is a structural diagram of a low-altitude UAV cross-field-of-view re-identification method that integrates observation spatial priors in an embodiment of the present invention, in which learnable spatial priors are used as attention biases in TransReID. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the present invention clearer, various exemplary embodiments described below will be referenced to the accompanying drawings, which form part of the exemplary embodiments, illustrating various exemplary embodiments that may be used to implement the present invention. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. It should be understood that they are merely examples of processes, methods, and apparatuses consistent with some aspects of the present invention disclosed as detailed in the appended claims, and other embodiments may be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and spirit of the present invention.
[0019] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," etc., indicate the orientation or positional relationship based on the accompanying drawings, and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the referred element must have a specific orientation, or be constructed and operated in a specific orientation. The terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. The term "multiple" means two or more. The terms "connected" and "linked" should be interpreted broadly, for example, they can be fixed connections, detachable connections, integral connections, mechanical connections, electrical connections, communication connections, direct connections, indirect connections through an intermediate medium, and can be the internal connection of two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more of the related listed items. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0020] To illustrate the technical solution described in this invention, specific embodiments are described below, showing only the parts related to the embodiments of this invention.
[0021] Example: like Figure 1As shown, this invention provides a low-altitude UAV cross-field-of-view re-identification method that integrates observation space priors, including the following steps: S100: Camera ,camera After testing the low-altitude drone, the camera ,camera The system consists of two independent fixed optoelectronic cameras, which are optoelectronic imaging devices installed in a fixed position. The camera's position and orientation remain unchanged, and the range of its captured image is limited, covering only a fixed area. Therefore, cross-field-of-view re-identification by the UAV is required to generate short trajectory segments for each camera. Short trajectory segments As candidate fragments Short trajectory segments are short, complete trajectories of the same low-altitude UAV target within a short, continuous video frame, under conditions of no occlusion or stable detection. These segments consist of continuously observed UAV image patches, detection boxes, and timestamps within the same camera, providing the smallest comparable time unit for cross-device association. Candidate segment pairs are combinations of segments that may belong to the same low-altitude UAV target. Spatiotemporal rules can be used for coarse screening to filter out segments that are impossible to match, retaining only theoretically possible matches, significantly reducing computational load for post-processing. S200: Camera-based Field of view, camera Three-dimensional field-of-view models were established for each field of view. 3D field of view model Since the position and field of view of the fixed photoelectric camera directly determine whether the UAV can move from one camera to another, this embodiment models the camera observation space as a three-dimensional field of view model. The area that the camera can effectively observe in the real three-dimensional physical space is represented by the three-dimensional field of view model to facilitate subsequent calculations and processing. S300: Based on candidate fragment pairs and 3D field of view model 3D field of view model Get the camera ,camera Spatial adjacency between fields of view (quantization camera) ,camera The indicators of proximity, connectivity, and traversability between two 3D fields of view in global world space, typically taking values [0,1], and the shortest reachability time on one side (referring to the time from which a low-altitude UAV target can reach the camera). Starting from the boundary of the field of view, following a path in physical space, the path reaches the camera. The initial spatial prior weights are calculated by multiplying the field of view's spatial adjacency, the shortest reachability time on one side, and the entrance region consistency. ;Right now ,in, This represents the initial spatial prior determined by the camera's 3D field of view (corresponding to the spatial adjacency of the field of view), temporal order (corresponding to the shortest reach time on one side), and entrance location (corresponding to the consistency of the entrance region). This weight comes from device geometry and expert rules and is not directly used as an identity label, but rather as the physical initialization basis for subsequent learnable spatial prior calibration modules. S400: The candidate segment pairs are processed by a TransReID encoder (which uses the basic architecture of TransReID as the backbone network for visual re-identification, employs a pure Transformer architecture for target re-identification, utilizes a global self-attention mechanism to learn more discriminative feature representations, and achieves accurate cross-camera target association through global modeling capabilities and cross-camera robustness). The two images are jointly encoded, which facilitates the introduction of spatial priors to modulate the interaction between the two images within the backbone attention mechanism. This allows for the extraction of UAV visual re-identification features that characterize the UAV's appearance contour, local morphology, and cross-view identity information. Unmanned aerial vehicle (UAV) visual re-identification features Due to the use of joint coding, the visual re-identification features of drones Unmanned aerial vehicle (UAV) visual re-identification features The two images reinforce each other in cross-image attention. S500: The initial spatial prior is set as a learnable spatial prior and used as the cross-image attention bias and fusion decision factor between the two images in the attention module of the TransReID encoder, outputting candidate segment pairs. The probability of belonging to the same UAV. In this embodiment, the consistency of the entrance area is incorporated into the spatial prior weight, so that the position of the UAV when it first enters the next camera's field of view is mutually verified with the deployment relationship of the fixed cameras; the initial spatial prior is used to... Calibration as a learnable space prior Furthermore, the prior-preserving loss is used to constrain it from deviating excessively from the physical geometry, enabling the spatial prior to adapt to real-world scene data while still being constrained by the fixed camera geometry. The learnable spatial prior is used as the cross-image attention bias between two images in the TransReID attention module, allowing the model to modulate the feature-level interaction of the two images according to spatial rationality during the visual feature extraction process and perceive the rationality of cross-device transfer before pairwise similarity is formed, rather than simply superimposing manual rules after visual matching. This improves the continuity and accuracy of drone identity association in complex low-altitude scenes and effectively avoids false or missed associations.
[0022] As an optional implementation, in step S100, the short trajectory segment Short trajectory segments The expressions are as follows: , ,in, Represents short trajectory segments The camera it belongs to Camera serial number, Represents short trajectory segments First time with a camera The time during which the field of view (the spatial range that a camera can capture at a given moment) is observed. Represents short trajectory segments Leave the camera Field of view or the time of the last observation Represents short trajectory segments The drone image patch or image patch sequence obtained by cropping from the detection frame. Represents short trajectory segments The detection box sequence is a sequence of labeled bounding boxes output frame by frame by frame by the low-altitude UAV target detector in a continuous video frame, arranged in chronological order. Represents short trajectory segments The camera it belongs to Camera serial number, Represents short trajectory segments First time with a camera The time during which the field of view is observed, Represents short trajectory segments Leave the camera Field of view or the time of the last observation Represents short trajectory segments The drone image patch or image patch sequence obtained by cropping from the detection frame. Represents short trajectory segments The sequence of detection boxes.
[0023] As an optional implementation, in step S200, the camera's observable area is determined by its position, orientation, field of view angle, and effective observation distance. In this embodiment, this is described using a three-dimensional field of view model. Three-dimensional field of view model 3D field of view model The expressions are as follows: in, Indicates camera Any point in the low-altitude space of the field of view, Indicates camera Any point in the low-altitude space of the field of view, Representing three-dimensional real space, Indicates camera Position in a unified coordinate system Indicates camera Oriented towards the unit vector, and They represent cameras respectively. The acceptable minimum working distance (lower limit of effective observation distance) and the maximum working distance (upper limit of effective observation distance) for imaging effect. Indicates camera The field of view is half an angle; the camera's field of view is actually divided into a horizontal field of view and a vertical field of view, but this is a simplified representation. This represents the angle between two vectors; Indicates camera Position in a unified coordinate system Indicates camera Oriented towards the unit vector, and They represent cameras respectively. The imaging effect is acceptable at both the closest and furthest working distances. Indicates camera Half-angle of the field of view. A three-dimensional field-of-view model of the two cameras is constructed using the above formula, allowing the camera deployment position, orientation, field of view angle, and effective observation distance to directly participate in cross-field-of-view target association judgment.
[0024] As an optional implementation, in step S300, the camera ,camera spatial adjacency between fields of view The expression is: ,in, ( , () indicates camera ,camera The shortest spatial distance or engineering calibration distance between two fields of view. When the two fields of view overlap or have continuous coverage, this distance can be 0 or a small value. This is a spatial scale parameter used to control the impact of distance changes on the spatial adjacency of the field of view. The closer the value is to 1, the easier it is for a drone to perform cross-field-of-view transfer between the two camera fields of view.
[0025] As an optional implementation, in step S300, the camera ,camera The shortest reachable time on one side between them includes: the shortest reachable time on one side of the non-overlapping field of view. Shortest reachable time on one side of overlapping field of view Both are soft weights for time rationality, and their expressions are as follows: , , , , in, Represents short trajectory segments Enter camera Time, Represents short trajectory segments Leave the camera Time, Represents short trajectory segments Short trajectory segments In terms of time intervals, to avoid excessive exclusion of flight behaviors such as hovering, detouring, and turning back by the maximum time window, this embodiment only uses the shortest reachable time as a unilateral soft constraint. This indicates the maximum reasonable flight speed that the drone can be set, preset based on different low-altitude drones. Indicates that the drone is from the camera Field of view shifted to camera The shortest reasonable time required for the field of view; For time scale parameters; if Greater than or equal to If so, the weight of the shortest reach time term on one side is not reduced. Less than If this indicates that the candidate target appears too early, the weight of the shortest reach time term on one side should be reduced. When the fields of view of the two cameras overlap or have continuous coverage (i.e., ...), ... , At the same time, the same drone may appear in two fields of view almost simultaneously. The value might be small, or even negative due to the detection timing. In such cases, using a single-sided constraint would incorrectly suppress reasonable co-view candidates. Therefore, a double-sided soft window is used for overlapping field-of-view pairs. By using a single-sided shortest reach time soft constraint, the weight is reduced only for candidate relationships that appear too early, without hard-suppressing delays caused by hovering, detours, or backtracking. Furthermore, a double-sided time soft window is used for overlapping or continuously covered fields of view to avoid incorrectly suppressing reasonable co-view candidates.
[0026] As an optional implementation, in step S300, specifically, if the camera The camera can be determined in the image plane. Transfer to camera The typical entry area is determined based on equipment deployment relationships or historical observation experience. An entry area consistency weight is introduced, incorporating entry area consistency into the spatial prior weights. This ensures that the location where the UAV first enters the next camera's field of view is mutually verified with the fixed camera deployment relationship. Entry Area Consistency Weight The expression is: ,in, Indicates from camera Transfer to camera At that time, in the camera Entry region weighting function in the image plane, and Represents short trajectory segments First time entering the camera The image coordinates at that time; if the entrance area has not yet been marked at the construction site, this item is set to a neutral weight, which is a default weight value that neither favors matching nor mismatch in the association between two segments; if the camera... The camera cannot be determined in the image plane. Transfer to camera For a typical entry region, the entry region consistency weight... The value is set to 1.
[0027] As an optional implementation, in step S400, the candidate segment pairs are processed by the TransReID encoder. When jointly encoding the two images, the initial joint token sequence is fed into the TransReID encoder. The expression (i.e., the concatenated Transformer input sequence) is: ,in, , These represent the category tokens of the two images, used to aggregate their respective global identity features; to , to These represent the data obtained by dividing two drone image blocks. Each block is marked; This represents a positional encoding used to preserve the relative position of image patches within an image; The camera encoding is used to characterize the domain differences brought about by different cameras and to identify the corresponding camera of the image to which each token belongs. After joint encoding by the TransReID encoder, the outputs of the category tokens of the two images are taken as short trajectory segments. Visual re-identification features of drones Short trajectory segments Visual re-identification features of drones .
[0028] As an optional implementation, in step S500, the formula for calculating the cross-image attention bias between two images in the attention module of the TransReID encoder is as follows: , , , , , , , , in, Indicates the dot product similarity between tokens within a joint sequence; Indicates each attention head and Feature dimensions; Used to normalize the dot product result; This represents the query matrix, used to represent the information actively sought by the current token. This represents the key matrix, used to represent the index features of the retrieved token; This represents the Value matrix, used to represent the aggregated content features; , , These represent the learnable linear mapping parameters; Represents the spatial prior bias matrix. This represents the feature matrix of the input token; This indicates the strength of the effect of spatial prior bias. Indicates prevention The operation is performed on extremely small positive numbers with unstable values, taking values only across image patches between two image patches and 0 within image patches; This represents the spatial prior weights after training and calibration. The parameter is A lightweight multilayer sensor. The sigmoid function represents the initial space prior. through Calibration as a learnable space prior The prior preservation loss is used to constrain it from deviating excessively from the physical geometry, so that the spatial prior can adapt to real scene data, but is still constrained by the fixed camera geometry. This represents the spatial prior feature vector of the candidate fragment pair, containing the initial spatial prior. Spatial adjacency Reasonableness of time Consistency of entrance area and normalized cross-field time interval , Represents short trajectory segments Short trajectory segments The time interval This indicates the reference time. In this embodiment, the learnable spatial prior is used as the cross-image attention bias between two images in the TransReID attention module. This allows the model to modulate the feature-level interaction of the two images according to spatial rationality during visual feature extraction, perceiving the rationality of cross-device transfer before pairwise similarity is formed, rather than simply superimposing manual rules after visual matching. Due to the bias... It only applies across image blocks. For any query token, the bias applied to the token in the same image is different from that applied to the token in another image in its attention row. This row is non-uniformly distributed along the key dimension, so it is not canceled out by the translation invariance of softmax along the key dimension. The spatial prior can be effectively applied to the attention weights. The larger the value, the less cross-image attention between the two images is suppressed, and the stronger the feature-level interaction between the two images within the backbone. The smaller the value, the more the interaction between the two images is suppressed, thus introducing interpretable spatial prior modulation during the visual feature extraction process.
[0029] As an optional implementation, candidate fragment pairs Association probability of belonging to the same drone The calculation formula is: , ,in, This represents the Sigmoid function. Indicates candidate fragment pairs The weighting coefficients of visual appearance similarity, Indicates candidate fragment pairs Visual similarity Represents cosine similarity. Represents short trajectory segments The visual re-identification features of drones Represents short trajectory segments The visual re-identification features of drones, The weight coefficients represent the spatial prior weights after training and calibration. This represents the spatial prior weights after training and calibration. This indicates the bias term.
[0030] As an optional implementation, the low-altitude UAV cross-field-of-view re-identification method uses manually confirmed cross-field-of-view segments (only a small number of samples are needed) as supervised samples for training. The total loss function during the training phase... The expression is: , , , , , , , in, The identity association classification loss is used to train the fusion discrimination results, which are then used to directly supervise the association probability. Does the candidate fragment pair express itself correctly? The probability that they belong to the same drone. =1 indicates a short trajectory segment With short trajectory segments Belonging to the same drone, =0 indicates a short trajectory segment With short trajectory segments They do not belong to the same drone. This represents a spatially weighted loss, used to make features of the same UAV more similar and features of different UAVs more distinct. This represents the negative sample interval threshold. The weights represent the spatially weighted loss measure. To represent the visual feature distance, and to further optimize the visual Re-ID feature space, candidate segment pairs are defined based on the UAV visual re-identification features output by the TransReID backbone. The feature distance, the smaller the value, the more similar the visual features. Represents cosine similarity. Indicates the weight of the loss metric for positive samples. This indicates the strength of positive samples that are difficult to control. Indicates the weight of the loss metric for negative samples. This indicates the strength of the negative samples that are difficult to control; for positive samples, their weights are negatively correlated with the spatial prior: when the spatial prior weights are calibrated after training... When the spatial prior is too low (i.e., for difficult real positive samples such as hovering, maneuvering, and turning back where the spatial prior is not significant), the closing intensity is actually increased, thus avoiding the suppression of complex maneuvering targets; for negative samples, the weight is positively correlated with the spatial prior: when the spatial prior weights after training and calibration are... For difficult negative samples with higher spatial priors (i.e., spatially plausible but distinct identities), they should be pushed further away. In this embodiment, differential spatial weighting is used for positive and negative samples in metric learning, so that difficult negative samples with plausible spatial priors but distinct identities are pushed away more strongly, while ensuring that true positive samples (complex maneuvering targets) with low spatial priors are not suppressed. To avoid the learnable spatial prior deviating completely from the physical geometry of the fixed photoelectric camera, a prior preservation loss is introduced. This represents the prior preservation loss, used to constrain the prior in the learnable space. However, it does not deviate excessively from the initial space prior. This allows the spatial prior to remain interpretable while adapting to the distribution of real samples. This is achieved through identity-associated classification loss. Spatial weighted loss Spatial weighted loss We jointly optimize visual features, learnable spatial priors, and associated probability outputs to prevent spatial priors from being misused as identity labels and to prevent learnable priors from deviating too much from the physical geometric foundation.
[0031] The embodiment is merely a specific example and does not indicate that this is the only way to implement the present invention.
[0032] The above description is merely a preferred embodiment of the present invention. Those skilled in the art will understand that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.
Claims
1. A cross-field-of-view re-identification method for low-altitude unmanned aerial vehicles (UAVs) that integrates observation space priors, characterized in that, Includes the following steps: S100: Camera ,camera After successively detecting the low-altitude drones, short trajectory segments were generated. Short trajectory segments As candidate fragments ; S200: Camera-based Field of view, camera Three-dimensional field-of-view models were established for each field of view. 3D field of view model S300: Based on candidate fragment pairs and 3D field of view model 3D field of view model Get the camera ,camera The initial spatial prior weights are calculated by multiplying the spatial adjacency of the field of view, the shortest reach time on one side, and the entrance region consistency. ; S400: Candidate fragment pairs are processed via the TransReID encoder. The two images are jointly encoded to extract drone visual re-identification features that characterize the drone's appearance outline, local morphology, and cross-view identity information. Unmanned aerial vehicle (UAV) visual re-identification features ; S500: Sets the initial spatial prior as a learnable spatial prior and uses it as the cross-image attention bias and fusion decision factor between two images in the attention module of the TransReID encoder, outputting candidate segment pairs. The probability of association between drones belonging to the same type.
2. The low-altitude UAV cross-field-of-view re-identification method based on fusion of spatial prior observations as described in claim 1, characterized in that, In step S100, short trajectory segments Short trajectory segments The expressions are as follows: , , in, Represents short trajectory segments The camera it belongs to Camera serial number, Represents short trajectory segments First time with a camera The time during which the field of view is observed, Represents short trajectory segments Leave the camera Field of view or the time of the last observation Represents short trajectory segments The drone image patch or image patch sequence obtained by cropping from the detection frame. Represents short trajectory segments The sequence of detection bounding boxes; Represents short trajectory segments The camera it belongs to Camera serial number, Represents short trajectory segments First time with a camera The time during which the field of view is observed, Represents short trajectory segments Leave the camera Field of view or the time of the last observation Represents short trajectory segments The drone image patch or image patch sequence obtained by cropping from the detection frame. Represents short trajectory segments The sequence of detection boxes.
3. The low-altitude UAV cross-field-of-view re-identification method based on fusion of spatial prior observations as described in claim 1, characterized in that, In step S200, the three-dimensional field-of-view model 3D field of view model The expressions are as follows: in, Indicates camera Any point in the low-altitude space of the field of view, Indicates camera Any point in the low-altitude space of the field of view, Representing three-dimensional real space, Indicates camera Position in a unified coordinate system Indicates camera Oriented towards the unit vector, and They represent cameras respectively. The imaging effect is acceptable at both the closest and furthest working distances. Indicates camera Half of the field of view, This represents the angle between two vectors; Indicates camera Position in a unified coordinate system Indicates camera Oriented towards the unit vector, and They represent cameras respectively. The imaging effect is acceptable at both the closest and furthest working distances. Indicates camera Half of the field of view.
4. The low-altitude UAV cross-field-of-view re-identification method based on fusion of spatial prior observations as described in claim 1, characterized in that, In the S300 procedure, the camera ,camera spatial adjacency between fields of view The expression is: , in, ( , () indicates camera ,camera The shortest spatial distance or engineering calibration distance between two fields of view. This is a spatial scale parameter used to control the impact of distance changes on the spatial adjacency of the field of view.
5. The low-altitude UAV cross-field-of-view re-identification method according to claim 1, characterized in that, In the S300 procedure, the camera ,camera The shortest reachable time on one side between them includes: the shortest reachable time on one side of the non-overlapping field of view. Shortest reachable time on one side of overlapping field of view The expressions for the two are as follows: , , , , in, Represents short trajectory segments Enter camera Time, Represents short trajectory segments Leave the camera Time, Represents short trajectory segments Short trajectory segments The time interval between events; This indicates the maximum reasonable flight speed that the drone can be set to. Indicates that the drone is from the camera Field of view shifted to camera The shortest reasonable time required for the field of view; For time scale parameters; if Greater than or equal to If so, the weight of the shortest reach time term on one side is not reduced. Less than If the candidate target appears too early, the weight of the shortest reach time term on one side will be reduced.
6. The low-altitude UAV cross-field-of-view re-identification method according to claim 1, characterized in that, In the S300 procedure, if the camera The camera can be determined in the image plane. Transfer to camera For a typical entry region, the entry region consistency weight... The expression is: , in, Indicates from camera Transfer to camera At that time, in the camera Entry region weighting function in the image plane, and Represents short trajectory segments First time entering the camera Image coordinates at time; If camera The camera cannot be determined in the image plane. Transfer to camera For a typical entry region, the entry region consistency weight... The value is set to 1.
7. The low-altitude UAV cross-field-of-view re-identification method according to claim 1, characterized in that, In step S400, the candidate fragment pairs are processed by the TransReID encoder. When jointly encoding the two images, the initial joint token sequence is fed into the TransReID encoder. The expression is: , in, , These represent the category tokens of the two images, used to aggregate their respective global identity features; to , to These represent the data obtained by dividing two drone image blocks. One patchtoken; This represents a positional encoding used to preserve the relative position of image patches within an image; This represents the camera code, used to characterize the domain differences brought about by different cameras, and to identify the corresponding camera of the image to which each token belongs; After joint encoding by the TransReID encoder, the output of the category tokens from the two images is taken as short trajectory segments. Visual re-identification features of drones Short trajectory segments Visual re-identification features of drones .
8. The low-altitude UAV cross-field-of-view re-identification method according to claim 1, characterized in that, In step S500, the formula for calculating the cross-image attention bias between two images in the attention module of the TransReID encoder is as follows: , , , , , , , , in, Indicates the dot product similarity between tokens within a joint sequence; Indicates each attention head and Feature dimensions; Used to normalize the dot product result; This represents the query matrix, used to represent the information actively sought by the current token. This represents the key matrix, used to represent the index features of the retrieved token; This represents the Value matrix, used to represent the aggregated content features; , , These represent the learnable linear mapping parameters; Represents the spatial prior bias matrix. This represents the feature matrix of the input token; This indicates the strength of the effect of spatial prior bias. Indicates prevention The operation is performed on extremely small positive numbers with unstable values, taking values only across image patches between two image patches and 0 within image patches; This represents the spatial prior weights after training and calibration. The parameter is A lightweight multilayer sensor. Represents the Sigmoid function; This represents the spatial prior feature vector of the candidate fragment pair, containing the initial spatial prior. Spatial adjacency Reasonableness of time Consistency of entrance area and normalized cross-field time interval , Represents short trajectory segments Short trajectory segments The time interval Indicates a reference time.
9. The low-altitude UAV cross-field-of-view re-identification method based on fusion of spatial prior observations as described in claim 8, characterized in that, Candidate fragment pairs Association probability of belonging to the same drone The calculation formula is: , , in, This represents the Sigmoid function. Indicates candidate fragment pairs The weighting coefficients of visual appearance similarity, Indicates candidate fragment pairs Visual similarity Represents cosine similarity. Represents short trajectory segments The visual re-identification features of drones Represents short trajectory segments The visual re-identification features of drones, The weight coefficients represent the spatial prior weights after training and calibration. This represents the spatial prior weights after training and calibration. This indicates the bias term.
10. The low-altitude UAV cross-field-of-view re-identification method according to claim 9, characterized in that, The low-altitude UAV cross-field-of-view re-identification method uses manually verified cross-field-of-view segments as supervised samples for training. The total loss function during the training phase is... The expression is: , , , , , , , in, This represents the identity association classification loss, used to directly supervise association probabilities. Does the candidate fragment pair express itself correctly? The probability that they belong to the same drone. =1 indicates a short trajectory segment With short trajectory segments Belonging to the same drone, =0 indicates a short trajectory segment With short trajectory segments They do not belong to the same drone; This represents a spatially weighted loss, used to make features of the same UAV more similar and features of different UAVs more distinct. This represents the negative sample interval threshold. The weights represent the spatially weighted loss measure. This represents the visual feature distance; the smaller the value, the more similar the visual features. Represents cosine similarity. Indicates the weight of the loss metric for positive samples. This indicates the strength of positive samples that are difficult to control. Indicates the weight of the loss metric for negative samples. Indicates the intensity of negative samples that are difficult to control; This represents the prior preservation loss, used to constrain the prior in the learnable space. However, it does not deviate excessively from the initial space prior. This allows spatial priors to remain interpretable while adapting to the distribution of real samples. This represents the weight of the prior preservation loss.