Spatial Target Association Method for Large Field-of-View Cameras Based on Spatio-Temporal Higher-Order Attribute Hypergraphs

Through the method based on high-order attribute hypergraph of space-time, space-time, space-time, high-order attribute hypergraphs, the problem of traditional technology being unable to effectively deal with multi-type targets is solved, and the precise target association and classification of large-field detection data is achieved.

CN119888207BActive Publication Date: 2025-06-13NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510369335.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-13
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

When traditional large field of view cameras deal with diversified space-time dynamic changes, they are unable to effectively adapt to the correlation needs of multiple types of goals, resulting in poor correlation performance.

Method used

The method based on high-order attribute hypergraph of space-time is adopted, and the static attributes of the image plane is obtained through the static attribute extraction unit, the attribute hypergraph and inference hypergraph are constructed, and the spatial-time attribute association module and the spatial-time inference model are used to achieve the association and classification of the goals.

Benefits of technology

It realizes the precise correlation and classification of goals of different types and states, adapts to the feature extraction and association requirements of multi-target groups, and improves the target correlation tracking effect of large-field detection data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888207B_ABST
    Figure CN119888207B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for spatial target association of a large field of view camera based on a spatio-temporal high-order attribute hypergraph, including: obtaining a large field of view image sequence by using the large field of view camera, and extracting a static attribute set corresponding to each frame of the large field of view image; constructing an attribute hypergraph based on the first three frames of the large field of view images, determining a potential target association matching matrix corresponding to each attribute hypergraph through a spatio-temporal attribute association module, and determining the high-order spatio-temporal attributes of the third frame of the large field of view image; for each frame of the large field of view image after the third frame, constructing an inference hypergraph based on the high-order spatio-temporal attributes of the previous frame of the large field of view image, and determining the high-order spatio-temporal attributes of the current frame of the large field of view image; using the inference hypergraph as an input, implementing the prediction of the potential target association matching matrix between adjacent frames of the large field of view images by using a spatio-temporal inference model, and implementing the continuous association and classification of the targets by using the potential target association matching matrix; the present invention can better adapt to the feature extraction and association requirements of multi-target groups.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a method for associating space targets of a large field-of-view camera based on a spatio-temporal high-order attribute hypergraph. Background Art

[0002] A large field-of-view camera has a wider observation range and stronger detection ability, and can often observe space targets on different orbits simultaneously, and then present different positions and morphological changes on the image plane. When using a large field-of-view camera to observe near-earth objects, affected by factors such as the position and attitude of the observation platform and the observed object, different targets often show different types and degrees of position and morphological changes in the imaging data. For example, a large field-of-view camera located in LEO (Low Earth Orbit) can often capture space objects with positions distributed from LEO to HEO (High Earth Orbit). Among them, HEO objects often show slow linear position movement and morphological changes in the observation data, while LEO objects closer to the observation platform show stronger non-linearity and time-variation. Traditional single-type processing strategies often cannot process these targets with different states simultaneously, resulting in poor association performance under large field-of-view conditions.

[0003] Reference 1 introduced time dynamic information on the basis of a traditional space static graph, and proposed a spatio-temporal graph neural network by combining a graph shift operator and a time shift operator; this method can process the topological structures of time and space simultaneously and has good generalization ability in dealing with diverse spatio-temporal dynamic changes. However, this method requires that the states of the same node in the spatio-temporal graph are known at different times and the time shift operator only operates on the same node, and cannot meet the application requirements of multi-type target association; Reference 1:

[0004] “Hadou S, Kanatsoulis C I, Ribeiro A. Space-time graph neuralnetworks[J]. arXiv preprint arXiv:2110.02880, 2021.”. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for associating space targets of a large field-of-view camera based on a spatio-temporal high-order attribute hypergraph to better adapt to the feature extraction and association requirements of multi-target groups.

[0006] To achieve the above task, the present invention adopts the following technical solutions:

[0007] A method for associating space targets of a large field-of-view camera based on a spatio-temporal high-order attribute hypergraph, including:

[0008] A large-field-of-view camera is used to obtain a large-field-of-view image sequence. For each frame of the large-field-of-view image in the large-field-of-view image sequence, a static property extraction unit is used to separately extract the image-plane static properties of each target, so as to obtain a static property set corresponding to each frame of the large-field-of-view image;

[0009] Pairwise combine the first three frames of the large-field-of-view images in order, and construct an attribute hypergraph for each combination based on the static property set; use the attribute hypergraphs corresponding to the first three frames of the large-field-of-view images, and determine the potential target association matching matrix corresponding to each attribute hypergraph through a spatio-temporal property correlation module; based on the static property set of the first three frames of the large-field-of-view images and the potential target association matching matrix, determine the high-order spatio-temporal properties of the third frame of the large-field-of-view image;

[0010] For each frame of the large-field-of-view image after the third frame, construct an inference hypergraph with the high-order spatio-temporal properties of the previous frame of the large-field-of-view image of the current frame, and determine the high-order spatio-temporal properties of the current frame of the large-field-of-view image; use the inference hypergraph as the input, use a spatio-temporal inference model to predict the potential target association matching matrix of adjacent frames of the large-field-of-view image, and use the potential target association matching matrix to realize the continuous association and classification of targets.

[0011] Furthermore, the image-plane static properties of each target include the target position, morphological features, and background information; wherein the morphological features include the maximum gray value of the area where the target is located, the standard deviation matrix and covariance matrix of the target offset Gaussian model; the background information includes the background gray value mean and standard deviation of the target neighborhood.

[0012] Furthermore, the first and second layers of the static property extraction unit are convolutional layers, the third to tenth layers are all C3 modules, and the last two layers are convolutional layers; the feature maps output by the fifth, seventh, ninth, and twelfth layers form a feature pyramid, which is regressed by a regression head, and the regression head outputs the image-plane static properties, so as to form a static property set corresponding to each frame of the large-field-of-view image; an activation function is set after each layer of the static property extraction unit;

[0013] The C3 module includes three convolutional layers connected in sequence. Among them, after the output feature of the first convolutional layer and the output feature of the third convolutional layer are superimposed, they are feature-fused with the output feature of the first convolutional layer after passing through another convolutional layer on the branch, and the fusion result is processed by another convolutional layer to obtain the feature map output by the C3 module.

[0014] Furthermore, the regression head adopts a decoupled design, including a shared convolutional layer and parallel location regression branch, shape regression branch, and confidence branch. The location regression branch, shape regression branch, and confidence branch all adopt the Conv(SiLU(BN(Conv(·)))) structure. Among them, the location regression branch is used to obtain the target location, the shape regression branch is used to obtain the shape features and background information of the target, and the confidence branch is used to predict the confidence of the target's existence, so as to determine whether it is a real target. After the feature map output by the feature pyramid passes through the shared convolutional layer, it is processed by the cooperation of the location regression branch, shape regression branch, and confidence branch, and finally the static attributes of the target on the image plane are obtained.

[0015] Furthermore, pair the first three large-field-of-view images in sequence, and construct an attribute hypergraph for each combination based on the static attribute set, including:

[0016] Construct the attribute hypergraph corresponding to the first large-field-of-view image and the second large-field-of-view image and the attribute hypergraph corresponding to the second large-field-of-view image and the third large-field-of-view image ; The attribute hypergraph is represented by the following formula:

[0017] ;

[0018] In the above formula, represents the attribute hypergraph constructed with the targets of the i-th large-field-of-view image and the (i + 1)-th large-field-of-view image as nodes; are the static attribute sets corresponding to the i-th large-field-of-view image and the (i + 1)-th large-field-of-view image respectively; are the sets of edges connecting different nodes in the i-th large-field-of-view image and the (i + 1)-th large-field-of-view image respectively, then represents the set of edges connecting the nodes of the i-th large-field-of-view image and the (i + 1)-th large-field-of-view image; Define thresholds and , for two nodes belonging to the same large-field-of-view image, if the Euclidean distance is less than , then an edge is established between the two nodes; for two nodes belonging to different large-field-of-view images, if the Euclidean distance is less than , then an edge is established between the two nodes.

[0019] Furthermore, determine the potential target association matching matrix corresponding to each attribute hypergraph through the spatio-temporal attribute association module, including:

[0020] The input of the spatiotemporal attribute association module is the attribute hypergraph. The spatiotemporal attribute association module includes a six-layer network structure, in which the first, third and fifth layers use spatial hypergraph convolution, and the second, fourth and sixth layers use temporal hypergraph convolution. An activation function is set after each layer of the network structure, and an association head is set after the sixth layer. The association head uses a dot product attention layer, which can be expressed as:

[0021] ;

[0022] in, is the potential target association matching matrix corresponding to the attribute hypergraph constructed using the i-th frame large field of view image and the i+1-th frame large field of view image, which contains the association relationship between each target in the two frames of large field of view images; and They represent the feature matrices of the targets in the i-th frame large field of view image and the i+1-th frame large field of view image, respectively. is the normalization factor, and the superscript T indicates transposition;

[0023] The structures of spatial hypergraph convolution and temporal hypergraph convolution are as follows:

[0024] ;

[0025] Where X represents the input of spatial hypergraph convolution or temporal hypergraph convolution, Y is the processed output, A is the adjacency matrix, which stores the edge connection and weight information, and W is the trainable parameter. is the activation function.

[0026] Furthermore, based on the static attribute set of the first three frames of large field of view images and the potential target association matching matrix, the high-order spatiotemporal attributes of the third frame of large field of view image are determined, including:

[0027] For each target in the third frame of the large field of view image, difference and second-order difference operations are performed based on the static image properties of the target at different times to obtain the first-order differential and second-order differential of the target; finally, the static image properties, first-order differential and second-order differential of the target are stacked to obtain its high-order spatiotemporal properties. The high-order spatiotemporal properties of all targets constitute the high-order spatiotemporal properties of the third frame of the large field of view image.

[0028] Furthermore, a reasoning hypergraph is constructed based on the high-order spatiotemporal attributes of the previous large field of view image of the current frame, and the high-order spatiotemporal attributes of the current large field of view image are determined, including:

[0029] The structure of the inference hypergraph is the same as that of the attribute hypergraph. The difference between the inference hypergraph and the attribute hypergraph in the construction method is the edge set connecting the nodes of the i-th frame large field of view image and the i+1-th frame large field of view image. The construction process is different; in the reasoning hypergraph middle, The construction method is as follows:

[0030] First, for the current large field of view image with i > 2, when constructing the inference hypergraph a hypothesized association relationship is constructed between the j-th node in the i-th large field of view image and any node in the (i + 1)-th large field of view image, that is, it is assumed that the two are the same potential target; second, calculate the high-order spatio-temporal attributes of the j-th node in the i-th large field of view image as its hypothesized high-order spatio-temporal attributes; finally, define the hyperparameter and judge the hypothesized high-order spatio-temporal attributes of the j-th node in the i-th large field of view image, and the hypothesized high-order spatio-temporal attributes of the node in the (i + 1)-th large field of view image, and determine whether the Euclidean distance between them is less than the hyperparameter ; if it is less, then take the hypothesized association relationship as the association relationship between the j-th node in the i-th large field of view image between the i-th large field of view image and the (i + 1)-th large field of view image, and connect an edge between the j-th node and the node; take the hypothesized high-order spatio-temporal attributes as the high-order spatio-temporal attributes of the j-th node, so as to obtain the high-order spatio-temporal attributes of the i-th large field of view image.

[0031] Furthermore, taking the inference hypergraph as the input, use the spatio-temporal inference model to realize the prediction of the potential target association matching matrix of adjacent large field of view images, including:

[0032] The spatio-temporal inference model is a graph convolutional neural network. The first layer, the fourth layer, and the seventh layer of the network structure of the spatio-temporal inference model are spatial hypergraph convolutions, the second layer, the fifth layer, and the eighth layer are temporal hypergraph convolutions, the third layer, the sixth layer, and the ninth layer are temporal sequence attention layers. An activation function is set after each layer of the network structure, and an association head and a classification head are set after the ninth layer; among them, the association head adopts a dot product attention layer, and the classification head is implemented by a softmax function;

[0033] Input the inference hypergraph into the spatio-temporal inference model, and finally obtain the potential target association matching matrix and the category of the target through the combination of the association head and the classification head, so as to realize the continuous association and classification of the target; the temporal sequence attention layer adopts a graph convolutional long short-term memory network.

[0034] A terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the large field of view camera spatial target association method based on the spatio-temporal high-order attribute hypergraph.

[0035] A computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the method for associating space targets of a large field-of-view camera based on a spatio-temporal high-order attribute hypergraph is implemented.

[0036] Compared with the prior art, the present invention has the following technical features:

[0037] Aiming at the problems that the motion patterns, image plane forms and other factors are diverse due to the various types of observed targets of a large field-of-view camera, which makes it difficult for traditional processing methods to effectively and comprehensively handle, and at the same time aiming at the problem that traditional spatio-temporal graph networks cannot match and associate unknown target groups, the present invention proposes a novel attribute representation and reasoning network architecture. The present invention proposes a target association strategy that fully explores deep spatio-temporal semantics, provides a new idea for coping with the tracking and association tasks of a large number of and multi-type targets, proposes a more comprehensive and sufficient high-order attribute representation, and at the same time constructs a spatio-temporal neighborhood hypergraph structure, and fully mines spatio-temporal information through spatio-temporal hypergraph convolution and attention mechanisms in space and time, realizing more effective target association tracking of large field-of-view detection data. Description of the Drawings

[0038] Figure 1 It is a schematic diagram of the principle of the method of the present invention;

[0039] Figure 2 It is a network structure diagram of the static attribute extraction unit in the present invention;

[0040] Figure 3 It is a comparison diagram of the experimental effects of an embodiment of the present invention and other algorithms. Detailed Embodiment

[0041] The present invention provides a method for associating space targets of a large field-of-view camera based on a spatio-temporal high-order attribute hypergraph. Given a sequence of continuous detection images as input, a high-order attribute representation of potential moving targets (light spots) is constructed to embed hypergraph nodes, a spatio-temporal hypergraph is constructed using spatio-temporal relationships, high-order spatio-temporal semantic information is extracted through a spatio-temporal hypergraph neural network and nested reasoning of multi-dimensional attributes is performed, so as to associate targets of different types and different states. The present invention realizes more precise feature guidance by constructing a high-order attribute representation, and at the same time introduces a spatio-temporal hypergraph structure, which is more suitable for the feature extraction and association requirements of a multi-target group in position while giving full play to the dynamic information processing ability of the spatio-temporal graph neural network. The specific implementation steps of the present invention are as follows:

[0042] Step 1, use a large field-of-view camera to obtain a sequence of large field-of-view images, and for each frame of large field-of-view image in the sequence of large field-of-view images, use a static attribute extraction unit to respectively extract the static attributes of each target on the image plane, so as to obtain a static attribute set corresponding to each frame of large field-of-view image.

[0043] Given a sequence of k consecutive large field-of-view images , where represents the i-th large field-of-view image; construct a static attribute extraction unit , and use the static attribute extraction unit to separately perform feature extraction to obtain the corresponding static attribute set .

[0044] Among them, the first and second layers of the static attribute extraction unit are convolutional layers, the third to tenth layers are all C3 modules, and the last two layers are convolutional layers; the feature maps output by the fifth, seventh, ninth, and twelfth layers form a feature pyramid FPN, which is regressed by a regression head, and the regression head outputs the static attributes on the image plane , thereby constituting the static attribute set corresponding to the i-th large field-of-view image ; is the static attribute on the image plane of the j-th target in the i-th large field-of-view image, and l is the number of targets. An activation function SiLU is set after each layer of the static attribute extraction unit .

[0045] ;

[0046] Among them, represents the target position in the static attribute on the image plane , represents the morphological feature, represents the background information, and respectively represent the centroid coordinates of the j-th target, is the maximum gray value of the area where the target is located, and respectively represent the standard deviation matrix and covariance matrix of the target offset Gaussian model, and respectively represent the background gray mean and standard deviation of the k n pixel neighborhood of the target, k n is a hyperparameter. In this solution, the superscript T represents the transpose operation, and the same applies hereinafter; in actual use, each attribute of is normalized to the range (0, 1).

[0047] The large field-of-view image is first cropped to a predetermined size (640, 640), and then downsampled to a size of (320, 320) by the first and second layers of the static attribute extraction unit , and then enters the C3 module for subsequent processing.

[0048] See Appendix Figure 2, in this solution, the regression head adopts a decoupled design, including a shared convolutional layer and parallel position regression branch, shape regression branch, and confidence branch. The position regression branch, shape regression branch, and confidence branch all adopt the Conv(SiLU(BN(Conv(·)))) structure; where the position regression branch is used to obtain the target position, that is in ; the shape regression branch is used to obtain the shape features and background information of the target, that is in and ; the confidence branch is used to predict the confidence of the target's existence, and then determine whether it is a real target; after the feature map output by the Feature Pyramid Network (FPN) passes through the shared convolutional layer, it is processed by the cooperation of the position regression branch, shape regression branch, and confidence branch, and finally the static attributes of the target on the image plane are obtained ; in the Conv(SiLU(BN(Conv(·)))) structure, Conv represents convolution, BN represents batch normalization, and SiLU represents the Swish activation function.

[0049] The C3 module includes three convolutional layers connected in sequence. After the output features of the first convolutional layer are superimposed with the output features of the third convolutional layer, they are fused with the output features of another convolutional layer on the branch after the output features of the first convolutional layer. The fusion result is then processed by another convolutional layer to obtain the feature map output by the C3 module.

[0050] In one embodiment of the present invention, the network parameters of the first to tenth layers of the static attribute extraction unit are shown in the following table:

[0051] Table 1: Network structure of the static attribute extraction unit.

[0052]

[0053] Step 2: Pairwise combine the first three large field-of-view images in order to construct an attribute hypergraph for each combination; use the attribute hypergraphs corresponding to the first three large field-of-view images to determine the potential target association matching matrix corresponding to each attribute hypergraph through the spatio-temporal attribute association module; based on the static attribute set of the first three large field-of-view images and the potential target association matching matrix, determine the high-order spatio-temporal attributes of the third large field-of-view image.

[0054] Step 2.1: Attribute hypergraph of the initial static attribute set.

[0055] Using the static attribute sets corresponding to the first three large field-of-view images , and as the initial static attribute set, pairwise combine them in chronological order to construct the attribute hypergraphs corresponding to the first large field-of-view image and the second large field-of-view image and the attribute hypergraphs corresponding to the second large field of view image and the third large field of view image ; The attribute hypergraph is represented by the following formula:

[0056] ;

[0057] In the above formula, represents the attribute hypergraph constructed with the targets of the i-th large field of view image and the (i + 1)-th large field of view image as nodes; represents the set of edges connecting different nodes in the i-th large field of view image, and represents the set of edges connecting the nodes of the i-th large field of view image and the (i + 1)-th large field of view image; Define the thresholds and For two nodes belonging to the same large field of view image, if the Euclidean distance between the two nodes is less than , then an edge is established between the two nodes, and the weight is , where e is the natural constant; For two nodes belonging to different large field of view images, if the Euclidean distance between the two nodes is less than , then an edge is established between the two nodes, and the weight is also

[0058] Step 2.2, Potential target association matching matrix.

[0059] After the attribute hypergraphs and are constructed, use the spatio-temporal attribute association module F M to process them to obtain the potential target association matching matrix and at the initial stage, so as to obtain the state change information of the same target between different frames.

[0060] The potential target association matching matrix can be obtained by the following formula:

[0061] ;

[0062] where is the potential target association matching matrix corresponding to the attribute hypergraph constructed using the large field of view image of the i-th frame and the large field of view image of the (i + 1)-th frame. Its row and column coordinates respectively represent the indices of the targets in two consecutive frames, and its elements are the similarities under the corresponding indices. For example, the element at the (m, n) position represents the similarity between the m-th target in the large field of view image of the i-th frame and the n-th target in the large field of view image of the (i + 1)-th frame. Select the value with the highest similarity under the row where the m-th target is located, and use this value as the association relationship of the m-th target between the large field of view image of the i-th frame and the large field of view image of the (i + 1)-th frame (that is, which target in the large field of view image of the (i + 1)-th frame corresponds to the m-th target).

[0063] Spatio-temporal attribute association module F M includes a six-layer network structure. Among them, the first, third, and fifth layers use spatial hypergraph convolution, and the second, fourth, and sixth layers use temporal hypergraph convolution. An activation function ReLU is set after each layer of the network structure, and an association head is set after the sixth layer. The association head uses a dot product attention layer, which is expressed as follows:

[0064] ;

[0065] and respectively represent the feature matrices of the targets in the large field of view image of the i-th frame and the large field of view image of the (i + 1)-th frame, that is, the output of the sixth layer of the network structure. Softmax is used for normalization; is the normalization factor, and its value is the node feature dimension.

[0066] Spatio-temporal attribute association module F M In, spatial hypergraph convolution is used to extract the spatial structure features contained at the same moment, that is, in the same frame, while temporal hypergraph convolution is used to integrate spatio-temporal features and extract deep temporal sequence information. The structures of spatial hypergraph convolution and temporal hypergraph convolution are as follows:

[0067] ;

[0068] Among them, X represents the input of spatial hypergraph convolution or temporal hypergraph convolution, Y is the processed output, A is the adjacency matrix, which stores the connection and weight information of the edges, W is the trainable parameter, is the activation function; Spatial hypergraph convolution and temporal hypergraph convolution are essentially only different in W. Spatial hypergraph convolution only processes the edges in the large field of view image of the same frame, while temporal hypergraph convolution only processes the edges in the large field of view images of different frames.

[0069] In an embodiment of the present invention, the provided spatio-temporal attribute association module F M has a network structure as shown in Table 2:

[0070] Table 2: Network structure of the spatio-temporal attribute association module.

[0071]

[0072] Step 2.3, the high-order spatio-temporal attributes of the third-frame large field-of-view image.

[0073] Using the static attribute set 、 and Combined with its corresponding potential target association matching matrix and , determine the initial high-order spatio-temporal attributes.

[0074] Based on the target association matching matrix and , the association relationship of the j-th target at different times (i.e., different frames) can be obtained. According to the static attribute set 、 and , the image-plane static attributes of the j-th target at different times can be obtained 、 and ; then for the j-th target in the third-frame large field-of-view image, perform difference and second-order difference operations based on the image-plane static attributes of the target at different times to obtain the first-order differential and the second-order differential ; finally, stack the image-plane static attributes, first-order differential, and second-order differential of the target to obtain the high-order spatio-temporal attributes of the target , then the high-order spatio-temporal attributes of the third-frame large field-of-view image :

[0075] ;

[0076] where represents the high-order spatio-temporal attributes of the j-th target obtained after the input of the third-frame large field-of-view image.

[0077] For example, for the image-plane static attributes 、 in the initial static attribute set 、 , taking the target position as an example, although the image-plane static attributes of each target in each frame of the large field-of-view image are extracted in step 1, for each frame of the large field-of-view image, the extracted image-plane static attributes are independent, and it is not clear which target in the previous frame of the large field-of-view image corresponds to which target in the next frame of the large field-of-view image; by using the association relationship of the j-th target in the position of the j-th target in the first-frame large field-of-view image can be obtained as , by taking the difference between the two and then differentiating, the first-order differential of the target position is obtained. The first-order differentials of other morphological features and background information in the static image plane attributes are also obtained in this way, thus constituting the first-order differential of the static image plane attributes of the j-th target in the third large field-of-view image. ; The principle of the second-order differential is the same and will not be elaborated here.

[0078] Step 3. For each large field-of-view image after the third frame, construct an inference hypergraph using the high-order spatio-temporal attributes of the previous large field-of-view image of the current frame, and determine the high-order spatio-temporal attributes of the current frame large field-of-view image; using the inference hypergraph as input, utilize the spatio-temporal inference model to predict the potential target association matching matrix between adjacent frame large field-of-view images, and utilize the potential target association matching matrix to achieve continuous association and classification of targets.

[0079] Step 3.1. Construction of the inference hypergraph.

[0080] For large field-of-view images after the third frame, start from the fourth large field-of-view image as the current frame to construct the inference hypergraph; the structure of the inference hypergraph is basically the same as the attribute hypergraph of the initial static attribute set in Step 2.1, denoted as , where i > 2.

[0081] ;

[0082] Among them: when i > 2, the static attribute extraction unit in Step 1 can be used to obtain the static attribute set corresponding to the i-th large field-of-view image , and the set of edges connecting different nodes in the i-th large field-of-view image is determined in the same way as in Step 2.1, that is, screening node pairs in the i-th large field-of-view image with an Euclidean distance less than the threshold and establishing edges.

[0083] The construction method of the inference hypergraph is basically the same as that of the attribute hypergraph of the initial static attribute set, and the main difference lies in the construction process of the set of edges connecting the nodes of the i-th large field-of-view image and the (i + 1)-th large field-of-view image; in the inference hypergraph , the construction method is as follows:

[0084] First, when constructing the inference hypergraph , establish a hypothetical association relationship between the j-th node (i.e., the target) in the i-th large field-of-view image and any node in the (i + 1)-th large field-of-view image, that is, assume that the two are the same potential target; secondly, calculate the high-order spatio-temporal attributes of the j-th node in the i-th large field-of-view image according to the method in Step 2.3 as its assumed high-order spatio-temporal attribute; finally, define the hyperparameter , and judge the assumed high-order spatio-temporal attribute of the j-th node in the large field-of-view image of the i-th frame , and the assumed high-order spatio-temporal attribute of the -th node in the large field-of-view image of the (i + 1)-th frame (when i > 3) or between the high-order spatio-temporal attributes (when i = 3), and determine whether the Euclidean distance is less than the hyperparameter ; if it is less, then regard the assumed association relationship as the association relationship between the j-th node in the large field-of-view image of the i-th frame between the i-th frame and the (i + 1)-th frame, and connect an edge between the j-th node and the -th node; regard the assumed high-order spatio-temporal attribute as the high-order spatio-temporal attribute of the j-th node, so as to obtain the high-order spatio-temporal attribute of the large field-of-view image of the i-th frame.

[0085] Step 3.2, nested reasoning.

[0086] This step takes the inference hypergraph as the input, and realizes the prediction of the potential target association matching matrix between the large field-of-view image of the i-th frame and the large field-of-view image of the (i + 1)-th frame through the spatio-temporal reasoning model ; by continuously generating for each subsequent large field-of-view image, the continuous association and classification of different targets can be realized by using .

[0087] ;

[0088] wherein is the potential target association matching matrix between the large field-of-view image of the i-th frame and the large field-of-view image of the (i + 1)-th frame predicted based on the inference hypergraph .

[0089] Repeating step 3, the corresponding potential target association matching matrix can be continuously generated for each subsequent large field-of-view image, realizing the continuous association of spatial targets.

[0090] The spatio-temporal reasoning model is a graph convolutional neural network. The first layer, the fourth layer, and the seventh layer of its network structure are spatial hypergraph convolutions, the second layer, the fifth layer, and the eighth layer are temporal hypergraph convolutions, the third layer, the sixth layer, and the ninth layer are temporal attention layers. An activation function ReLU is set after each layer of the network structure, and an association head and a classification head are set after the ninth layer; among them, the association head adopts the dot product attention layer mentioned in step 2.2, and the classification head is implemented by using the softmax function. Finally, the potential target association matching matrix and the categories of the targets (LEO, GEO, non-targets, etc.), thereby achieving continuous association and classification of the targets.

[0091] Spatial hypergraph convolution and temporal hypergraph convolution were described in the previous step 2.2; and the temporal attention layer uses a graph convolutional long short-term memory network, which is expressed as follows:

[0092] ;

[0093] Among them, and respectively represent the spatial structure features and deep temporal information extracted by the spatial hypergraph convolution and temporal hypergraph convolution before the temporal attention layer, represents the predicted deep temporal information; and respectively represent the forget gate and the input gate.

[0094] In an embodiment of the present invention, the spatio-temporal reasoning model has a network structure as shown in Table 3.

[0095] Table 3: Network structure of the spatio-temporal reasoning model.

[0096]

[0097] As Figure 3 shown, in an embodiment of the present invention, for large field of view detection data, the F1 score of the present invention is improved by 0.04 - 0.17 and the MOTA is improved by 0.03 - 0.16 compared with other association strategies, which proves the effectiveness of the present invention.

[0098] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A large-field-of-view camera space target association method based on spatiotemporal high-order attribute hypergraph, characterized in that: include: A large-field-of-view image sequence is acquired by using a large-field-of-view camera, and for each frame of the large-field-of-view image in the large-field-of-view image sequence, a static attribute extraction unit is used to extract image plane static attributes of each target, thereby obtaining a static attribute set corresponding to each frame of the large-field-of-view image; Combine the first three frames of large field of view images in pairs in order, and construct an attribute hypergraph for each combination based on a static attribute set; Using the attribute hypergraphs corresponding to the first three frames of large field of view images, the potential target association matching matrix corresponding to each attribute hypergraph is determined through the spatiotemporal attribute association module, including: The input of the spatiotemporal attribute association module is the attribute hypergraph. The spatiotemporal attribute association module includes a six-layer network structure, in which the first, third and fifth layers use spatial hypergraph convolution, and the second, fourth and sixth layers use temporal hypergraph convolution. An activation function is set after each layer of the network structure, and an association head is set after the sixth layer. The association head uses a dot product attention layer, which can be expressed as: Among them, softmax is used for normalization, M i,i+1 is the potential target association matching matrix corresponding to the attribute hypergraph constructed using the i-th frame large field of view image and the i+1-th frame large field of view image, which contains the association relationship between each target in the two frames of large field of view images; F i With F i+1 They represent the characteristic matrices of the targets in the i-th frame large field of view image and the i+1-th frame large field of view image, respectively. D′ is the normalization factor, and the superscript T indicates transposition. The structures of spatial hypergraph convolution and temporal hypergraph convolution are as follows: Y = σ(AXW); Where X represents the input of spatial hypergraph convolution or temporal hypergraph convolution, Y is the processed output, A is the adjacency matrix, which stores the edge connection and weight information, W is the trainable parameter, and σ is the activation function; Based on the static attribute set of the first three frames of large field of view images and the potential target association matching matrix, the high-order spatiotemporal attributes of the third frame of large field of view image are determined, including: For each target in the third frame of the large field of view image, differential and second-order differential operations are performed based on the static properties of the target's image plane at different times to obtain the first-order differential and second-order differential of the target; finally, the static properties of the target's image plane, the first-order differential and the second-order differential are stacked to obtain its high-order spatiotemporal properties. The high-order spatiotemporal properties of all targets constitute the high-order spatiotemporal properties of the third frame of the large field of view image; For each large field of view image after the third frame, an inference hypergraph is constructed based on the high-order spatiotemporal attributes of the previous large field of view image of the current frame, and the high-order spatiotemporal attributes of the current frame are determined. With the inference hypergraph as input, the spatiotemporal inference model is used to predict the potential target association matching matrix of adjacent large field of view images, and the potential target association matching matrix is ​​used to achieve continuous association and classification of targets.

2. The method for associating targets in large-field-of-view camera space based on spatiotemporal high-order attribute hypergraph according to claim 1, characterized in that: The static properties of each target’s image plane include target position, morphological features, and background information; The morphological features include the maximum grayscale of the target area, the standard deviation matrix and covariance matrix of the target offset Gaussian model; the background information includes the background grayscale mean and standard deviation of the target neighborhood.

3. The method for associating targets in a large field of view camera space based on a spatiotemporal high-order attribute hypergraph according to claim 1, characterized in that: The first and second layers of the static attribute extraction unit are convolutional layers, the third to tenth layers are C3 modules, and the last two layers are convolutional layers; the feature maps output by the fifth, seventh, ninth, and twelfth layers constitute a feature pyramid, which is regressed by a regression head, and the image plane static attributes are output by the regression head, thereby forming a static attribute set corresponding to each frame of a large field of view image; an activation function is set after each layer of the static attribute extraction unit; The C3 module includes three convolutional layers connected in sequence. The output features of the first convolutional layer are superimposed with the output features of the third convolutional layer, and then fused with the output features of the first convolutional layer after passing through another convolutional layer on the branch. The fusion result is processed by another convolutional layer to obtain the feature map output by the C3 module.

4. The method for associating targets in a large field of view camera space based on a spatiotemporal high-order attribute hypergraph according to claim 3, characterized in that: The regression head adopts a decoupled design, including a shared convolutional layer and parallel position regression branches, morphological regression branches and confidence branches, and the position regression branch, morphological regression branch and confidence branch all adopt a Conv(SiLU(BN(Conv(·)))) structure; The position regression branch is used to obtain the target position, the morphological regression branch is used to obtain the morphological features and background information of the target, and the confidence branch is used to predict the confidence of the target existence, and then determine whether it is a real target; After the feature map output by the feature pyramid passes through the shared convolution layer, it is processed by the position regression branch, the morphology regression branch and the confidence branch to finally obtain the static properties of the target image surface.

5. The method for associating targets in large-field-of-view camera space based on spatiotemporal high-order attribute hypergraph according to claim 1, characterized in that: Combine the first three frames of large field of view images in pairs in order, and construct the attribute hypergraph of each combination based on the static attribute set, including: Construct the attribute hypergraph G corresponding to the first frame of large field of view image and the second frame of large field of view image 1,2 And the attribute hypergraph G corresponding to the second frame large field of view image and the third frame large field of view image 2,3 ; The attribute hypergraph is represented by the following formula: G i,i+1 =(S i ,S i+1 ,E i ,E i+1 ,E i,i+1 ),i=1,2; In the above formula, G i,i+1 represents the attribute hypergraph constructed with the targets of the i-th frame large field of view image and the i+1-th frame large field of view image as nodes; S i , S i+1 are the static attribute sets corresponding to the i-th frame large field of view image and the i+1-th frame large field of view image respectively; E i , E i+1 are the sets of edges connecting different nodes in the i-th frame large field of view image and the i+1-th frame large field of view image, respectively, and E i,i+1 It represents the edge set connected between the nodes of the i-th frame large field of view image and the i+1-th frame large field of view image; define the threshold l S With l T , for two nodes belonging to the same frame of large field of view image, if the Euclidean distance d between the two nodes is less than l S , then an edge is established between the two nodes; for two nodes belonging to different frames of large field of view images, if the Euclidean distance d between the two nodes is less than l T , an edge is established between two nodes.

6. The method for associating targets in a large field of view camera space based on a spatiotemporal high-order attribute hypergraph according to claim 5, characterized in that: The reasoning hypergraph is constructed based on the high-order spatiotemporal attributes of the previous large field of view image of the current frame, and the high-order spatiotemporal attributes of the current large field of view image are determined, including: The structure of the inference hypergraph is the same as that of the attribute hypergraph. The difference between the inference hypergraph and the attribute hypergraph in the construction method is the edge set E between the nodes of the i-th frame large field of view image and the i+1-th frame large field of view image. i,i+1 The construction process is different; in the reasoning hypergraph In, E i,i+1 The construction method is as follows: First, for the current frame large field of view image with i>2, when constructing the inference hypergraph When the jth node in the i-th frame of the large field of view image is connected to any j′th node in the i+1-th frame of the large field of view image, it is assumed that the two are the same potential target; secondly, the high-order spatiotemporal attributes of the jth node in the i-th frame of the large field of view image are calculated. As its assumed high-order spatiotemporal properties; finally, define the hyperparameter k h , determine the assumed high-order spatiotemporal properties of the jth node in the i-th frame of the large field of view image Is the Euclidean distance between the assumed high-order spatiotemporal attributes of the j′th node in the i+1th frame of the large field of view image less than the hyperparameter k? h ; If it is less than, the assumed association relationship is taken as the association relationship between the jth node in the i-th frame large field image and the i+1-th frame large field image, and an edge is connected between the jth node and the j′th node; the assumed high-order spatiotemporal attribute As the high-order spatiotemporal attributes of the j-th node, the high-order spatiotemporal attributes of the i-th frame large field of view image are obtained.

7. The method for associating targets in large-field-of-view camera space based on spatiotemporal high-order attribute hypergraph according to claim 1, characterized in that: Taking the inference hypergraph as input, the spatiotemporal inference model is used to predict the potential target association matching matrix of adjacent frame large field of view images, including: The spatiotemporal reasoning model is a graph convolutional neural network. The first, fourth and seventh layers of the spatiotemporal reasoning model network structure are spatial hypergraph convolutions, the second, fifth and eighth layers are temporal hypergraph convolutions, the third, sixth and ninth layers are temporal attention layers, an activation function is set after each layer of the network structure, and an association head and a classification head are set after the ninth layer; wherein the association head adopts a dot product attention layer, and the classification head adopts a softmax function to implement; The reasoning hypergraph is input into the spatiotemporal reasoning model, and finally the potential target association matching matrix and the target category are obtained through the combination of the association head and the classification head, thereby realizing the continuous association and classification of the target; the temporal attention layer adopts a graph convolutional long short-term memory network.

8. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor; characterized in that: When the processor executes the computer program, the method for associating targets in a large-field-of-view camera space based on a spatiotemporal high-order attribute hypergraph according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Segmentation method based on limited temporal-spatial resolution class-independent attribute dynamic scene

    CN108053420A

  • Associated information mining method based on space-time hypergraph

    CN116306924A