A method for detecting and identifying high-energy instantaneous events in remote sensing images
By performing time channel weighting processing on remote sensing images and improving the Yolo-v7 object detection model, combining object detection of long input streams and short input streams and target tracking processing of ByteTrack, the problem of high-energy instantaneous event detection and recognition in remote sensing images is solved, and efficient and real-time detection and recognition effects are achieved.
Patent Information
- Application Number
- CN202410335005.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-03-22
AI Technical Summary
In the detection and recognition of high-energy instantaneous event, the prior art has problems such as small event targets, complex backgrounds, susceptible to noise interference such as clouds, inaccurate recognition of timing textures, and inability to realize real-time detection.
A high-energy instantaneous event detection and recognition method for remote sensing images is adopted, including collecting remote sensing images and performing time channel weighting processing, building an improved Yolo-v7 object detection model, target detection through long input streams and short input streams, and post-processing using the target tracker ByteTrack to obtain the object description information.
Accurate detection and identification of instantaneous high-energy events in low signal-to-noise ratio remote sensing images is realized, real-time and accuracy of detection is improved, false alarm rate is reduced, and models can be widely deployed in edge computing systems.
Smart Images

Figure CN118247655B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image recognition, and in particular relates to a remote sensing image high-energy instantaneous event detection and recognition method. Background Art
[0002] In recent years, with the rapid development of deep learning-based image processing technology, target detection and recognition for remote sensing images have been widely used in industries such as agriculture, transportation, and geology. Unlike conventional images, remote sensing images are usually captured from a bird's-eye view in space, which often have the characteristics of small target proportion, unbalanced position distribution, complex background, and low signal-to-noise ratio, which brings great challenges to image detection model methods.
[0003] Traditional feature extraction algorithms, such as corner detection and edge extraction, have difficulties in detecting high-energy instantaneous events in remote sensing images, such as unclear target features, large data volume and high computational complexity. However, a high-dimensional neural network parameter model is obtained by using data-driven training and fitting, which enhances the detection and recognition capabilities of targets in remote sensing images to a certain extent.
[0004] Among the target detection models based on neural networks, the mainstream algorithms include Faster-RCNN and Yolo series. Among them, the Yolo series model algorithm is a single-stage end-to-end target detection model with fast calculation speed, high recognition accuracy, and easy fusion of multi-scale features. It is suitable for detecting small targets in remote sensing images in production environments. Different from conventional remote sensing targets such as buildings and environments, high-energy instantaneous events are a type of event target that occurs in a time sequence in remote sensing images. For example, rocket launches appear as high-energy bright spots with certain changing rules in remote sensing images. This bright spot has less texture in pixel space, but is reflected in the texture change in the time dimension. In addition, the characteristics of small event targets, complex backgrounds, and susceptibility to noise interference such as clouds bring great difficulties to detection. On the other hand, due to the real-time requirements of the detection task, the scale and detection speed of the algorithm model need to be considered. A detection model with too small a scale has fewer parameters, which may lead to insufficient fitting ability, but a model with too large a scale affects the real-time performance of the algorithm. Summary of the invention
[0005] The present invention provides a method for detecting and identifying high-energy instantaneous events in remote sensing images, which solves the problems in the prior art of detecting and identifying high-energy instantaneous events, such as small event targets, complex backgrounds, susceptibility to interference from noise such as clouds, inaccurate recognition of temporal textures, and failure to achieve real-time detection.
[0006] In order to solve the above technical problems, the technical solution of the present invention is: a method for detecting and identifying high-energy instantaneous events in remote sensing images, comprising the following steps:
[0007] S1, collect remote sensing images, perform time channel weighted processing on the remote sensing images, and obtain feature maps arranged in three-dimensional space;
[0008] S2, processing the feature graph arranged in three-dimensional space to obtain a long input stream and a short input stream;
[0009] S3. Build and improve the Yolo-v7 target detection model;
[0010] S4, input the long input stream and the short input stream into the improved Yolo-v7 target detection model, perform target detection, and obtain the target detection result positioning;
[0011] S5, match and cache the target detection result positioning through the target tracker ByteTrack to obtain the cached target information;
[0012] S6. Calculate the target description information based on the cached target information to complete the real-time detection and identification of high-energy instantaneous events.
[0013] Furthermore, the specific steps of step S1 are: collecting a remote sensing image, sliding the entire three-dimensional space of the remote sensing image through a sliding window filter, and obtaining a feature map arranged in the three-dimensional space, wherein the sliding window filter moves on the height of the remote sensing image, the width of the remote sensing image, and the time channel of the remote sensing image, and obtains a numerical value at each position element by element according to the multiplication and addition weighted calculation of the sliding window filter parameters.
[0014] Furthermore, the specific steps of step S2 are:
[0015] S21, taking the detected frame in the feature map arranged in the three-dimensional space as the 0th frame;
[0016] S22, taking the feature maps of the three-dimensional space arrangement of the -6th frame, the -4th frame, the -2th frame, the -1th frame, the 0th frame, the 1st frame and the 3rd frame, and arranging them in time order to form a three-dimensional image sequence to obtain a long input stream;
[0017] S22, taking the feature maps of the three-dimensional space arrangement of the -1th frame, the 0th frame and the 1st frame, and arranging them in time order to form a three-dimensional image sequence to obtain a short input stream.
[0018] Furthermore, the improved Yolo-v7 target detection model in step S3 includes a sampling module and a feature cascade fusion module;
[0019] The sampling module includes two convolution-normalization layers and three maximum pooling layers connected in sequence, wherein the convolution kernel size of the convolution-normalization layer is 3 and the convolution step size is 2.
[0020] Furthermore, the target detection result positioning in step S4 includes position information, category information and confidence information of the detection result box.
[0021] Furthermore, the specific steps of step S4 are:
[0022] S41, inputting the long input stream and the short input stream in parallel into the sampling module to obtain feature graph sequences of three different spatiotemporal scales;
[0023] S42: Input the feature map sequences of three different spatiotemporal scales into the feature cascade fusion module to obtain the position information, category information and confidence information of the detection result box.
[0024] Furthermore, the specific steps of step S42 are:
[0025] A1. Input the feature map sequences of three different spatiotemporal scales into the local cross-level feature scaling network to obtain feature description information of three different sizes;
[0026] A2. Input the feature description information of three different sizes into the sigmoid activation function for prediction to obtain the relative position and size information of the target prediction box;
[0027] A3. Generate a series of target candidate frames in the original image area based on the relative position and size information of the target prediction frame;
[0028] A4. A non-maximum suppression operation is performed on the target candidate frame to retain the detection result frame that meets the detection target among the overlapping target frames, wherein the detection result frame includes the position information, category information and confidence information of the detection result frame.
[0029] Furthermore, the local cross-level feature scaling network in step A1 includes four 3×3 convolutions and one 1×1 convolution, wherein the four 3×3 convolutions are respectively a first 3×3 convolution with a channel number of 2C, a second 3×3 convolution with a channel number of C, a third 3×3 convolution with a channel number of C, and a fourth 3×3 convolution with a channel number of C, and the number of channels of the 1×1 convolution is 4C;
[0030] The input of the first 3×3 convolution is the input of the local cross-level feature scaling network, the first 3×3 convolution outputs a first convolution feature map and a second convolution feature map, the second convolution feature map is the input of the second 3×3 convolution, the output of the second 3×3 convolution is the input of the third 3×3 convolution, the third 3×3 convolution outputs a third convolution feature map, the third convolution feature map is the input of the fourth 3×3 convolution, the fourth 3×3 convolution outputs a fourth convolution feature map, the first convolution feature map, the second convolution feature map, the third convolution feature map and the fourth convolution feature map are jointly input into a 1×1 convolution, and the output of the 1×1 convolution is the output of the local cross-level feature scaling network.
[0031] Furthermore, the specific steps of step S5 are:
[0032] S51, dividing the detection result box into high-scoring detection results and low-scoring detection results according to the confidence information of the detection result box;
[0033] S52, matching the track that successfully matched the high-scoring detection result at the previous moment through the target tracker ByteTrack, obtaining the track that successfully matched at the current moment and the track that failed to match, and caching them, wherein if there is no track that has been successfully matched at the previous moment, the cached high-scoring detection result is used as the starting point of the candidate track;
[0034] S53, matching the unmatched trajectories with the low-score detection results, filtering the background, caching, and obtaining cached target information.
[0035] Furthermore, the description information of the target in step S6 includes the location, category, area, duration, maximum energy moment and maximum energy value of the target.
[0036] The beneficial effects of the present invention are as follows: (1) by introducing a time channel and performing weighted processing on the time channel to form a spatiotemporal convolution, the feature extraction of the spatial texture change in the time series is realized, thereby realizing the detection and recognition of instantaneous high-energy events in low signal-to-noise ratio remote sensing images;
[0037] (2) The feature graph sequence obtained by the long input stream focuses on the information of spatial texture changes before and after the time sequence, which helps the model to better identify and classify the categories of high-energy events. The feature graph sequence obtained by the short input stream focuses on the high-energy detected targets that last for a short period of time, allowing the model to better locate the target in the image space.
[0038] (3) The improved Yolo-v7 target detection model adopts a fully convolutional network framework, which can cope with inputs of different sizes without retraining during the reasoning process, taking into account both detection accuracy and detection efficiency. Since the model parameters are small, the calculation and reasoning speed is guaranteed while enabling it to be more widely deployed in edge computing systems. Therefore, the improved Yolo-v7 target detection model is small in scale and fast in speed, and can be applied to tasks such as space-based reconnaissance and remote sensing monitoring.
[0039] (4) The local cross-level scaling network performs convolutions of different scales on some channel feature maps and then cascades and fuses them to obtain a new feature map sequence, that is, feature description information of three different sizes. The local cross-level scaling network expands the width and depth of the convolutional neural network, while also ensuring the effective propagation of gradients during training, thereby ensuring the detection effect while improving the speed of training and inference.
[0040] (5) By post-processing the model detection results through the target tracker ByteTrack, it is possible to more completely track targets that may have low detection confidence scores due to noise interference or motion occlusion, obtain more descriptive information about the target, and exclude some detection results that do not have target features, thereby reducing the false alarm rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 The present invention is a flowchart of a method for detecting and identifying high-energy instantaneous events in remote sensing images.
[0042] Figure 2 This is a structural diagram of the local cross-level feature scaling network of the present invention.
[0043] Figure 3 This is a diagram showing the detection and recognition effect of high-energy instantaneous events in remote sensing images according to the present invention. DETAILED DESCRIPTION
[0044] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.
[0045] Example
[0046] like Figure 1 As shown, the present invention provides a method for detecting and identifying high-energy instantaneous events in remote sensing images, comprising the following steps:
[0047] S1, collect remote sensing images, perform time channel weighted processing on the remote sensing images, and obtain feature maps arranged in three-dimensional space;
[0048] S2, processing the feature graph arranged in three-dimensional space to obtain a long input stream and a short input stream;
[0049] S3. Build and improve the Yolo-v7 target detection model;
[0050] S4, input the long input stream and the short input stream into the improved Yolo-v7 target detection model, perform target detection, and obtain the target detection result positioning;
[0051] S5, match and cache the target detection result positioning through the target tracker ByteTrack to obtain the cached target information;
[0052] S6. Calculate the target description information based on the cached target information to complete the real-time detection and identification of high-energy instantaneous events.
[0053] The specific steps of step S1 are: collecting a remote sensing image, sliding the entire three-dimensional space of the remote sensing image through a sliding window filter, and obtaining a feature map arranged in the three-dimensional space, wherein the sliding window filter moves on the height of the remote sensing image, the width of the remote sensing image, and the time channel of the remote sensing image, and obtains a value element by element at each position according to the multiplication and addition weighted calculation of the sliding window filter parameters.
[0054] In this embodiment, the time channel is introduced, and the convolution input has an additional dimension of the number of channels C, which is a three-dimensional matrix of H×W×C, where H represents the height and W represents the width. The sliding window filter moves in three directions (i.e., the height, width and time channel of the image), and slides through the entire three-dimensional space through the filter to obtain a feature map arranged in the three-dimensional space. The method of the present invention can realize the recognition of targets larger than 4×4 pixels in a million-pixel image.
[0055] The specific steps of step S2 are:
[0056] S21, taking the detected frame in the feature map arranged in the three-dimensional space as the 0th frame;
[0057] S22, taking the feature maps of the three-dimensional space arrangement of the -6th frame, the -4th frame, the -2th frame, the -1th frame, the 0th frame, the 1st frame and the 3rd frame, and arranging them in time order to form a three-dimensional image sequence to obtain a long input stream;
[0058] S22, taking the feature maps of the three-dimensional space arrangement of the -1th frame, the 0th frame and the 1st frame, and arranging them in time order to form a three-dimensional image sequence to obtain a short input stream.
[0059] The improved Yolo-v7 target detection model in step S3 includes a sampling module and a feature cascade fusion module;
[0060] The sampling module includes two convolution-normalization layers and three maximum pooling layers connected in sequence, wherein the convolution kernel size of the convolution-normalization layer is 3 and the convolution step size is 2.
[0061] The target detection result positioning in step S4 includes the location information, category information and confidence information of the detection result box.
[0062] The specific steps of step S4 are:
[0063] S41, inputting the long input stream and the short input stream in parallel into the sampling module to obtain feature graph sequences of three different spatiotemporal scales;
[0064] S42: Input the feature map sequences of three different spatiotemporal scales into the feature cascade fusion module to obtain the position information, category information and confidence information of the detection result box.
[0065] The specific steps of step S42 are:
[0066] A1. Input the feature map sequences of three different spatiotemporal scales into the local cross-level feature scaling network to obtain feature description information of three different sizes;
[0067] A2. Input the feature description information of three different sizes into the sigmoid activation function for prediction to obtain the relative position and size information of the target prediction box;
[0068] A3. Generate a series of target candidate frames in the original image area based on the relative position and size information of the target prediction frame;
[0069] A4. A non-maximum suppression operation is performed on the target candidate frame to retain the detection result frame that meets the detection target among the overlapping target frames, wherein the detection result frame includes the position information, category information and confidence information of the detection result frame.
[0070] The local cross-level feature scaling network in step A1 includes four 3×3 convolutions and one 1×1 convolution, wherein the four 3×3 convolutions are respectively a first 3×3 convolution with a channel number of 2C, a second 3×3 convolution with a channel number of C, a third 3×3 convolution with a channel number of C, and a fourth 3×3 convolution with a channel number of C, and the number of channels of the 1×1 convolution is 4C;
[0071] The input of the first 3×3 convolution is the input of the local cross-level feature scaling network, the first 3×3 convolution outputs a first convolution feature map and a second convolution feature map, the second convolution feature map is the input of the second 3×3 convolution, the output of the second 3×3 convolution is the input of the third 3×3 convolution, the third 3×3 convolution outputs a third convolution feature map, the third convolution feature map is the input of the fourth 3×3 convolution, the fourth 3×3 convolution outputs a fourth convolution feature map, the first convolution feature map, the second convolution feature map, the third convolution feature map and the fourth convolution feature map are jointly input into a 1×1 convolution, and the output of the 1×1 convolution is the output of the local cross-level feature scaling network.
[0072] In this embodiment, by improving the original Yolo-v7, the size and number of the backbone network convolution kernels and the setting of the anchor box are adjusted according to the characteristics of small pixel size and unclear spatial texture of the task target, thereby obtaining an improved Yolo-v7 target detection model.
[0073] Among them, the long and short input streams are input into the sampling module in parallel to obtain three feature map sequences of different spatiotemporal scales, that is, three feature maps of one quarter, one eighth, and one sixteenth of the original size, and the corresponding channel numbers C are 128, 256, and 512 respectively.
[0074] The feature cascade fusion module outputs possible target anchor box information, that is, the location information, category information, and confidence information of the detection result box. The local cross-level feature scaling network in the feature cascade fusion module transforms the feature size and splices it according to the channel, such as Figure 2 As shown, it includes 4 3×3 convolutions and one 1×1 convolution, wherein the feature map input to the local cross-level feature scaling network is convolved through the first 3×3 convolution to obtain the first convolution feature map and the second convolution feature map with the number of channels C, and the second convolution feature map is input to the second 3×3 convolution, and then input to the third 3×3 convolution for convolution to obtain the third convolution feature map, and the third convolution feature map is input to the fourth 3×3 convolution for convolution to obtain the fourth convolution feature map, and finally the first convolution feature map, the second convolution feature map, the third convolution feature map and the fourth convolution feature map are input into the 1×1 convolution with the number of channels 4C to obtain the local cross-level feature scaling network input feature map, that is, feature description information of three different sizes is obtained.
[0075] The local cross-level feature scaling network performs convolutions of three feature map sequences of different spatiotemporal scales on them at different scales and then cascades them to obtain a new feature map sequence, thus obtaining feature description information of three different sizes. This ensures the effective propagation of gradients during training, ensuring the detection effect while improving the speed of training and inference.
[0076] The specific steps of step S5 are:
[0077] S51, dividing the detection result box into high-scoring detection results and low-scoring detection results according to the confidence information of the detection result box;
[0078] S52, matching the track that successfully matched the high-scoring detection result at the previous moment through the target tracker ByteTrack, obtaining the track that successfully matched at the current moment and the track that failed to match, and caching them, wherein if there is no track that has been successfully matched at the previous moment, the cached high-scoring detection result is used as the starting point of the candidate track;
[0079] S53, matching the unmatched trajectories with the low-score detection results, filtering the background, caching, and obtaining cached target information.
[0080] The description information of the target in step S6 includes the location, category, area, duration, maximum energy moment and maximum energy value of the target.
[0081] In this embodiment, the target detection result positioning is divided into two parts, high-scoring detection result and low-scoring detection result, by the target tracker ByteTrack, and the two parts are processed separately, wherein the high-scoring detection result is matched with the track that was successfully matched at the previous moment, and the track after the successful matching at the current moment and the track that was not successfully matched due to obstruction, motion blur, or size change are obtained. Then, these unsuccessfully matched tracks are matched with the low-scoring detection results, and the background is filtered out. Therefore, the target tracker ByteTrack can more completely track targets that may have a low detection confidence score due to noise interference or motion occlusion.
[0082] At the same time, through the target tracker ByteTrack, possible targets in the detection results are numbered and cached, so that more descriptive information about the target can be obtained, such as the maximum energy moment and maximum energy value of the target. At the same time, some detection results that do not have target characteristics can be excluded, reducing the false alarm rate.
[0083] In this embodiment, in order to solve the problem of less real data in the training process of improving the Yolo-v7 target detection model, the Unreal 5 engine is used to produce simulated data. Different terrains are created and used by the Unreal Engine, with realistic lighting rendering, object surface textures, and particle motion systems.
[0084] This embodiment constructs a mountainous terrain, places missile models at random positions in the terrain and triggers them randomly during the simulation process, and places the camera view at a sufficient height perpendicular to the ground to record instantaneous high-energy events during missile launch, flight, and collision strike during the simulation. In the training process of the improved Yolo-v7 target detection model, in order to save the manpower cost of manual labeling, an iterative update method is adopted. After labeling a small amount of data, the model can be trained and updated. After the update, the improved Yolo-v7 target detection model is used to infer a new part of the data to be trained. The tester further modifies these inference results and can quickly adjust the model according to actual needs to ultimately achieve better detection effects, such as Figure 3 As shown, the box is a high-energy transient event target that was successfully detected.
Claims
1. A method for detecting and identifying high-energy instantaneous events in remote sensing images, characterized in that: The following steps are involved: S1, collect remote sensing images, perform time channel weighted processing on the remote sensing images, and obtain feature maps arranged in three-dimensional space; S2. Process the feature graph arranged in three-dimensional space to obtain long input stream and short input stream, which are specifically: S21, taking the detected frame in the feature map arranged in the three-dimensional space as the 0th frame; S22, taking the feature maps of the three-dimensional space arrangement of the -6th frame, the -4th frame, the -2th frame, the -1th frame, the 0th frame, the 1st frame and the 3rd frame, and arranging them in time order to form a three-dimensional image sequence to obtain a long input stream; S22, taking the feature maps of the three-dimensional space arrangement of the -1th frame, the 0th frame and the 1st frame, and arranging them in time order to form a three-dimensional image sequence to obtain a short input stream; S3. Build and improve the Yolo-v7 target detection model; S4. Input the long input stream and the short input stream into the improved Yolo-v7 target detection model to perform target detection and obtain the target detection result positioning, which is specifically: S41, inputting the long input stream and the short input stream in parallel into the sampling module to obtain feature graph sequences of three different spatiotemporal scales; S42, inputting the feature map sequences of three different spatiotemporal scales into the feature cascade fusion module to obtain the position information, category information and confidence information of the detection result box; S5, match and cache the target detection result positioning through the target tracker ByteTrack to obtain the cached target information; S6. Calculate the target description information based on the cached target information to complete the real-time detection and identification of high-energy instantaneous events; The target detection result positioning in step S4 includes the location information, category information and confidence information of the detection result box; The improved Yolo-v7 target detection model includes a sampling module and a feature cascade fusion module; The sampling module includes two convolution-normalization layers and three maximum pooling layers connected in sequence, wherein the convolution kernel size of the convolution-normalization layer is 3 and the convolution step size is 2; The local cross-level feature scaling network in the feature cascade fusion module includes four 3×3 convolutions and one 1×1 convolution, wherein the four 3×3 convolutions are respectively a first 3×3 convolution with a channel number of 2C, a second 3×3 convolution with a channel number of C, a third 3×3 convolution with a channel number of C, and a fourth 3×3 convolution with a channel number of C, and the number of channels of the 1×1 convolution is 4C; The input of the first 3×3 convolution is the input of the local cross-level feature scaling network, the first 3×3 convolution outputs a first convolution feature map and a second convolution feature map, the second convolution feature map is the input of the second 3×3 convolution, the output of the second 3×3 convolution is the input of the third 3×3 convolution, the third 3×3 convolution outputs a third convolution feature map, the third convolution feature map is the input of the fourth 3×3 convolution, the fourth 3×3 convolution outputs a fourth convolution feature map, the first convolution feature map, the second convolution feature map, the third convolution feature map and the fourth convolution feature map are jointly input into a 1×1 convolution, and the output of the 1×1 convolution is the output of the local cross-level feature scaling network.
2. The method for detecting and identifying high-energy instantaneous events in remote sensing images according to claim 1, characterized in that: The specific steps of step S1 are: collecting a remote sensing image, sliding the entire three-dimensional space of the remote sensing image through a sliding window filter, and obtaining a feature map arranged in the three-dimensional space, wherein the sliding window filter moves on the height of the remote sensing image, the width of the remote sensing image, and the time channel of the remote sensing image, and obtains a value element by element at each position according to the multiplication and addition weighted calculation of the sliding window filter parameters.
3. The method for detecting and identifying high-energy instantaneous events in remote sensing images according to claim 1, characterized in that: The specific steps of step S42 are: A1. Input the feature map sequences of three different spatiotemporal scales into the local cross-level feature scaling network to obtain feature description information of three different sizes; A2. Input the feature description information of three different sizes into the sigmoid activation function for prediction to obtain the relative position and size information of the target prediction box; A3. Generate a series of target candidate boxes in the original image area based on the relative position and size information of the target prediction box; A4. A non-maximum suppression operation is performed on the target candidate frame to retain the detection result frame that meets the detection target among the overlapping target frames, wherein the detection result frame includes the position information, category information and confidence information of the detection result frame.
4. The method for detecting and identifying high-energy instantaneous events in remote sensing images according to claim 3, characterized in that: The specific steps of step S5 are: S51, dividing the detection result box into high-scoring detection results and low-scoring detection results according to the confidence information of the detection result box; S52, matching the track that successfully matched the high-scoring detection result at the previous moment through the target tracker ByteTrack, obtaining the track that successfully matched at the current moment and the track that failed to match, and caching them, wherein if there is no track that has been successfully matched at the previous moment, the cached high-scoring detection result is used as the starting point of the candidate track; S53: Match the unmatched trajectories with the low-score detection results, filter the background, cache, and obtain cached target information.
5. The method for detecting and identifying high-energy instantaneous events in remote sensing images according to claim 1, characterized in that: The description information of the target in step S6 includes the location, category, area, duration, maximum energy moment and maximum energy value of the target.
Citation Information
Patent Citations
Camera monitoring network target matching method based on camera network topology
CN116824490A
Efficient retrieval of a target from an image in a collection of remotely sensed data
US20220319144A1