Space target image reconstruction method, device and equipment in dynamic observation scene
By constructing a dynamic sequence multi-exposure dataset and utilizing an unsupervised learning neural network, the problem of inaccurate image reconstruction in dynamic observation scenes is solved, efficient HDR image reconstruction and preservation of multi-perspective geometric relationships are achieved, and the effectiveness of space situational awareness tasks is improved.
Patent Information
- Application Number
- CN202411462263.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-18
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-10-18
AI Technical Summary
Existing technologies have difficulty effectively handling complex target overlapping areas in dynamic observation scenarios, resulting in inaccurate image reconstruction. In addition, traditional methods are sensitive to changes in multi-perspective geometric relationships in spatial situational awareness tasks, affecting tasks such as three-dimensional reconstruction and pose estimation.
A space target image reconstruction method based on unsupervised learning is adopted. By constructing a dynamic sequence multi-exposure dataset, a neural network consisting of a feature extraction unit, a feature deformable alignment unit, and a feature fusion reconstruction unit is used to generate high dynamic range images and preserve the multi-view geometric relationships of the observed targets.
It achieves efficient HDR image reconstruction in dynamic scenes, effectively preserves the multi-perspective geometric relationship of the observed target, and improves the visual effect of the image and the computer processing effect of subsequent space situation awareness tasks.
Smart Images

Figure CN119418163B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of space target imaging, in particular to a space target image reconstruction method, device and equipment in a dynamic observation scene. BACKGROUND
[0002] In recent years, with the gradual development of space resources, the importance of developing space-based optical space target situation awareness technology has gradually increased. Space-based imaging means plays an important role in spacecraft maintenance and debris removal tasks, and space situation awareness technology can effectively monitor space debris and ensure the normal operation and maintenance of on-orbit spacecraft. In space-based optical imaging, due to the complex space lighting environment and the large movement speed of the imaging platform and the observation target in the intersection process, the lighting angle and the scene brightness change greatly, therefore, in the actual scene, a rolling exposure mechanism is often used for imaging, that is, the exposure amount of the camera is changed in a very short time to obtain a continuous space target multi-exposure image sequence, so that the image sequence contains various detail information of the observation target under different exposure degrees.
[0003] For space image sequences, although the exposure degrees are different, the dynamic range of each image is limited. Dynamic range refers to the ratio of the maximum luminance to the minimum luminance in an image or scene. The range of luminance in a real scene can reach 8-10 orders of magnitude, and the dynamic range observable by the human visual system is also very wide, while the dynamic range captured by the camera during actual imaging is relatively low. Multi-exposure fusion (MEF) is a main method for obtaining a high dynamic range image (HDR). This method is based on the complementary and redundant information between low dynamic range (LDR) images with different exposure degrees, and uses image fusion technology to accurately represent high dynamic scenes on a single image, which can effectively enhance the detail information of the image and improve the visual effect of the image. Space multi-exposure fusion aims to enhance the texture details of the observation target from the subjective visual, and lays a good foundation for subsequent computer processing of space situation awareness tasks.
[0004] However, the traditional dynamic multi-exposure fusion method based on the removal of motion pixels is difficult to process complex target overlapping areas, and a lot of effective information will be lost. In addition, due to the particularity of the subsequent tasks of space target situation awareness technology, such as three-dimensional reconstruction, attitude estimation, etc., the multi-view geometric relationship of the observation target changes very sensitively, and the traditional method based on registration is easy to cause the change of the representation of the observation target (such as the rotation angle of the sail plate), thereby causing the problem of inaccurate space target image reconstruction. SUMMARY
[0005] Therefore, it is necessary to provide a spatial target image reconstruction method, device and equipment for a dynamic observation scene, which can better preserve the original multi-view geometric relationship of an observation target and realize efficient HDR image reconstruction in a dynamic scene.
[0006] A spatial target image reconstruction method for a dynamic observation scene, the method comprising:
[0007] obtaining a dynamic sequence multi-exposure data set of a spatial target, wherein the dynamic sequence multi-exposure data set comprises a plurality of track intersection observation image sequences of a plurality of spatial targets under different initial poses.
[0008] sequentially dividing the sequence images in the dynamic sequence multi-exposure data set into a plurality of groups of low dynamic range sub-image sequences, each of the low dynamic range sub-image sequences comprising three sequence images with different spatial target poses and different exposure levels, and generating corresponding static sub-image sequences according to each group of low dynamic range sub-image sequences.
[0009] training a spatial target image reconstruction network using the plurality of groups of low dynamic range sub-image sequences and the corresponding static sub-image sequences to obtain a trained spatial target image reconstruction network, the spatial target image reconstruction network comprising a feature extraction unit, a feature deformable alignment unit and a feature fusion reconstruction unit, wherein the feature extraction unit transforms each of the sequence images from an image domain to a corresponding feature space to obtain a feature map, the feature deformable alignment unit aligns a non-reference image feature map with a reference image feature map to generate a corresponding aligned feature map to eliminate and weaken the influence of spatial information of a dynamic scene on fusion effect, and the feature fusion reconstruction unit fuses and reconstructs the aligned feature maps corresponding to different exposure levels to obtain a high dynamic range image.
[0010] obtaining observation sequence images of a spatial target in a dynamic observation scene, and generating a high dynamic range image of the spatial target according to the observation sequence images using the trained spatial target image reconstruction network.
[0011] In one embodiment, when the dynamic sequence multi-exposure data set is constructed:
[0012] obtaining three-dimensional models of different types of spatial targets, and simulating based on the three-dimensional models, adjusting the pose of the three-dimensional model at an initial position and performing rolling exposure rendering to obtain a multi-exposure image sequence of the same spatial target under different initial poses.
[0013] In one embodiment, generating corresponding static sub-image sequences according to each group of low dynamic range sub-image sequences comprises:
[0014] The spatial target pose in the sequence image located in the middle of the low dynamic range sub-image sequence is taken as the pose of the spatial target in three sequence images in the corresponding static sub-image sequence, and meanwhile, the three sequence images in the static sub-image sequence retain the exposure degree of the three sequence images in the corresponding low dynamic range sub-image sequence.
[0015] In one embodiment, in the feature extraction unit:
[0016] The low dynamic range sub-image sequence in the high dynamic range image domain is subjected to gamma transformation correction according to the sequence image located in the middle to obtain a low dynamic range sub-image sequence in the high dynamic range image domain;
[0017] Each image in the low dynamic range sub-image sequence in the high dynamic range image domain is connected into a multi-channel low dynamic range sub-image sequence according to the channel dimension;
[0018] Stride convolution is introduced to the multi-channel low dynamic range sub-image sequence to perform pyramid feature extraction to obtain feature maps corresponding to each sequence image.
[0019] In one embodiment, in the feature deformable alignment unit:
[0020] The feature map corresponding to the sequence image with the exposure time located in the middle of the low dynamic range sub-image sequence is taken as the reference feature map, and the other two feature maps are taken as non-reference feature maps;
[0021] Based on the reference feature map, the bias information of the other two non-reference feature maps is calculated respectively, and the non-reference feature maps and the corresponding bias information are subjected to deformable convolution to obtain aligned feature maps after alignment;
[0022] The reference feature map is directly taken as the aligned feature map.
[0023] In one embodiment, in the feature fusion reconstruction unit:
[0024] After the aligned feature maps are combined, DenseNet network is used for feature fusion and reconstruction to obtain the high dynamic range image.
[0025] In one embodiment, when the spatial target image reconstruction network is trained, the loss function used includes a structural similarity loss function, a mean square error loss function, an alignment loss function and a feature point extraction loss function;
[0026] The alignment loss function is calculated from the feature maps obtained by using an encoder to extract features from each low dynamic range image in the static sub-image sequence and the corresponding aligned feature maps;
[0027] The feature point extraction loss function is calculated by the low dynamic sub-image sequence and the corresponding reconstructed high dynamic range image using a scale invariant feature transform operator.
[0028] The application also provides a device for reconstructing an image of a space target in a dynamic observation scene, the device comprising:
[0029] A dynamic sequence multi-exposure data set acquisition module is configured to acquire a dynamic sequence multi-exposure data set of a space target, wherein the dynamic sequence multi-exposure data set comprises a plurality of orbit intersection observation image sequences of the space target under different initial attitudes and rolling exposure.
[0030] A static sub-image sequence generation module is configured to divide the sequence images in the dynamic sequence multi-exposure data set into a plurality of groups of low dynamic sub-image sequences one by one, each of the low dynamic sub-image sequences comprising three sequence images with different space target attitudes and different exposure levels from low to high, and to generate corresponding static sub-image sequences according to each group of low dynamic sub-image sequences.
[0031] A space target image reconstruction network training module is configured to divide the sequence images in the dynamic sequence multi-exposure data set into a plurality of groups of low dynamic range sub-image sequences one by one, each of the low dynamic range sub-image sequences comprising three sequence images with different space target attitudes and different exposure levels from low to high, and to generate corresponding static sub-image sequences according to each group of low dynamic range sub-image sequences.
[0032] A dynamic target image reconstruction module is configured to acquire observation sequence images of a space target in a dynamic observation scene, and to generate a high dynamic range image of the space target according to the observation sequence images by using the trained space target image reconstruction network.
[0033] A computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:
[0034] A dynamic sequence multi-exposure data set of a space target is acquired, wherein the dynamic sequence multi-exposure data set comprises a plurality of orbit intersection observation image sequences of the space target under different initial attitudes and rolling exposure.
[0035] The sequence images in the dynamic sequence multi-exposure data set are divided into a plurality of groups of low dynamic range sub-image sequences one by one, each of the low dynamic range sub-image sequences comprising three sequence images with different space target attitudes and different exposure levels from low to high, and corresponding static sub-image sequences are generated according to each group of low dynamic range sub-image sequences.
[0036] The spatial target image reconstruction network is trained by using a plurality of groups of low dynamic range sub-image sequences and corresponding static sub-image sequences, to obtain a trained spatial target image reconstruction network, the spatial target image reconstruction network comprising a feature extraction unit, a feature deformable alignment unit and a feature fusion reconstruction unit, wherein the feature extraction unit transforms each of the sequence images from an image domain to a corresponding feature space to obtain a feature map, the feature deformable alignment unit aligns a non-reference image feature map with a reference image feature map to generate a corresponding aligned feature map, so as to eliminate and weaken the influence of spatial information of a dynamic scene on the fusion effect, and the feature fusion reconstruction unit performs information fusion and reconstruction on the aligned feature maps corresponding to different exposure levels to obtain a high dynamic range image.
[0037] The observation sequence images of a spatial target in a dynamic observation scene are acquired, and a high dynamic range image of the spatial target is generated according to the observation sequence images by using the trained spatial target image reconstruction network.
[0038] A computer readable storage medium having stored thereon a computer program, the computer program being executed by a processor to implement the following steps:
[0039] A dynamic sequence multi-exposure data set of a spatial target is acquired, wherein the dynamic sequence multi-exposure data set comprises a plurality of orbit intersection observation image sequences of a plurality of spatial targets under different initial attitudes and rolling exposure.
[0040] The sequence images in the dynamic sequence multi-exposure data set are sequentially divided into a plurality of groups of low dynamic range sub-image sequences in groups of three, each of the low dynamic range sub-image sequences comprising three sequence images with different exposure levels and different spatial target attitudes, and a corresponding static sub-image sequence is generated according to each group of low dynamic range sub-image sequences.
[0041] The spatial target image reconstruction network is trained by using a plurality of groups of low dynamic range sub-image sequences and corresponding static sub-image sequences, to obtain a trained spatial target image reconstruction network, the spatial target image reconstruction network comprising a feature extraction unit, a feature deformable alignment unit and a feature fusion reconstruction unit, wherein the feature extraction unit transforms each of the sequence images from an image domain to a corresponding feature space to obtain a feature map, the feature deformable alignment unit aligns a non-reference image feature map with a reference image feature map to generate a corresponding aligned feature map, so as to eliminate and weaken the influence of spatial information of a dynamic scene on the fusion effect, and the feature fusion reconstruction unit performs information fusion and reconstruction on the aligned feature maps corresponding to different exposure levels to obtain a high dynamic range image.
[0042] An observation sequence image of a spatial target in a dynamic observation scene is acquired, and a high dynamic range image of the spatial target is generated from the observation sequence image by using a trained spatial target image reconstruction network.
[0043] The spatial target image reconstruction method, device and equipment described above, by sequentially dividing the sequence images in the dynamic sequence multi-exposure dataset into multiple groups of low dynamic range sub-image sequences, each group of low dynamic range sub-image sequences including three sequence images with different spatial target poses and different exposure levels, and generating corresponding static sub-image sequences according to each group of low dynamic range sub-image sequences, training the spatial target image reconstruction network by using the multiple groups of low dynamic range sub-image sequences and the corresponding static sub-image sequences, obtaining the trained spatial target image reconstruction network, and using the trained spatial target image reconstruction network to generate the high dynamic range image of the spatial target from the observation sequence image of the spatial target in the dynamic observation scene, the method can realize efficient HDR image reconstruction in a dynamic scene. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 A flowchart of the spatial target image reconstruction method in a dynamic observation scene for an embodiment;
[0045] Figure 2 A simulation step diagram of the spatial target multi-exposure image sequence for an embodiment;
[0046] Figure 3 A diagram of part of the sequence images in the dynamic sequence multi-exposure data for an embodiment;
[0047] Figure 4 A structural diagram of the spatial target image reconstruction network for an embodiment;
[0048] Figure 5 A diagram of the low dynamic range sub-image sequence and the static sub-image sequence for an embodiment;
[0049] Figure 6 A comparison diagram of the target HDR images obtained by using the method and other methods in an experiment;
[0050] Figure 7 A comparison diagram of the SIFT feature point matching effect in the target HDR images obtained by using the method and other methods for an embodiment;
[0051] Figure 8 A structural block diagram of the spatial target image reconstruction device in a dynamic observation scene for an embodiment;
[0052] Figure 9 Figure 1 is a diagram of the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0053] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.
[0054] In existing space-based optical imaging technology, the traditional dynamic multi-exposure fusion method based on removing moving pixels cannot handle complex target overlapping areas and will lose a lot of effective information. At the same time, due to the particularity of the subsequent tasks of space target situation awareness technology, such as three-dimensional reconstruction, attitude estimation, etc., the multi-view geometric relationship of the observed target changes very sensitively, and the traditional method based on registration is easy to cause the change of the observed target representation (such as the turning angle of the sail plate, etc.). Further, when deep learning is used for space target image reconstruction, a method is proposed to use convolutional neural network to solve the ghost problem in dynamic scene multi-exposure fusion, but most of them use supervised learning mode, which needs true value HDR image as guidance. The unsupervised method based on GAN network may appear pseudo-texture when reconstructing, and the unsupervised method based on CNN lacks a module for processing a large number of moving pixels in the learning framework, and usually uses a direct addition strategy for feature maps when fusing and reconstructing, which is not suitable for the case where there is a complex overlapping area between space image frames. At the same time, the existing deep learning multi-exposure general dataset camera and static background are in a relatively static state, which is quite different from the imaging scene of the space target image sequence, and the open source space target dataset is generated based on a single exposure condition, both of which cannot meet the actual application requirements of space target image sequence multi-exposure fusion without true value in real scene of space situation awareness.
[0055] In view of the above problems, as shown in Figure 1 A space target image reconstruction method in a dynamic observation scene is provided, which specifically includes the following steps:
[0056] Step S100, acquiring a dynamic sequence multi-exposure dataset of a space target, wherein the dynamic sequence multi-exposure dataset includes a plurality of orbit intersection observation image sequences of the space target under different initial attitudes.
[0057] Step S110, dividing the sequence images in the dynamic sequence multi-exposure dataset into a plurality of groups of low dynamic range sub-image sequences in turn in groups of three, each of the low dynamic range sub-image sequences including three sequence images with different space target attitudes and exposure levels from low to high, and generating corresponding static sub-image sequences according to each group of low dynamic range sub-image sequences.
[0058] Step S120, training the spatial target image reconstruction network by using multiple sets of low dynamic range sub-image sequences and corresponding static sub-image sequences, to obtain a trained spatial target image reconstruction network, the spatial target image reconstruction network comprising a feature extraction unit, a feature deformable alignment unit and a feature fusion reconstruction unit, wherein the feature extraction unit transforms each sequence image from an image domain to a corresponding feature space to obtain a feature map, the feature deformable alignment unit aligns the non-reference image feature map with the reference image feature map to generate a corresponding aligned feature map, so as to eliminate and weaken the influence of spatial information of the dynamic scene on the fusion effect, and the feature fusion reconstruction unit fuses and reconstructs the aligned feature maps corresponding to different exposure levels to obtain a high dynamic range image.
[0059] Step S130, obtaining observation sequence images of a spatial target in a dynamic observation scene, and generating a high dynamic range image of the spatial target according to the observation sequence images by using the trained spatial target image reconstruction network.
[0060] In the embodiment, an unsupervised multi-exposure fusion method for space target space-based optical images in a dynamic observation scene is proposed, which realizes efficient HDR image reconstruction in a dynamic scene while well preserving the original multi-view geometric relationship of the observed target. An unsupervised learning framework suitable for multi-exposure fusion of space target space-based optical images is constructed, and a new space target rolling exposure dataset is also proposed, which uses blender to simulate the observation and imaging of space targets at the intersection of orbits, to solve the problem of insufficient training data during neural network training.
[0061] In steps S100 and S110, the process of constructing a training dataset is performed, wherein the spatial target is a man-made satellite.
[0062] In step S100, when constructing the dynamic sequence multi-exposure dataset: three-dimensional models of different types of space targets are obtained, and simulation is performed based on the three-dimensional models, and by adjusting the attitude of the three-dimensional model at the initial position and performing rolling exposure, multiple sequence images of the same space target under different attitudes are obtained.
[0063] Considering that it is difficult to obtain actual observation images of space targets, most space target image datasets are generated by software rendering or semi-physical simulation, such as the BUAA-SID dataset and the SPEED+ dataset. However, the existing simulation datasets do not consider the actual situation of imaging the observed target by the rolling exposure mechanism at the natural intersection of orbits. At the same time, the lighting conditions of the observed target change when it is in different regions of the orbit, and the reflectivity of different components of the observed target also differs.
[0064] Therefore, in the embodiment, a track intersection observation space target multi-exposure image dataset is created in a simulated manner, and a simulation step of generating a space target image is as shown in Figure 2 .
[0065] Since the intersection time of the imaging platform and the observation target is short in actual cases, the attitude of the space target does not change much from the beginning of observation to the end of observation. In order to obtain as much training data as possible, a multi-exposure image sequence of different attitudes of the same target is obtained by adjusting the initial position of the space target. This dataset is named "Multi-exposure Spacecraft Image Dataset" (MES Dataset), that is, a dynamic sequence multi-exposure dataset, and part of the images are as shown in Figure 3 .
[0066] Further, in order to adapt to the training of the subsequent neural network, the space target is reconstructed according to three sequence images of different exposure levels and different attitudes of the same space target in the neural network, and in step S110, the sequence images in the dynamic sequence multi-exposure dataset are divided into multiple groups of low dynamic range sub-image sequences one by one, each low dynamic range sub-image sequence includes three sequence images with different exposure levels from low to high and different attitudes of the space target. Referring to the multiple groups of low dynamic range sub-image sequences as shown in Figure 3 , for example, three low, medium, and high exposure images in the same row at the initial intersection observation time, and it can be further seen that the attitudes of the space targets in the three images are also different.
[0067] In one embodiment, in order to better train the neural network, the dynamic sequence multi-exposure dataset includes 120 rolling exposure image sequences of 12 space target models in 4 different attitudes, a total of 48 image sequences, including 1920 multi-exposure pairs, that is, 5760 images. In this embodiment, in order to train the neural network subsequently, the corresponding static sub-image sequences are also generated according to each group of low dynamic range sub-image sequences to calculate the alignment loss function, which will be introduced hereinafter.
[0068] Specifically, the attitude of the space target in the sequence image in the middle of the low dynamic range sub-image sequence is taken as the attitude of the space target in the three sequence images in the corresponding static sub-image sequence, and at the same time, the three sequence images in the static sub-image sequence retain the exposure levels of the three sequence images in the corresponding low dynamic range sub-image sequence. In this way, each group of low dynamic range sub-image sequences will correspond to a group of static sub-image sequences.
[0069] In this application, an unsupervised learning neural network framework, that is, a space target image reconstruction network, is also proposed, and the structure is as shown in Figure 4As shown. From the perspective of function, this network consists of three parts: feature extraction unit, feature deformable alignment unit and feature fusion reconstruction unit. The feature extraction unit transforms the image from the image domain to the corresponding feature space, the feature deformable alignment unit is responsible for aligning the extracted feature map one by one, eliminating and weakening the influence of spatial information of dynamic scene on the fusion effect, and finally, the feature fusion reconstruction unit realizes the information fusion and reconstruction of the aligned features corresponding to different exposures, and outputs the final HDR image.
[0070] In this embodiment, in the feature extraction unit: the low dynamic range sub-image sequence in the high dynamic range image domain is corrected by gamma transformation according to the medium exposure time image, then each image in the low dynamic range sub-image sequence in the high dynamic range image domain is connected in the channel dimension to form a multi-channel low dynamic range sub-image sequence, and the multi-channel low dynamic range sub-image sequence is introduced Stride convolution for pyramid feature extraction to obtain the feature map corresponding to each sequence image.
[0071] Specifically, generally speaking, although the LDR image sequence can well distinguish the exposure degree and saturation area of the image, it is limited in measuring the motion area. The image in the linear space of the HDR domain simulates the real illumination field and places the image sequence scene in a unified measurement, so it has absolute advantages in identifying uncalibrated and motion areas, therefore, the input dynamic scene LDR rolling exposure space target space optical image ordered sequence, i.e. low dynamic range sub-image sequence The gamma transformation correction is applied, and the following formula is used:
[0072]
[0073] The corresponding image in the HDR domain is obtained by formula (1) Where t i represents the exposure time of the i-th image, and γ represents the gamma correction parameter.
[0074] Then, each image in the image sequence is inputted to be connected in the channel dimension to form a 6-channel input sequence Where [·;·] represents the channel dimension connection operator.
[0075] Further, the input sequence Stride convolution is introduced for pyramid feature extraction, which is converted into a 64-channel feature map, denoted as:
[0076]
[0077] In the feature deformable alignment unit in this embodiment: taking the feature map corresponding to the middle sequence image in the low dynamic range sub-image sequence as the reference feature map, and taking the other two feature maps as non-reference feature maps, based on the reference feature map, the bias information of the other two non-reference feature maps is calculated respectively, the non-reference feature maps and the corresponding bias information are subjected to deformable convolution to obtain the spatially aligned aligned feature maps, and the reference feature map is directly taken as the aligned feature map.
[0078] The feature deformable alignment unit in the method introduces deformable convolution, so as to be suitable for the complex pose changes of the observed target in the space image, and to retain the multi-view geometric information consistent with the reference image as much as possible in the HDR reconstruction. The deformable convolution has the advantage of arbitrary sampling, and can realize the correction of basic geometric deformations such as translation, rotation, scaling, skewing and perspective, and is suitable for solving the problem of non-matching of the image sequence of the observed target in the space target monitoring scene. Meanwhile, the anisotropic bias information can fit a dense optical flow field, and through the direction and amplitude of the bias, the direction and amplitude of the object deformation are adjusted, and the non-rigid deformation situation is solved.
[0079] Further, the feature deformable alignment unit calculates the bias information O i of the non-reference feature E r (i.e. E2) together with the reference feature E i 1), and sends the above information (i.e. the non-reference feature map and the corresponding bias information) into the corresponding deformable convolution dc i (·) to calculate the spatially aligned aligned feature map A i :
[0080]
[0081] In formula (3), p is the position in the feature map, ω j is the weight in the convolution kernel Ω, Δp j and Δm j respectively represent the learnable bias and the modulation intensity of the jth position in the convolution kernel Ω.
[0082] In the feature fusion reconstruction unit in this embodiment: after merging each aligned feature map, a DenseNet network is used for feature fusion and reconstruction to obtain a high dynamic range image.
[0083] Further, through the feature extraction unit and the feature deformable alignment unit, the feature maps corresponding to each exposure length are obtained, and the texture information contained in the adjacent frames is also rich and important for the dynamic information of the fixed exposure length lens. Considering that the typical exposure step is 3, the concatenation operation is used as the feature merging method to obtain the merged feature E c =[A1;E r; A3] and sent to DenseNet for fusion reconstruction to obtain the final HDR image, denoted as:
[0084] H = DenseNet(E c ) (4)
[0085] In this embodiment, when training the spatial target image reconstruction network, the loss function adopted includes a structural similarity index loss function, a mean square error loss function, an alignment loss function and a feature point extraction loss function, denoted as:
[0086]
[0087] In formula (5), θ represents the parameters in the network, D represents the training data set, and α, β, λ and μ are balance parameters greater than 0.
[0088] Due to the lack of true value images, the SSIM is a structural similarity index (Structural Similarity Index Metric, SSIM) that focuses on contrast and structural changes, and the MSE is a mean square error (Mean Squared Error, MSE) that focuses on intensity distribution difference constraints. Therefore, the similarity constraints of the fusion image and the input image are realized from two aspects of structural similarity and intensity distribution:
[0089]
[0090]
[0091] In formula (6) and formula (7), ω i The calculation of the weight is calculated by using the VGG network to extract information and save the degree of information, and then normalized by the softmax function to make The weight calculation part in Figure 4 can be expressed as f(·), and the weight The calculation process of the weight
[0092]
[0093] Further, the method also proposes a corresponding feature alignment loss function for the feature deformable alignment unit, introduces a deformable convolution to fit the complex geometric deformation of the observed target, and extracts the image exposure information while not changing the representation of the observed target in the reference image. At the same time, a loss function based on a SIFT operator is designed to optimize the fusion result of the network, so as to obtain better feature extraction effect and make it more suitable for attitude estimation and other tasks in the space detection system.
[0094] In this embodiment, in order to achieve better registration effect of spatial target image information and retain its multi-view geometric information as strictly as possible, a feature alignment loss function is also proposed. The function is calculated by using an encoder to extract features from each sequence image in the static sub-image sequence and the corresponding aligned feature map.
[0095] Specifically, if it is known that the input dynamic LDR image sequence, i.e., the low dynamic range sub-image sequence {I1, I2, I3}, corresponds to the static LDR image sequence, i.e., the static sub-image sequence Right now, The observed target pose is consistent with that of the I2 image, and the exposure level is the same as that of the I i Therefore, it is only necessary to limit the image features in dynamic scenes. After correction by the feature deformable alignment unit, its output feature The features extracted by the encoder from the corresponding static image sequence Get as close as possible to achieve feature alignment:
[0096]
[0097] However, for existing dynamic datasets, a large amount of data does not contain corresponding static scene data. Therefore, most of them proposed a method of using reference images for exposure conversion to generate pseudo static scene image sequences. And use the self-supervision method to learn and update the network parameters. However, when preparing MESDataset, this method can directly produce a static dataset, ensuring the consistency of image information and exposure level. Take the rendered image of the Dawn model as an example, as shown in Figure 5, where Figure 5 (a)-(c) is a set of low dynamic range sub-image sequences, Figure 5 (d)-(f) are the corresponding static sub-image sequences.
[0098] Therefore, the alignment loss function can be expressed as:
[0099]
[0100] In this embodiment, a feature point extraction loss function is proposed to optimize the computer processing of HDR reconstruction results for tasks such as subsequent pose estimation in spatial detection. This loss function is constructed using the SIFT (Scale Invariant Feature Transform) operator based on a sequence of low dynamic range sub-images and the corresponding reconstructed high dynamic range images.
[0101] The SIFT operator is an algorithm for feature extraction in image processing and computer vision, which can detect key points in different scale spaces and is invariant to image rotation, scaling, and brightness changes. The SIFT operator is used for key point detection on each exposure image in the network. These key points should have unique local features and be insensitive to exposure changes.
[0102] The SIFT feature point set and the number of each set of dynamic LDR image sequences {I1, I2, I3}, i.e., low dynamic range sub-image sequences, and the corresponding reconstructed HDR image H are calculated using the following formula:
[0103]
[0104] K H = SIFT(H) (12)
[0105]
[0106] n H = num(K H ) (14)
[0107] In formula (11) to (14), SIFT(·) represents the SIFT function for extracting image feature points, represents the set of feature points in the i-th LDR image, K H represents the set of feature points in the HDR image, num(·) is used to calculate the number of feature points in the set, represents the number of feature points in the i-th LDR image, n H represents the number of feature points in the HDR image.
[0108] Further, in order to make the reconstructed spatial target HDR image have better computer processing effect in subsequent spatial monitoring tasks, the fused image should contain as many non-repeated feature points as possible, therefore, the intersection of the feature point sets of the LDR images is calculated and repeated detection is performed, and the number of feature points is compared with the set of feature points of the HDR image that has undergone repeated detection. Therefore, the feature point extraction loss function is constructed as follows:
[0109]
[0110] In formula (15), rd(·) represents repeated detection of the feature point set, represents the set of feature points with the largest number of feature points after repeated detection of each feature point set of the LDR sequence. From the numerical calculation, it can be understood that when the number of non-repeated feature points of the HDR image is the same as the number of non-repeated feature points in the LDR image, When the number of non-repeated feature points of the HDR image is equal to the number of all non-repeated feature points in the LDR image, That is, the less the number of non-repeated feature points of the HDR image, the closer the value of the loss function to 1, and the more the number of non-repeated feature points of the HDR image, the closer the value of the loss function to 0.
[0111] In this paper, the effectiveness of the method is also proved by experimental results. In the experiment, 16 rolling exposure observation sequences of 4 space targets in the MES dataset are used, a total of 640 groups of multi-exposure images with a size of 800*800.
[0112] Figure 6 and Figure 7 The fusion results and SIFT feature point extraction effects of the method and the comparative algorithm on the MES dataset test set are respectively shown.
[0113] From Figure 6 It can be seen that when the multi-exposure fusion of the space target space-based optical image sequence is performed on the MES test set, in addition to the SPD-MEF algorithm and the method proposed in this paper, the HDR reconstruction results of other comparative algorithms all have serious artifact phenomena at the edge of the sailboard and the important components in the middle. However, the brightness change of the fusion results of the SPD-MEF algorithm at the sailboard and the main body appears uneven and incoherent, while the fusion results of the method proposed in this paper retain higher contrast and texture details.
[0114] From Figure 7 It can be seen that compared with the comparative algorithm, the method proposed in this paper can obtain better feature matching effect, especially at the outer sailboard with large attitude change, more feature points can be extracted, and at the same time the subjective visual effect is retained, better computer processing results can be provided for subsequent space situation awareness tasks.
[0115] In the above space target image reconstruction method in the dynamic observation scene, a space target space-based optical image multi-exposure fusion method based on unsupervised learning in the dynamic monitoring scene is proposed, and a space target rolling exposure dataset is proposed. The method can realize high-quality HDR space target image reconstruction in a dynamic scene. By introducing deformable convolution in the feature alignment module and constructing a feature alignment loss function and a feature point extraction loss function, the original multi-view geometric relationship of the observed target can be effectively retained, and better feature extraction effect can be obtained to serve subsequent tasks such as attitude estimation in the space situation awareness system.
[0116] It should be understood that although Figure 1The steps in the flowchart are shown in sequence as indicated by the arrows, but the steps are not necessarily performed in the order indicated by the arrows. Unless otherwise explicitly stated herein, there is no strict order of execution of the steps, and the steps can be performed in other orders. Also, Figure 1 At least some of the steps in the flowchart can include multiple sub-steps or multiple stages, which are not necessarily performed at the same time, but can be performed at different times, and the order of the sub-steps or stages is not necessarily sequential, but can be performed in rotation or alternation with other steps or sub-steps or stages of other steps.
[0117] In one embodiment, as shown in Figure 8 Fig. 1, there is provided an image reconstruction device for spatial targets in a dynamic observation scene, comprising: a dynamic sequence multi-exposure data acquisition module 200, a static sub-image sequence generation module 210, a spatial target image reconstruction network training module 220, and a dynamic target image reconstruction module 230, wherein:
[0118] The dynamic sequence multi-exposure data set acquisition module 200 is configured to acquire a dynamic sequence multi-exposure data set of spatial targets, wherein the dynamic sequence multi-exposure data set includes multiple orbit intersection observation image sequences of multiple spatial targets under different initial attitudes.
[0119] The static sub-image sequence generation module 210 is configured to divide the sequence images in the dynamic sequence multi-exposure data set three at a time into multiple groups of low dynamic sub-image sequences, each of the low dynamic sub-image sequences including three sequence images with different spatial target attitudes and exposure levels from low to high, and to generate corresponding static sub-image sequences according to each group of low dynamic sub-image sequences.
[0120] The spatial target image reconstruction network training module 220 is configured to divide the sequence images in the dynamic sequence multi-exposure data set three at a time into multiple groups of low dynamic range sub-image sequences, each of the low dynamic range sub-image sequences including three sequence images with different spatial target attitudes and exposure levels from low to high, and to generate corresponding static sub-image sequences according to each group of low dynamic range sub-image sequences.
[0121] The dynamic target image reconstruction module 230 is configured to acquire observation sequence images of spatial targets in a dynamic observation scene, and to generate a high dynamic range image of the spatial target according to the observation sequence images using the trained spatial target image reconstruction network.
[0122] The specific limitations of the image reconstruction device for space targets in a dynamic observation scene can refer to the limitations of the image reconstruction method for space targets in a dynamic observation scene, which will not be repeated here. Each module in the above image reconstruction device for space targets in a dynamic observation scene can be realized by software, hardware, and combinations thereof. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor calls and executes the operations corresponding to each module.
[0123] In one embodiment, a computer device, which can be a terminal, has an internal structure diagram as shown in Figure 9 The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected by a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is used to communicate with external terminals through network connections. The computer program is executed by the processor to implement an image reconstruction method for space targets in a dynamic observation scene. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball, or touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0124] Those skilled in the art can understand that Figure 9 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0125] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the following steps:
[0126] obtaining a dynamic sequence multi-exposure data set of space targets, wherein the dynamic sequence multi-exposure data set includes a plurality of orbit intersection observation image sequences of a plurality of space targets exposed by rolling at different initial attitudes;
[0127] the sequence images in the dynamic sequence multi-exposure data set are sequentially divided into multiple groups of low dynamic range sub-image sequences, each of the low dynamic range sub-image sequences includes three sequence images with different spatial target poses and different exposure levels from low to high, and a corresponding static sub-image sequence is generated according to each group of low dynamic range sub-image sequences;
[0128] The spatial target image reconstruction network is trained by using the multiple groups of low dynamic range sub-image sequences and the corresponding static sub-image sequences, to obtain a trained spatial target image reconstruction network, the spatial target image reconstruction network includes a feature extraction unit, a feature deformable alignment unit and a feature fusion reconstruction unit, wherein the feature extraction unit transforms each of the sequence images from an image domain to a corresponding feature space to obtain a feature map, the feature deformable alignment unit aligns a non-reference image feature map with a reference image feature map to generate a corresponding aligned feature map, so as to eliminate and weaken the influence of spatial information of a dynamic scene on fusion effect, and the feature fusion reconstruction unit performs information fusion and reconstruction on the aligned feature maps corresponding to different exposure levels to obtain a high dynamic range image.
[0129] The observation sequence images of a spatial target in a dynamic observation scene are obtained, and a high dynamic range image of the spatial target is generated according to the observation sequence images by using the trained spatial target image reconstruction network.
[0130] In one embodiment, a computer readable storage medium is provided, and a computer program is stored on the computer readable storage medium, and the computer program is executed by a processor to implement the following steps:
[0131] A dynamic sequence multi-exposure data set of a spatial target is obtained, wherein the dynamic sequence multi-exposure data set includes multiple orbit intersection observation image sequences of multiple spatial targets under different initial poses and rolling exposure;
[0132] The sequence images in the dynamic sequence multi-exposure data set are sequentially divided into multiple groups of low dynamic range sub-image sequences, each of the low dynamic range sub-image sequences includes three sequence images with different spatial target poses and different exposure levels from low to high, and a corresponding static sub-image sequence is generated according to each group of low dynamic range sub-image sequences;
[0133] The spatial target image reconstruction network is trained by using a plurality of low dynamic range sub-image sequences and corresponding static sub-image sequences, to obtain a trained spatial target image reconstruction network, the spatial target image reconstruction network comprising a feature extraction unit, a feature deformable alignment unit and a feature fusion reconstruction unit, wherein the feature extraction unit transforms each of the sequence images from an image domain to a corresponding feature space to obtain a feature map, the feature deformable alignment unit aligns the non-reference image feature map with the reference image feature map to generate a corresponding aligned feature map, so as to eliminate and weaken the influence of spatial information of a dynamic scene on the fusion effect, and the feature fusion reconstruction unit fuses and reconstructs the aligned feature maps corresponding to different exposure degrees to obtain a high dynamic range image.
[0134] The observation sequence images of a spatial target in a dynamic observation scene are acquired, and the trained spatial target image reconstruction network is used to generate a high dynamic range image of the spatial target according to the observation sequence images.
[0135] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the processes of the above-mentioned embodiments can be included. Any reference to a memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM) and the like.
[0136] The technical features of the above embodiments can be combined in any manner. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the present disclosure.
[0137] The above-described embodiments are merely illustrative of several embodiments of the present application, which are described in more detail and in a specific manner, but should not be construed as limiting the scope of the patent. It should be noted that for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for reconstructing a space target image in a dynamic observation scene, characterized in that: The method comprises: Acquire a dynamic sequence multi-exposure dataset of a space target, wherein the dynamic sequence multi-exposure dataset includes a plurality of orbital intersection observation image sequences of a plurality of space targets under rolling exposures at different initial postures; Sequentially dividing the sequence images in the dynamic sequence multi-exposure dataset into a plurality of low dynamic range sub-image sequences in groups of three, each of the low dynamic range sub-image sequences including three sequence images with varying exposure levels from low to high and with different spatial object poses, and generating a corresponding static sub-image sequence based on each group of low dynamic range sub-image sequences; A spatial target image reconstruction network is trained using multiple sets of low dynamic range sub-image sequences and corresponding static sub-image sequences to obtain a trained spatial target image reconstruction network, wherein the spatial target image reconstruction network includes a feature extraction unit, a feature deformable alignment unit, and a feature fusion reconstruction unit, wherein the feature extraction unit transforms each of the sequence images from the image domain to the corresponding feature space to obtain a feature map, the feature deformable alignment unit aligns the non-reference image feature map with the reference image feature map to generate corresponding aligned feature maps to eliminate and weaken the influence of the spatial information of the dynamic scene on the fusion effect, and the feature fusion reconstruction unit fuses and reconstructs the aligned feature maps corresponding to different exposure levels to obtain a high dynamic range image; An observation sequence image of a space target in a dynamic observation scene is obtained, and a high dynamic range image of the space target is generated according to the observation sequence image using a trained space target image reconstruction network.
2. The space target image reconstruction method according to claim 1, characterized in that: When constructing the dynamic sequence multi-exposure dataset: Three-dimensional models of different types of space targets are obtained, and simulation is performed based on the three-dimensional models. By adjusting the posture of the three-dimensional model at the initial position and performing rolling exposure rendering, a multi-exposure image sequence corresponding to the same space target in different initial postures is obtained.
3. The space target image reconstruction method according to claim 2, characterized in that: Generating a corresponding static sub-image sequence according to each group of low dynamic range sub-image sequences includes: The posture of the space target in the sequence image located in the middle of the low dynamic range sub-image sequence is used as the posture of the space target in the three sequence images in the corresponding static sub-image sequence. At the same time, the three sequence images in the static sub-image sequence retain the exposure level of the three sequence images in the corresponding low dynamic range sub-image sequence.
4. The space target image reconstruction method according to claim 3, characterized in that: In the feature extraction unit: Correcting the low dynamic range sub-image sequence by gamma transformation according to the sequence image located in the middle, to obtain a low dynamic range sub-image sequence in a high dynamic range image domain; Connect each image in the low dynamic range sub-image sequence in the high dynamic range image domain into a multi-channel low dynamic range sub-image sequence according to the channel dimension; Stride convolution is introduced into the multi-channel low dynamic range sub-image sequence to perform pyramid feature extraction, and a feature map corresponding to each sequence image is obtained.
5. The space target image reconstruction method according to claim 4, characterized in that: In the feature deformable alignment unit: The feature map corresponding to the medium exposure time image in the low dynamic range sub-image sequence is used as the reference image feature map, and the other two feature maps are used as the non-reference image feature maps; Based on the reference image feature map, the bias information of the other two non-reference image feature maps is calculated respectively, and the non-reference image feature maps and the corresponding bias information are subjected to deformable convolution to obtain spatially aligned aligned feature maps; The reference image feature map is directly used as the aligned feature map.
6. The space target image reconstruction method according to claim 5, characterized in that: In the feature fusion reconstruction unit: After merging the aligned feature maps, the DenseNet network is used to perform feature fusion and reconstruction to obtain the high dynamic range image.
7. The space target image reconstruction method according to any one of claims 1 to 6, characterized in that: When training the spatial target image reconstruction network, the loss functions used include a structural similarity loss function, a mean square error loss function, an alignment loss function, and a feature point extraction loss function; The alignment loss function is calculated by using an encoder to extract features from each low dynamic range image in the static sub-image sequence, and the corresponding aligned feature map; The feature point extraction loss function is calculated by using a scale-invariant feature transformation operator on the low dynamic range sub-image sequence and the corresponding reconstructed high dynamic range image.
8. A device for reconstructing images of space targets in dynamic observation scenes, characterized in that: The device comprises: A dynamic sequence multi-exposure dataset acquisition module is used to acquire a dynamic sequence multi-exposure dataset of a space target, wherein the dynamic sequence multi-exposure dataset includes a plurality of orbital intersection observation image sequences of rolling exposures of a plurality of space targets at different initial postures; a static sub-image sequence generation module, configured to sequentially divide the sequence images in the dynamic sequence multi-exposure dataset into a plurality of low dynamic range sub-image sequences in groups of three, each of the low dynamic range sub-image sequences including three sequence images with varying exposure levels from low to high and with different spatial target poses, and to generate a corresponding static sub-image sequence based on each group of low dynamic range sub-image sequences; A spatial target image reconstruction network training module is used to train the spatial target image reconstruction network using multiple sets of low dynamic range sub-image sequences and corresponding static sub-image sequences to obtain a trained spatial target image reconstruction network, wherein the spatial target image reconstruction network includes a feature extraction unit, a feature deformable alignment unit, and a feature fusion reconstruction unit, wherein the feature extraction unit transforms each of the sequence images from the image domain to the corresponding feature space to obtain a feature map, the feature deformable alignment unit aligns the non-reference image feature map with the reference image feature map to generate a corresponding aligned feature map to eliminate and weaken the influence of the spatial information of the dynamic scene on the fusion effect, and the feature fusion reconstruction unit fuses and reconstructs the aligned feature maps corresponding to different exposure levels to obtain a high dynamic range image; The dynamic target image reconstruction module is used to obtain the observation sequence images of the space target in the dynamic observation scene, and use the trained space target image reconstruction network to generate a high dynamic range image of the space target based on the observation sequence images.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Dynamic scene HDR reconstruction method based on deep learning
CN111242883A
Dynamic scene three-dimensional reconstruction method, apparatus and system, server, and medium
WO2019161813A1