An image dynamic acquisition method for spatio-temporal-frequency complementarity of a scene
By presetting multiple image parameters in traditional optical cameras and building a fully connected neural network, dynamically adjusting the camera parameters to obtain the initial image sequence and the optimal focus distance, the problem that traditional cameras cannot actively obtain all the information of the scene is solved, and more efficient dynamic image acquisition is achieved.
Patent Information
- Application Number
- CN202510413766.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-03
AI Technical Summary
Traditional optical cameras cannot obtain all information in the scene by actively adjusting camera parameters in complex environments, and cannot effectively utilize space-time frequency domain information, resulting in incomplete acquisition of scene information.
By presetting the initial value of camera imaging parameters of multiple components of the imaging system, the initial image sequence is obtained, the optimal focus distance and information fusion weighted graph is determined, and a fully connected neural network is built to predict camera imaging parameters, and dynamic adjustment and image acquisition are realized.
It realizes the active acquisition of missing scene information under dynamically changing imaging conditions, maximizes the acquisition of all the information in the scene, and improves the efficiency and effect of dynamic image acquisition.
Smart Images

Figure CN119919617B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of optoelectronic imaging, and particularly to an image dynamic acquisition method for spatio-temporal-frequency complementarity of a scene. Background Art
[0002] Due to its high imaging quality, stable system operation and other characteristics, traditional optical cameras have been widely used in space environment imaging tasks. The way to obtain scene information mainly depends on one or multiple imaging under fixed camera parameters. That is, during the imaging process, parameters such as the exposure time, gain, aperture, and focal length of the camera are all set values based on manual experience. However, when shooting in a complex environment, only a fixed number of images can be taken with fixed camera parameters, and it is impossible to maximize the acquisition of information in the scene by actively adjusting the camera parameters. Moreover, the method of obtaining scene information based on image fusion only utilizes spatial domain information and cannot effectively extract and fuse time domain, spatial domain, and frequency domain information, resulting in limited ability to obtain scene information. Therefore, the traditional method of a camera obtaining scene information is single or multiple fixed shootings with fixed parameters, which not only cannot actively obtain the missing scene information by dynamically controlling the camera parameters, but also cannot effectively utilize spatio-temporal-frequency domain information, leading to problems such as incomplete acquisition of scene information and poor environmental adaptability of the camera.
[0003] Document 1 proposed a deep perception enhancement network for multi-exposure fusion, which consists of two modules: image detail extraction and color mapping, and can obtain results with rich detail information and good visual perception. However, this static method will cause problems such as artifacts, uneven brightness and darkness of the image, and incomplete acquisition of effective scene information due to factors such as light changes and target movement in the actual dynamic shooting task, and does not consider the feasibility of actively obtaining dynamic effective information by adjusting camera parameters in the actual shooting process. Document 1 is as follows:
[0004] "Han D, Li L, Guo X J, et al. Multi-exposure image fusion via deepperceptual enhancement[J]. Information Fusion, 2022, 79:248-262." Summary of the Invention
[0005] The purpose of the present invention is to provide an image dynamic acquisition method for spatio-temporal-frequency complementarity of a scene to solve the problem that traditional optical imaging methods cannot actively control the camera to target and shoot missing information in the scene under dynamically changing imaging conditions, thus unable to maximize the acquisition of all information in the scene.
[0006] To achieve the above task, the present invention adopts the following technical solutions:
[0007] An image dynamic acquisition method for spatio - temporal - frequency complementarity of a scene, comprising:
[0008] Presetting initial values of imaging parameters of cameras in multiple imaging systems, and each camera respectively uses each set of initial imaging parameter values to collect images of a target, obtaining an initial image sequence;
[0009] Based on the initial image sequence, determining the optimal focusing distance of the camera; after adjusting the camera with the optimal focusing distance, re - obtaining a new initial image sequence with the imaging parameter initial values;
[0010] For the images in the new initial image sequence, determining an information quantity fusion weighted graph of the images;
[0011] The imaging system continuously collects images of the target, calculates the scene irradiance value corresponding to each image, determines the scene irradiance distribution and obtains a scene irradiance graph, thereby determining the optimal information quantity fusion weighted graph;
[0012] Based on the images that the imaging system has currently captured, calculating the information quantity fusion weighted graph of each image and fusing them; comparing the fusion result with the optimal information quantity fusion weighted graph, and calculating the positive difference of the missing information;
[0013] Constructing a fully - connected neural network, where the input layer of the fully - connected neural network is the average missing information quantity obtained by averaging the positive difference of the missing information over the entire image, and the output is the imaging parameters of the camera; training the fully - connected neural network, and saving the trained fully - connected neural network for predicting the imaging parameters of the camera.
[0014] Further, based on the initial image sequence, determining the optimal focusing distance of the camera, comprising:
[0015] Calculating the object distance of the target in each image in the initial image sequence according to the thin - lens imaging principle, and determining the imaging range of the camera; using the local gradient method to calculate the clarity index of each image, constructing and solving a system of equations according to the relationship between the clarity index and the object distance, taking the position of the extreme point of the object distance as the optimal imaging distance of the target, obtaining the corresponding image clarity index, and thereby determining the optimal focusing distance of the camera.
[0016] Further, for the images in the new initial image sequence, determining the information quantity fusion weighted graph of the images, comprising:
[0017] By calculating the scene irradiance value and defining the signal-to-noise ratio within the local region of the image, the gradient magnitude is obtained and the image information content in the spatial domain is determined; by performing a discrete Fourier transform on the image and combining it with a high-pass filter to calculate the high-frequency component map, and then the image information content in the frequency domain is determined; the image information content in the spatial domain and the image information content in the frequency domain are weighted and combined to obtain the weighted image information content fusion map.
[0018] Further, the process of calculating the scene irradiance value and defining the signal-to-noise ratio within the local region of the image, obtaining the gradient magnitude and determining the image information content in the spatial domain includes:
[0019] For each image in the new initial image sequence, first calculate the corresponding grayscale image, and then perform an inverse gamma transformation on the grayscale image to convert the image into a linear image;
[0020] Normalize the linear image to obtain the normalized image; then establish an irradiance model based on the corresponding imaging parameters and camera response function of each image in the new initial image sequence. The irradiance model uses gamma transformation to obtain the scene irradiance value of the image;
[0021] For the images in the new initial image sequence, calculate the mean and standard deviation in the local image region centered on each pixel point, define the signal-to-noise ratio in the local image region, and then use the gradient kernel to calculate the gradient magnitude of the scene irradiance value; multiply the signal-to-noise ratio by the gradient magnitude of the scene irradiance value to obtain the image information content in the spatial domain.
[0022] Further, the process of performing a discrete Fourier transform on the image and combining it with a high-pass filter to calculate the high-frequency component map, and then determining the image information content in the frequency domain includes:
[0023] First, perform a discrete Fourier transform on each image in the new initial image sequence;
[0024] Set the cut-off frequency and define the high-pass filter, use the high-pass filter to filter the result of the discrete Fourier transform, and perform an inverse Fourier transform on the filtered result to obtain the high-frequency component map;
[0025] Take the magnitude of the high-frequency component map as the image information content in the frequency domain.
[0026] Further, the process of determining the scene irradiance distribution and obtaining the scene irradiance map, and thus determining the optimal information content fusion weighted map includes:
[0027] Obtain the scene irradiance distribution by averaging all the scene irradiance values; scale and normalize the scene irradiance distribution to obtain the scene irradiance map, calculate the image information content in the spatial domain and the image information content in the frequency domain corresponding to the scene irradiance map, and perform weighting on the two to obtain the optimal information content fusion weighted map.
[0028] Further, after calculating the positive difference of the missing information, the following steps are further included:
[0029] Calculate the relative missing ratio for the positive difference of the missing information and the weighted fusion graph of the optimal information amount, and determine whether to continue image acquisition according to the relative missing ratio.
[0030] Further, the step of calculating the relative missing ratio for the positive difference of the missing information and the weighted fusion graph of the optimal information amount, and determining whether to continue image acquisition according to the relative missing ratio includes:
[0031] Perform full-image averaging processing on the positive difference of the missing information and the weighted fusion graph of the optimal information amount respectively to obtain the corresponding average missing information amounts and; define the ratio of the two as the relative missing ratio. If the relative missing ratio is less than the preset ratio threshold, the image acquisition is completed; otherwise, the imaging system continues to acquire images until the relative missing ratio is not less than the preset ratio threshold.
[0032] Further, the first layer of the fully connected neural network is a convolutional layer based on a convolutional neural network, the activation function is ReLU, and the feature map output by the convolutional layer is flattened into a vector;
[0033] The second and third layers of the fully connected neural network are both hidden layers using fully connected layers;
[0034] The last layer of the fully connected neural network is a fully connected layer, the activation function is Softplus, and the imaging parameters of the camera are output.
[0035] Further, when the fully connected neural network is trained, the loss function used is defined as the square of the difference between the weighted fusion graph of the information amount calculated from the image taken by the camera using the imaging parameters output by the fully connected neural network and the positive difference of the missing information.
[0036] A terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor; when the processor executes the computer program, the image dynamic acquisition method for scene spatio-temporal-frequency complementarity is implemented.
[0037] A computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the image dynamic acquisition method for scene spatio-temporal-frequency complementarity is implemented.
[0038] Compared with the prior art, the present invention has the following technical features:
[0039] The method of the present invention can start the calculation by using a small number of images taken with initial parameters, actively obtain the missing information in the scene, and maximize the acquisition of all the information in the scene, thereby providing a new solution for the dynamic acquisition of images. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a schematic flowchart of the method of the present invention;
[0041] Figure 2 is a schematic diagram of the application effect in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0042] The present invention provides an image dynamic acquisition method for spatio-temporal-frequency complementarity of a scene, including the following steps:
[0043] Step 1, preset the initial values of the imaging parameters of the cameras in the multi-component imaging system, and the cameras respectively use each set of initial imaging parameter values to collect images of the target to obtain an initial image sequence.
[0044] To reasonably estimate the best imaging distance of the target, three sets of initial imaging parameter values are set; the imaging parameters include exposure time, gain, and aperture size; where:
[0045] The first set of initial imaging parameter values is: set the focus distance to 10 meters, the exposure time to 1 / 500 second, the gain (ISO) to 100, and the aperture size f value to f / 11, and the image shows a short focus distance and local underexposure.
[0046] The second set of initial imaging parameter values is: set the focus distance to 25 meters, the exposure time to 1 / 60 second, the gain (ISO) to 200, and the aperture size f value to f / 4, and the image shows a medium focus distance and normal exposure.
[0047] The third set of initial imaging parameter values is: set the focus distance to 50 meters, the exposure time to 1 / 10 second, the gain (ISO) to 400, and the aperture size f value to f / 2.8, and the image shows a long focus distance and local overexposure.
[0048] By setting the above imaging parameters, an initial image sequence composed of locally underexposed images , normally exposed images and locally overexposed images
[0049] can be obtained, so as to maximize the acquisition of the basic information in the scene where the target is located.
[0050] Step 2, based on the initial image sequence, determine the best focus distance of the camera; after adjusting the camera with the best focus distance, re-obtain a new initial image sequence with the initial imaging parameter values.
[0050] Among them, the object distance of the target in each image in the initial image sequence is calculated according to the thin lens imaging principle, and the imaging range of the camera is determined; the clarity index of each image is calculated using the local gradient method, and the equation group is constructed and solved according to the relationship between the clarity index and the object distance to determine the extreme point position of the object distance as the optimal imaging distance of the target, and obtain the corresponding image clarity index, thereby determining the optimal focusing distance of the camera, as follows:
[0051] Extracting partially underexposed images , Normal exposure image , Partially overexposed images Focus distance to obtain imaging plane distance , according to the formula:
[0052] ;
[0053] The above formula is based on the principle of thin lens imaging. If it is to be applied to the calculation of imaging distance of a general camera, the variables in the above formula need to be redefined; among them, is the i-th image in the initial image sequence The corresponding focus distance, ; For images The object distance from the target to the object principal plane, For images The image distance from the image plane principal plane to the imaging plane; the object plane principal plane and the image plane principal plane are used to define the main plane where the target is located and the plane where the imaging position is located, that is, the light emitted by the point on the object plane principal plane will form a corresponding image point on the image plane principal plane after passing through the optical system, so the object distance can be solved for:
[0054] ;
[0055] The above formula is used to calculate each image in the initial image sequence. , , At different image distances The corresponding object distance is , , ; Image distance This can be calculated by looking at the camera lens calibration data and using the focus distance.
[0056] Similarly, using the above formula, the imaging range of the camera can be obtained, which is expressed as ,in and Represents the minimum and maximum focusing distances, and are the corresponding minimum and maximum image distances respectively. Therefore, the camera imaging range can be expressed as .
[0057] Due to the change of the value after adjusting the focusing distance, different values are presented, along with different degrees of defocus blur. Therefore, a method for evaluating image sharpness needs to be designed to solve the imaging distance that can make the target image the clearest.
[0058] This solution designs a method for calculating the image sharpness index based on the local gradient method. First, the initial image sequence is converted into a grayscale image, and the formula is:
[0059] ;
[0060] In the above formula, , and represent the values at the pixel coordinates in the red, green, and blue channels of the image respectively. The gradients in the horizontal and vertical directions are calculated using the gradient kernels sobel x and sobel y respectively:
[0061] ;
[0062] The grayscale image is convolved two-dimensionally using the gradient kernel to obtain:
[0063] ;
[0064] ;
[0065] In the formula, * is the two-dimensional convolution operation. The calculation formula for the gradient amplitude of each pixel is as follows:
[0066] ;
[0067] Then, the mean statistical method is used to calculate the image sharpness index:
[0068] ;
[0069] In the formula, is the sharpness index of the image , ROI is the area to be calculated in the image, N is the number of all pixel points in ROI. Here, 75% of the image area centered on the image is selected as ROI; from this, the sharpness index of each image , , Clarity index , , .
[0070] Image clarity index and object distance The relationship can be expressed as:
[0071] ;
[0072] In the formula is the coefficient to be solved. Using each image in the initial image sequence , , The corresponding object distance and image clarity index form three groups of data , , Substitute into the above equation to construct a system of equations:
[0073] ;
[0074] The coefficient can be obtained through the above system of equations.
[0075] For , the position of its extreme point can be calculated by the following formula:
[0076] ;
[0077] Take the position of the extreme point calculated above as the estimated target optimal imaging distance. At this time, the value of the corresponding image clarity index is the largest; according to this value , the corresponding optimal focusing distance can be obtained using the principle of similar triangles, that is:
[0078] ;
[0079] At this time, the imaging system controls the camera to adjust the focusing distance to and then the target can be clearly focused.
[0080] After focusing is completed, the camera re - shoots images according to the initial imaging parameter values at the optimal focusing distance to obtain a new initial image sequence , and then proceed to the next step.
[0081] If the calculated target optimal imaging distance is not within the camera imaging range , then loop back to step 1 to reacquire the initial image sequence and recalculate according to step 2.
[0082] Step 3: for each image in the new initial image sequence, determine the information fusion weighted graph of the image.
[0083] In this step, for the images in the new initial image sequence, the gradient amplitude is obtained and the image information in the spatial domain is determined by calculating the scene irradiance value and defining the signal-to-noise ratio in the local area of the image; the high-frequency component map is calculated by discrete Fourier transforming the image and combining it with a high-pass filter, and then the image information in the frequency domain is determined; the image information in the spatial domain and the image information in the frequency domain are weighted and combined to obtain a weighted image information fusion map.
[0084] (3.1) Image information in the spatial domain.
[0085] The new initial image sequence is recorded as ; For each image, first calculate the corresponding grayscale image , and then perform an inverse gamma transform on the grayscale image to convert the image into a linear image :
[0086] ;
[0087] in, is the gamma transform parameter.
[0088] For linear images Normalization is performed so that the pixel interval in the image is normalized to [0,1], and we get ; Then according to the new initial image sequence The corresponding imaging parameters for each image, including exposure time , Gain ,aperture And the camera response function Build an irradiance model; here it is simplified to a simple gamma transform ,in is the amount of incident light, which can be expressed as irradiance and exposure time , Gain , aperture size relationship, , so the scene irradiance value The estimation formula is:
[0089] ;
[0090] in is the inverse function of the camera response function.
[0091] Calculate the scene irradiance value After that, use the method combining local signal-to-noise ratio and local gradient to calculate the initial image sequence in the spatial domain image information content of each image in the, the specific method is as follows:
[0092] In the image , taking the pixel point as the center of the local image area (for example, a 5*5 window) to calculate the mean and standard deviation, and define the signal-to-noise ratio in the local image area , the specific calculation formula is:
[0093] ;
[0094] Among them, is the image mean value in the local image area, is the standard deviation, is a constant to prevent division by zero.
[0095] Then use sobel x and sobel y gradient kernels to calculate the gradient magnitude of the scene irradiance value , to obtain edge and texture information; since the algorithm details have been listed above, it is simplified to the following formula here and ensure that its value is positive:
[0096] ;
[0097] In the formula represents the gradient calculation process.
[0098] After obtaining the signal-to-noise ratio in the local image area and the scene irradiance value , the spatial domain image information content is expressed as:
[0099] ;
[0100] The spatial domain image information content contains the information weight of each pixel in the local image area.
[0101] If you want to obtain the image information content evaluation index of the entire image , you can perform an average operation on the pixels of the image based on the information content of the local image area:
[0102] ;
[0103] In the formula represents the image The number of all pixels in
[0104] Image information quantity evaluation index It can reflect the detailed information and reliability contained in the images taken under different camera parameters. The higher this index, the more detailed information and higher signal-to-noise ratio the image has after camera parameter adjustment. It represents the richness of information of each pixel in the whole image.
[0105] (3.2) Image information quantity in the frequency domain.
[0106] First, for the image Calculate the discrete Fourier transform , and the calculation formula is:
[0107] ;
[0108] In the formula, and respectively represent the length and width values of the image , is the pixel coordinates in and are the horizontal and vertical direction frequency components respectively, is the imaginary unit, and e is the natural constant; the corresponding inverse transform formula is:
[0109] ;
[0110] Set a high-pass filter Define it as follows:
[0111] ;
[0112] In the formula is the cut-off frequency, , where represents the pixel value distance unit in the frequency domain; calculate the filtered frequency domain expression as:
[0113] ;
[0114] Calculate After the inverse Fourier transform of, the high-frequency component map is obtained, and the formula is as follows:
[0115] ;
[0116] In the formula, IFFT is the inverse Fourier transform.
[0117] Then define the image information quantity in the frequency domain For this high-frequency component diagram The amplitude value is as follows:
[0118] .
[0119] (3.3) Information fusion weighted diagram.
[0120] Next, according to the image information amount in the spatial domain and the image information amount in the frequency domain calculate the information fusion weighted diagram, and set the weights of the spatial domain and the frequency domain and both to be 0.5. The calculation formula of the information fusion weighted diagram is as follows:
[0121] .
[0122] Step 4: The imaging system continuously performs image acquisition on the target, calculates the scene irradiance value corresponding to each image, determines the scene irradiance distribution and obtains the scene irradiance diagram, and thus determines the optimal information fusion weighted diagram.
[0123] Specifically, after obtaining the scene irradiance value of each image, the scene irradiance distribution is obtained by averaging all the scene irradiance values; after scaling and normalizing the scene irradiance distribution, the scene irradiance diagram is obtained, the image information amount in the spatial domain and the image information amount in the frequency domain corresponding to the scene irradiance diagram are calculated, and the two are weighted to obtain the optimal information fusion weighted diagram.
[0124] The imaging system continuously performs image acquisition. First, calculate the scene irradiance value corresponding to each image . Since the camera is always in continuous shooting mode and shooting the same scene, it is considered here that the scene irradiance value is in a basically unchanged state; take the average value of the scene irradiance value of each image to obtain a relatively stable scene irradiance distribution .
[0125] In an ideal situation, the optimal camera parameter setting should make the image dynamic range contain as much as possible the entire dynamic range of the scene. Usually, the average brightness value of the normalized image pixels is close to 0.5, so that the maximum detail contrast and higher signal-to-noise ratio in the area where the pixel is located can be obtained. Therefore, based on the scene irradiance distribution define the scaling coefficient :
[0126] ;
[0127] where The function is to take the average value of all pixels of the irradiance diagram.
[0128] Using the zoom factor Distribute the scene irradiance Scaling and normalization to the range 0 to 1:
[0129] ;
[0130] in, Represents a scene irradiance map that maximizes detail and information; The function is a normalization operation.
[0131] The scene irradiance map is calculated using the local signal-to-noise ratio and local gradient combination method described above. The corresponding image information in the spatial domain ; Similarly, use the method in (3.2) to calculate the scene irradiance map The corresponding image information in the frequency domain , and weight the two based on the method in (3.3), and the result is recorded as the optimal information fusion weighted graph .
[0132] Step 5: based on the images currently captured by the imaging system, calculate the information fusion weighted graph of each image and fuse them; compare the fusion result with the optimal information fusion weighted graph, and calculate the positive difference of the missing information.
[0133] The imaging system has currently captured images, each image The generated information fusion weighted graph is denoted as In order to make full use of the information captured in different areas of each image, the pixel-by-pixel maximum fusion method is used to obtain the fusion result. , the formula is as follows:
[0134] ;
[0135] In the formula The function is to take the maximum operation among all values; in the above formula, each pixel If at least one historical image captures a high amount of information, the region gets a high value in the fusion map, whereas if none of the images adequately captures the details of the region, the fusion value is low.
[0136] The fusion results Fusion weighted graph with optimal information content Do a pixel-by-pixel comparison and calculate the positive difference of the missing information , the formula is as follows:
[0137] ;
[0138] The positive difference of the missing information reflects the amount of missing information in the overall historical image.
[0139] Calculate the relative missing ratio for the positive difference of the missing information and the best information volume fusion weighted map, and determine whether to continue image acquisition based on the relative missing ratio:
[0140] The positive difference of the missing information and the best information volume fusion weighted map are respectively subjected to full-image averaging processing to obtain the corresponding average missing information volume and ; Define the ratio of the two as the relative missing ratio:
[0141] ;
[0142] In the above formula, if , it can be determined that the image acquisition is completed and there is no need to supplement the missing information, and the next step can be carried out; otherwise, new images need to be continuously acquired until .
[0143] Step 6, construct a fully connected neural network. The input layer of the fully connected neural network is the average missing information volume obtained by subjecting the positive difference of the missing information to full-image averaging processing, and the output is the imaging parameters of the camera; train the fully connected neural network and save the trained fully connected neural network for predicting the imaging parameters of the camera, so that the camera can be adjusted according to the predicted imaging parameters and images can be obtained.
[0144] After obtaining the positive difference of the missing information , it can be understood as the information volume map of the new image to be taken. Therefore, as long as the image obtained by correctly adjusting the camera parameters can provide the same information volume map, the information volume required for the scene can be complemented.
[0145] Based on the above idea, design a simple fully connected neural network. The input of the fully connected neural network is the average missing information volume obtained by subjecting the positive difference of the missing information to full-image averaging processing , and the output is the imaging parameters of the camera, that is, the exposure time , gain , aperture size .
[0146] In the training stage, use the imaging parameters output by the fully connected neural network to control the imaging system to obtain new images, and then calculate the corresponding information volume fusion weighted map for the new images, and use it and the positive difference of the missing information Find the loss function to constrain the training of the fully connected neural network;
[0147] In the actual application stage, the currently obtained average missing information amount is input into the trained fully connected neural network , and the corresponding imaging parameters can be output; the imaging system uses these imaging parameters to control the camera and capture a new image, and the required missing information amount value can be included in the obtained new image, thereby realizing the acquisition of images with complementary spatio-temporal domains of the scene.
[0148] The structure of the fully connected neural network is designed as follows:
[0149] The input layer of the fully connected neural network is the average missing information amount , with a size of , where and are the length and width of the average missing information amount , and 1 is the number of channels;
[0150] The first layer of the fully connected neural network is a convolutional layer based on a convolutional neural network, and the activation function is ReLU, with the following mathematical expression:
[0151] ;
[0152] The feature map output by the convolutional layer is flattened into a vector , where the flattening operation is:
[0153] ;
[0154] The second layer is a hidden layer, which is a fully connected layer with 16 neurons, and the activation function is ReLU, with the following mathematical expression:
[0155] ;
[0156] where and are the weights and biases of the training parameters in the hidden layer of the second layer respectively, and
[0157] The third layer is a hidden layer, which is a fully connected layer with 32 neurons, and the activation function is ReLU, with the following mathematical expression:
[0158] ;
[0159] where and are the weights and biases of the training parameters in the hidden layer of the third layer respectively, The feature map output by the hidden layer of the third layer.
[0160] The output layer of the fully connected neural network is a fully connected layer with 3 neurons, and the activation function is Softplus to ensure that the output is positive:
[0161] ;
[0162] Among them, represents the output of the fully connected neural network, which is the imaging parameter and includes the exposure time , gain , aperture size ; and are the weights and biases of the output layer respectively.
[0163] When the fully connected neural network is training, its loss function is as follows:
[0164] ;
[0165] Among them, represents the information volume fusion weighted map calculated by the camera using the imaging parameter to shoot the image, represents the positive difference of the missing information. The above formula can constrain the fully connected neural network to select the correct imaging parameters to shoot the image with the required information volume value.
[0166] The method of the present invention has been applied to the space environment imaging task. Due to the special conditions such as complex lighting, extreme temperature difference, microgravity and high radiation in the space environment, a single or multiple image acquisitions with fixed parameters cannot contain all the information of the photographed target. The present invention realizes the method of obtaining all the effective information of the target by actively controlling the camera parameters to shoot the target for an indefinite number of times, as Figure 2 shown; in this example, the imaging system can start the calculation by using the images taken with the initial values of three groups of imaging parameters. By calculating the information volume map missing in all the images, the camera parameters are actively controlled for targeted shooting, and after 20 cycles of shooting, it stops, realizing the acquisition of the maximum information volume of the shooting target.
[0167] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for dynamic image acquisition for scene time-space-frequency complementarity, characterized in that: include: Preset the initial values of imaging parameters of cameras in multiple imaging systems, and use each group of initial values of imaging parameters to collect images of the target to obtain an initial image sequence; Based on the initial image sequence, determining the optimal focus distance of the camera; after adjusting the camera at the optimal focus distance, reacquiring a new initial image sequence with the initial values of the imaging parameters; For an image in a new initial image sequence, determining an information fusion weighted graph of the image; The imaging system continuously collects images of the target, calculates the scene irradiance value corresponding to each image, determines the scene irradiance distribution and obtains the scene irradiance map, thereby determining the optimal information fusion weighted map; Based on the images currently captured by the imaging system, the information fusion weighted graph of each image is calculated and fused; Compare the fusion result with the optimal information fusion weighted graph, and calculate the positive difference of the missing information, including: The imaging system has currently captured images, each image The generated information fusion weighted graph is denoted as ; The fusion result is obtained by using the pixel-by-pixel maximum fusion method , the formula is as follows: ; In the formula The function is to take the maximum operation among all values; The fusion results Fusion weighted graph with optimal information content Do a pixel-by-pixel comparison and calculate the positive difference of the missing information , the formula is as follows: ; The relative missing ratio is calculated by fusion weighted graph of the positive difference of missing information and the optimal amount of information, and the need to continue image acquisition is determined based on the relative missing ratio: Positive difference for missing information , optimal information fusion weighted graph Perform full-image averaging processing to obtain the corresponding average amount of missing information and ; Define the ratio of the two as the relative missing proportion: ; In the above formula, if When , it can be determined that the image acquisition is complete and there is no need to supplement the missing information, and the next step can be processed; otherwise, it is necessary to continue to acquire new images until ; A fully connected neural network is constructed. The input layer of the fully connected neural network is the average amount of missing information obtained by averaging the positive difference of the missing information over the entire image, and the output is the imaging parameters of the camera. The fully connected neural network is trained, and the trained fully connected neural network is saved for predicting the camera imaging parameters.
2. The method for dynamic image acquisition for scene time-space-frequency complementarity according to claim 1, characterized in that: Based on the initial image sequence, determine the optimal focus distance of the camera, including: According to the thin lens imaging principle, the object distance of the target in each image in the initial image sequence is calculated, and the imaging range of the camera is determined; the clarity index of each image is calculated using the local gradient method, and a group of equations is constructed and solved according to the relationship between the clarity index and the object distance to determine the extreme point position of the object distance as the optimal imaging distance of the target, and obtain the corresponding image clarity index, thereby determining the optimal focusing distance of the camera.
3. The method for dynamic image acquisition for scene time-space-frequency complementarity according to claim 1, characterized in that: For an image in a new initial image sequence, determine the information fusion weighted graph of the image, including: By calculating the scene irradiance value and defining the signal-to-noise ratio in the local area of the image, the gradient amplitude is obtained and the image information in the spatial domain is determined; by performing discrete Fourier transform on the image and combining it with a high-pass filter to calculate the high-frequency component map, the image information in the frequency domain is determined; the image information in the spatial domain and the image information in the frequency domain are weighted and combined to obtain the image information fusion weighted map.
4. The method for dynamic image acquisition for scene time-space-frequency complementarity according to claim 3, characterized in that: By calculating the scene irradiance value and defining the signal-to-noise ratio in the local area of the image, the gradient amplitude is obtained and the image information content in the spatial domain is determined, including: For each image in the new initial image sequence, first calculate the corresponding grayscale image, then perform an inverse gamma transform on the grayscale image to convert the image into a linear image; The linear image is normalized to obtain a normalized image; then an irradiance model is established according to the corresponding imaging parameters and camera response function of each image in the new initial image sequence, and the irradiance model adopts gamma transformation to obtain the scene irradiance value of the image; For the images in the new initial image sequence, the mean and standard deviation are calculated in the local image area centered on each pixel, and the signal-to-noise ratio in the local image area is defined. Then, the gradient amplitude of the scene irradiance value is calculated using the gradient kernel; the signal-to-noise ratio is multiplied by the gradient amplitude of the scene irradiance value to obtain the image information in the spatial domain.
5. The method for dynamic image acquisition for scene time-space-frequency complementarity according to claim 3, characterized in that: By performing a discrete Fourier transform on the image and combining it with a high-pass filter to calculate the high-frequency component map, the image information content in the frequency domain is determined, including: First, a discrete Fourier transform is performed on each image in the new initial image sequence; Set the cutoff frequency and define a high-pass filter, use the high-pass filter to filter the discrete Fourier transform result, and perform inverse Fourier transform on the filtered result to obtain a high-frequency component map; The amplitude of the high-frequency component map is used as the image information in the frequency domain.
6. The method for dynamic image acquisition for scene time-space-frequency complementarity according to claim 1, characterized in that: Determine the scene irradiance distribution and obtain the scene irradiance map, thereby determining the optimal information fusion weighted map, including: The scene irradiance distribution is obtained by averaging all scene irradiance values; the scene irradiance distribution is scaled and normalized to obtain a scene irradiance map, the image information in the spatial domain and the image information in the frequency domain corresponding to the scene irradiance map are calculated, and the two are weighted to obtain the optimal information fusion weighted map.
7. The method for dynamic image acquisition for scene time-space-frequency complementarity according to claim 1, characterized in that: The first layer of the fully connected neural network is a convolutional layer based on a convolutional neural network, the activation function is ReLU, and the feature map output by the convolutional layer is flattened into a vector; The second and third layers of the fully connected neural network are both hidden layers that use fully connected layers; The last layer of the fully connected neural network is the fully connected layer, the activation function is Softplus, and the imaging parameters of the camera are output.
8. A computer-readable storage medium storing a computer program; characterized in that: When the computer program is executed by a processor, the method for dynamic image acquisition for scene time-space-frequency complementarity according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Image recognition method and device, training method and device, electronic equipment and storage medium
CN116758618A
Nerve radiation field-based single tree image three-dimensional reconstruction method and device
CN118154770A