Low-light environment imaging enhancement method based on multi-frame synthesis

By employing optical flow estimation and wavelet transform decomposition and fusion algorithms, the problems of image noise suppression and detail preservation in low-light animal fast-moving scenes were solved, achieving high-quality nighttime animal observation image enhancement.

CN120876310APending Publication Date: 2025-10-31GUANGZHOU GOMO SHIJI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510819342.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing single-frame enhancement techniques are insufficient in noise suppression under low light conditions, and multi-frame synthesis methods have difficulty maintaining the consistency of fur texture details and the accuracy of body contours in fast-moving animal scenes, affecting the visual effect of the image and the accuracy of animal feature recognition.

Method used

The algorithm analyzes animal movement trends using optical flow estimation, segments the image into head, trunk, and limb regions, performs geometric correction using different registration parameters, and processes high and low frequency components using wavelet transform decomposition and fusion algorithms to detect texture direction consistency and repair contour breaks, generating an enhanced image.

Benefits of technology

It improves the clarity and detail of nighttime animal observation images, provides higher quality image data, and offers a reliable image recognition foundation for animal monitoring and research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876310A_ABST
    Figure CN120876310A_ABST
Patent Text Reader

Abstract

The invention provides a low-light environment imaging enhancement method based on multi-frame synthesis, and the method comprises the steps: carrying out the independent geometric correction of each region according to the optimal registration transformation matrix of each region, and obtaining the image data of each region after geometric correction; weight distribution is carried out on the high-frequency texture components according to the texture consistency evaluation result, if the texture direction deviation angle is smaller than a preset angle threshold value, the weight coefficient of the frame is improved, the high-frequency components of the multiple frames are fused, and a high-frequency synthesis component with enhanced texture details is obtained; reconstructing the texture detail enhanced high-frequency synthetic component and the contour retentivity optimized low-frequency component, and performing multi-scale fusion processing on the reconstructed image to obtain a final nighttime animal observation synthetic image; in an image quality evaluation module of the camera APP, animal key feature points are extracted from the synthesized image, the reliability of animal feature recognition is verified by calculating the stability of feature point descriptors, and a verification result of the reliability of animal feature recognition is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, and in particular to a method for enhancing low-light environment imaging based on multi-frame synthesis. Background Technology

[0002] With the increasing demand for animal, especially pet, photography, high-quality low-light imaging technology has become a core driving force for the development of camera apps. Current nighttime animal observation imaging methods used in camera apps mainly rely on single-frame enhancement techniques, which often suffer from insufficient noise suppression when processing images in low-light environments. While traditional multi-frame synthesis methods can improve image quality to some extent, they often produce unsatisfactory results when dealing with scenes of fast-moving animals, particularly in preserving animal details. To overcome these limitations, camera apps need more precise multi-frame processing algorithms to handle complex nighttime animal observation scenes. The rapid movement of animals during daily activities causes significant displacement changes between consecutive frames, directly affecting the accuracy of inter-frame registration. Insufficient registration accuracy further leads to inconsistencies in the texture details of animal fur during multi-frame synthesis, manifested as significant differences in the degree of fur edge blurring between different frames. More complexly, when registration algorithms attempt to compensate for these displacements, they often cause distortion of the animal's body contours during synthesis. This distortion not only affects the visual effect of the image but, more importantly, reduces the accuracy of subsequent animal feature recognition. How to maintain the inter-frame consistency of fur texture details and the accuracy of body contours while ensuring the reliability of animal feature recognition in the final synthesized image under conditions of rapid animal movement has become a key issue in the development of multi-frame synthetic imaging technology for nighttime animal observation. Summary of the Invention

[0003] This invention provides a low-light environment imaging enhancement method based on multi-frame synthesis, mainly including: calculating the pixel displacement of animals between consecutive frames using an optical flow estimation algorithm, analyzing the motion vector distribution of various parts of the animal's body in adjacent frames to obtain the overall motion trend of the animal; segmenting the animal image into three processing units: head region, trunk region, and limb region, using different registration parameters for the differences in motion characteristics of each region, calculating the pixel displacement consistency in each region, and determining the registration transformation matrix for each region; applying the registration transformation matrix to each region for geometric correction, processing the pixel displacement changes at the region boundaries, and generating corrected regional image data; decomposing the corrected regional image data into high-frequency texture components and low-frequency contour components using wavelet transform, detecting the consistency of animal hair texture direction in the high-frequency components, and analyzing the continuity of animal body contour in the low-frequency contour components; assigning weights according to texture direction consistency, fusing the high-frequency components from multiple frames to generate texture-enhanced high-frequency synthetic components; performing edge detection on the low-frequency contour components, repairing contour breaks or blurred areas, and generating low-frequency components with optimized contour preservation; reconstructing the high-frequency synthetic components and low-frequency components using inverse wavelet transform, processing the reconstructed image using a multi-scale fusion algorithm, and generating a synthetic image for nighttime animal observation.

[0004] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0005] This invention discloses a low-light environment imaging enhancement method based on multi-frame synthesis. For consecutive frame animal images, an optical flow estimation algorithm is used to calculate pixel displacement, analyze the motion characteristics of different animal parts, and segment the image into head, torso, and limb regions, performing geometric corrections on each. The image is decomposed using wavelet transform, and high-frequency texture components undergo orientation consistency detection and adaptive fusion, while low-frequency contour components undergo edge detection and repair. Finally, the image is reconstructed and fused to obtain the enhanced image. This invention also employs the SIFT algorithm to extract feature points, evaluate the quality of the synthesized image, and dynamically adjust the fusion parameters. This method effectively improves the clarity and detail of nighttime animal observation images, providing higher-quality image data for animal monitoring and research. Attached Figure Description

[0006] Figure 1 This is a flowchart of a low-light environment imaging enhancement method based on multi-frame synthesis according to the present invention.

[0007] Figure 2 This is a schematic diagram of a low-light environment imaging enhancement method based on multi-frame synthesis according to the present invention. Detailed Implementation

[0008] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0009] like Figure 1 -2, This embodiment of a low-light environment imaging enhancement method based on multi-frame synthesis may specifically include:

[0010] S101. Calculate the pixel displacement of the animal between consecutive frames. By analyzing the motion vector distribution of various parts of the animal's body in adjacent frames, obtain the overall motion trend and local deformation information of the animal. At the same time, based on the physiological structure characteristics of the animal, identify the differences in motion characteristics of the head region being relatively stable, the trunk region having moderate amplitude motion, and the limbs region having large amplitude swinging motion.

[0011] The Lucas-Kanade optical flow algorithm is used to perform pixel-by-pixel calculations on adjacent video frames to obtain the horizontal and vertical displacement of each pixel in the animal image. A motion vector field is generated based on the magnitude and direction of the pixel displacement. K-means clustering analysis is then performed on the motion vector field to classify pixels into three categories based on their motion amplitude: stable regions with motion amplitude less than a first preset threshold, moderate motion regions with motion amplitude between the first and second preset thresholds, and violent motion regions with motion amplitude greater than the second preset threshold. This yields a motion feature distribution map of different body parts of the animal. Based on the spatial distribution of stable regions in the motion feature distribution map, connected components of the stable regions are extracted. Hough circle detection is performed on each connected component. An accumulator is used to count the distance from edge points to the candidate circle center. If the accumulator peak value exceeds a preset detection threshold, the connected component is determined to contain a circular structure. The area ratio and compactness of the circular structure are calculated. If both the area ratio and compactness meet preset conditions, the region is marked as the head region, and the head center coordinates and contour boundary information are obtained. Starting from the head center coordinates, the positional relationship of the centroid of each region relative to the head center in the motion feature distribution map is calculated. The largest connected region adjacent to the head region and belonging to the moderate motion region is identified as the trunk region. The overall motion direction of the trunk is obtained by calculating the average value of the motion vectors of each pixel in the trunk region, and the degree of trunk deformation is obtained by calculating the standard deviation of the motion vectors, thus obtaining the trunk motion feature parameters. Based on the trunk region boundary and the motion feature distribution map, the connected regions connected to the trunk region in the violent motion region are extracted as the limb regions. The change sequence of motion vectors in the limb regions over time is calculated, and the frequency components of the motion vectors are analyzed by Fourier transform. If there is a significant dominant frequency component, it is determined to be a regular swaying walking state. If the rate of change of the motion vector direction exceeds a preset threshold and the amplitude continues to increase, it is judged to be a running or jumping state. Combining the motion features of the head, trunk, and limbs, the complete motion state information of the animal is obtained.

[0012] For example, the Lucas-Kanade optical flow algorithm constructs optical flow constraint equations to solve for the motion velocity of each pixel by establishing constant brightness and small displacement assumptions between adjacent frames.

[0013] Specifically, the algorithm establishes a local window for each pixel and uses the least squares method to solve for the motion vectors of all pixels within the window, thereby obtaining the displacement components of that point in the x and y directions. This pixel-by-pixel calculation method can capture subtle motion changes on the animal's body surface, providing an accurate data foundation for subsequent motion region segmentation. When processing the motion vector field, K-means clustering analysis first calculates the magnitude of the motion vector of each pixel, and then automatically divides the pixels into three categories based on the magnitude.

[0014] In one possible implementation, the algorithm initializes three cluster centers, corresponding to low, medium, and high movement amplitudes, respectively. The cluster center positions are iteratively updated until convergence yields the final classification result. This adaptive classification method can automatically adjust the threshold based on the movement characteristics of different animals, avoiding the limitations of manually setting fixed thresholds. Hough circle detection plays a crucial role in identifying the head region.

[0015] It should be noted that this algorithm uses a parameter space voting mechanism to map edge points in the image space to the parameter space. Each edge point votes for possible center positions, and the position with the most votes in the accumulator is the detected center.

[0016] For example, when a running deer is detected, its head region exhibits a smaller range of motion compared to the torso and limbs due to the stabilizing effect of the neck muscles. The Hough transform can accurately identify this approximately circular stable region. The motion feature distribution map, as the core data structure connecting various processing steps, not only records the motion category of each pixel but also preserves the original motion vector information.

[0017] Preferably, by calculating the connectivity between adjacent regions, discrete pixels can be organized into body parts with semantic meaning. The identification of the torso region relies on its spatial adjacency to the head and its moderate range of motion; this identification method based on both spatial and motion constraints improves the accuracy of part segmentation. Fourier transform, when analyzing the periodicity of limb movements, converts the time-domain motion vector sequence to the frequency domain for analysis.

[0018] For example, when an animal walks normally, the swinging of its limbs exhibits a regular periodicity, manifesting as a distinct dominant frequency peak in the frequency spectrum. However, when an animal suddenly accelerates and runs, the irregularity of the movement causes the spectral energy to disperse, and the dominant frequency characteristic disappears. This frequency domain analysis method can quantitatively describe the regularity of movement patterns, providing a reliable basis for accurately judging the animal's behavioral state. By integrating head stability, moderate-amplitude trunk movements, and large-amplitude periodic swinging of the limbs, a comprehensive description of the animal's complete movement state is formed. This multi-level motion analysis method significantly improves the accuracy and robustness of animal behavior recognition.

[0019] S102. In the real-time processing module of the camera APP, the animal image is divided into three independent processing units: head region, trunk region, and limb region, based on the overall movement trend of the animal and the differences in the obtained movement characteristics. Different registration parameters are used for the differences in movement characteristics of each region. By calculating the displacement consistency coefficient of the pixels in each region, the optimal registration transformation matrix of each region is determined.

[0020] Based on the overall movement trend and movement characteristic differences of the animal, a region growing algorithm based on movement amplitude is adopted in the real-time processing module of the camera APP. Starting from the pixel with the smallest movement amplitude in the movement trend data as the head seed point, the expansion starts outward. When the difference in movement amplitude between adjacent pixels exceeds a preset threshold, the expansion stops, thus obtaining the head region boundary. The center of the pixel cluster area with medium movement amplitude is used as the trunk seed point to expand and obtain the trunk region boundary. The edge of the pixel cluster area with the largest movement amplitude is used as the limb seed point to expand and obtain the limb region boundary, generating a segmentation result containing three independent processing units. For each processing unit in the segmentation result, the motion vectors of all pixels within the unit are extracted between adjacent frames. The variance of the motion vectors in the horizontal and vertical directions is calculated. The principal and secondary directions of motion are obtained through singular value decomposition. If the motion variance of the head unit is less than a first threshold, a two-dimensional rigid transformation containing only horizontal and vertical translation parameters is assigned to it. If the motion variance of the torso unit is between the first and second thresholds, a rigid transformation containing translation and rotation angle parameters is assigned to it. If the motion variance of the limb units is greater than the second threshold, an affine transformation containing translation, rotation, and scaling factors is assigned to it. Based on the transformation type assigned to each processing unit, the current position coordinates and target position coordinates of all pixels within the unit are substituted into the corresponding transformation equation. The transformation parameter value that minimizes the sum of the squares of the Euclidean distances between the transformed positions of all pixels and the target position is solved using the least squares method. The distance between the predicted position of each pixel after transformation and the actual target position is calculated. The displacement consistency coefficient of the unit is obtained by summing the reciprocals of the distances of all pixels and dividing by the total number of pixels. The transformation parameters of each unit are iteratively optimized based on the displacement consistency coefficient. The partial derivatives of the displacement consistency coefficient with respect to each transformation parameter are calculated. The parameter values ​​are adjusted in the positive direction of the partial derivatives to increase the coefficients. The optimization stops when the coefficient change between two adjacent iterations is less than the convergence threshold. The optimal registration transformation matrix of each head processing unit, torso processing unit, and limb processing unit is output.

[0021] For example, the application of region growing algorithms in animal image segmentation is based on a core principle: pixels with similar motion characteristics often belong to the same body part.

[0022] Specifically, the algorithm starts from a seed point, examines the pixels within its 8-neighborhood, calculates the difference in motion amplitude between the neighboring pixels and the seed point, and when the difference is less than a preset threshold, the neighboring pixel is included in the current region and used as a new growth point to continue expansion. This gradual expansion method can adaptively determine the region boundary, avoiding the oversegmentation or undersegmentation problems that may be caused by traditional fixed threshold segmentation.

[0023] In one possible implementation, the selection of seed points directly affects the segmentation results. Head seed points are determined by scanning the entire motion field to find the center of the connected region with the smallest motion amplitude, because the animal's head typically remains relatively stable during movement to maintain visual positioning. Trunk seed points are selected from the middle segment of the motion amplitude histogram, reflecting the moderate motion characteristics of the trunk as the main body. Limb seed points are selected from the high motion amplitude regions at the image edges, corresponding to the large swinging features of the limbs. Singular value decomposition plays a role in dimensionality reduction and feature extraction when processing motion vectors.

[0024] It should be noted that the motion vectors within each processing unit constitute a two-dimensional matrix, where each row represents a pixel, and the two columns represent the motion components in the horizontal and vertical directions, respectively. Through singular value decomposition, the first singular vector indicates the primary motion direction of the unit, and the corresponding singular value reflects the motion intensity along that direction. The second singular vector, perpendicular to the primary direction, represents the secondary motion component. The choice between rigid transformation and affine transformation reflects the accurate modeling of the motion characteristics of different body parts.

[0025] For example, the head primarily undergoes translational movements, so only two translation parameters are needed to accurately describe its motion. The torso, in addition to translation, also experiences some rotation, requiring an additional rotation angle parameter. Due to joint movement and perspective effects, the limbs may exhibit scale changes in addition to translation and rotation, thus requiring complete affine transformation parameters. Solving for these transformation parameters using the least squares method involves constructing an overdetermined system of equations.

[0026] For example, for a head unit containing 100 pixels, the current position and target position of each pixel constitute a constraint equation, forming an overdetermined system with 200 equations but only 2 unknown translation parameters. The optimal transformation parameter estimate is obtained by minimizing the sum of squared residuals of all constraint equations. The calculation of the displacement consistency coefficient reflects a quantitative assessment of the transformation quality. When the transformation parameters are accurate, all pixels within the unit can be well aligned to the target position after transformation, the distance between the predicted and actual positions is small, and the corresponding reciprocal is large, resulting in a high consistency coefficient. This design makes the coefficient value positively correlated with the registration quality, facilitating subsequent optimization. The iterative optimization process guides the parameter adjustment direction by calculating the partial derivatives of the objective function with respect to each parameter. Each time, the parameters are updated along the direction that increases the consistency coefficient the most, gradually approaching the optimal solution, ultimately generating an accurate registration transformation matrix for each body part.

[0027] S103. Based on the optimal registration transformation matrix of each region, perform independent geometric correction on each region to obtain geometrically corrected regional image data.

[0028] Based on the optimal registration transformation matrix for each region, the original coordinates of each pixel within the region are calculated using the transformation matrix to obtain its floating-point coordinate position in the target image. Bilinear interpolation is then used to determine the pixel value at this position. The pixel values ​​of the four nearest-neighbor integer coordinates around the floating-point coordinates are obtained, and horizontal and vertical interpolation weights are calculated based on the fractional part of the floating-point coordinates. The pixel values ​​of the four neighboring points are multiplied by their corresponding weights and summed to obtain the geometrically corrected pixel value. This process is repeated for all pixels to complete the initial correction image of the region. For each pixel in the initial correction image, its offset from its original position is calculated. The rate of change of offset is obtained by the difference between the offsets of adjacent pixels. If the rate of change of offset of a pixel exceeds a preset boundary threshold, the point is determined to be near the region boundary and exhibit discontinuity. The coordinates of all pixels meeting this condition are recorded in the boundary discontinuity point set. For each pixel in the set of boundary discontinuities, obtain its pixel value in the current region and adjacent regions. Establish a circular window centered on this point with a preset radius. Calculate the distance from each position within the window to the center and normalize it to the interval 0 to 1 as a mixing weight. Calculate the weighted average of the pixel values ​​in the current region and the adjacent regions to obtain the feathered transition pixel value. Replace the original pixel value at the corresponding position in the boundary discontinuity set with the feathered transition pixel value, and combine it with the unfeathered pixel values ​​within the region to output the geometrically corrected regional image data.

[0029] For example, the application of transformation matrices in geometric correction involves a mathematical transformation process that maps the original image coordinates to the target coordinates.

[0030] Specifically, for an affine transformation matrix that includes translation, rotation, and scaling, the coordinates of each original pixel are multiplied by this 3×3 matrix to obtain new coordinates in the target image. These new coordinates are usually in floating-point form; for example, the original integer coordinates (100, 150) might be transformed into floating-point coordinates such as (102.3, 148.7). Bilinear interpolation solves the problem of determining the pixel value at floating-point coordinates.

[0031] In one possible implementation, for the floating-point coordinates (102.3, 148.7), the algorithm first determines four surrounding integer coordinates: (102, 148), (103, 148), (102, 149), and (103, 149). Then, interpolation weights are calculated based on the fractional parts of the floating-point coordinates, 0.3 and 0.7. Horizontally, the weights for the two points to the left are 0.7, and for the two points to the right are 0.3; vertically, the weights for the two points above are 0.3, and for the two points below are 0.7. Through this bidirectional linear interpolation, the final pixel value equals the weighted average of the values ​​of its four neighboring pixels. The calculation of the offset change rate reflects the severity of local image deformation.

[0032] It should be noted that when different transformation matrices are used to correct different parts of an animal's body, discontinuities will occur at the boundaries between regions.

[0033] For example, the head region might only undergo a small translation, while the adjacent torso region might be rotated at a larger angle, resulting in a noticeable positional deviation between previously adjacent pixels after correction. This discontinuity can be quantified by calculating the difference in offset between adjacent pixels. The boundary threshold is set based on the requirement of image continuity. When the rate of change in offset exceeds a preset threshold, it indicates a significant visual break at that location, requiring special processing. This threshold is typically determined based on the human eye's sensitivity to image discontinuities, ensuring that obvious boundary problems are detected while avoiding misjudgments of normal texture changes. Feathering achieves a smooth transition between regions through gradient blending.

[0034] For example, for a pixel marked as having a discontinuous boundary, the algorithm constructs a circular window around it. Each position within the window receives a different blending weight based on its distance from the center. Positions closer to the center retain more pixel features from the current region; positions farther away utilize more pixel information from adjacent regions. This gradual transition avoids the visual abruptness caused by hard boundaries. By combining the feathered boundary pixels with the internal correction pixels, the final generated regional image data maintains the independent motion correction effect for each body part while achieving a natural visual integrity, effectively solving the boundary discontinuity problem caused by regional processing.

[0035] S104. Decompose the geometrically corrected regional animal image data into high-frequency texture components and low-frequency contour components. Perform texture direction consistency detection on the animal hair texture information in the high-frequency components, determine the texture direction deviation angle between adjacent frames, and obtain the texture consistency evaluation result.

[0036] Two-dimensional discrete wavelet transform is used to perform multi-scale decomposition on geometrically corrected regional animal image data. The image is decomposed into horizontal high-frequency sub-bands containing detailed information, vertical high-frequency sub-bands, diagonal high-frequency sub-bands, and low-frequency approximation sub-bands containing overall shape information. The three high-frequency sub-bands are merged to form a high-frequency texture component matrix, and the low-frequency approximation sub-band is used as a low-frequency contour component matrix. A multi-directional Gabor filter bank is applied to the high-frequency texture component matrix, with the filter directions increasing from the horizontal direction to the vertical direction at preset angle intervals. The filter response value of each pixel in each direction is calculated, and the direction with the largest response value is selected as the main texture direction of that point, generating a texture direction map containing the main texture directions of all pixels. The texture direction map is compared with the texture direction map of the adjacent frame acquired at the previous time step, and the angle difference of the texture direction values ​​of the corresponding pixels is calculated. If the absolute value of the angle difference is less than a preset consistency threshold, the pixel is marked as texture consistent. The proportion of texture consistent points and the distribution range of angle differences among all pixels are counted to form the texture consistency evaluation result.

[0037] For example, the two-dimensional discrete wavelet transform in image processing is similar to breaking down a complete painting into different levels of detail.

[0038] Specifically, wavelet transform performs convolution operations on the image using a set of high-pass and low-pass filters, decomposing the original image into four sub-bands. The horizontal high-frequency sub-band captures vertical edge information, the vertical high-frequency sub-band extracts horizontal edge features, the diagonal high-frequency sub-band contains diagonal texture details, and the low-frequency approximation sub-band preserves the overall contour and brightness distribution of the image. This decomposition method is particularly suitable for processing animal images because animal fur is mainly reflected in high-frequency details, while the body contour is mainly present in low-frequency components.

[0039] In one possible implementation, the merging process of the three high-frequency subbands employs a weighted fusion method. Considering that animal fur may exhibit texture features in different directions, the horizontal, vertical, and diagonal subbands reflect texture information in different directions, respectively. By calculating the energy distribution of each subband, higher-energy subbands are assigned greater weights to ensure that the merged high-frequency texture component matrix can comprehensively reflect the multi-directional features of the fur. The Gabor filter bank design is based on the receptive field model of the biological visual system.

[0040] It should be noted that each Gabor filter consists of the product of a sine wave and a Gaussian envelope, where the direction of the sine wave determines the filter's sensitivity to texture in a specific direction. The filtering response reaches its maximum when the filter direction aligns with the texture direction in the image. By setting multiple Gabor filters with different directions, evenly distributed from horizontal to vertical, hair texture in any direction can be detected. The calculation of the angular difference between texture directions involves processing the cyclic angle.

[0041] For example, if the texture direction of a pixel in the current frame is 170 degrees, and the corresponding pixel in an adjacent frame is 10 degrees, a direct subtraction would yield a difference of 160 degrees. However, the actual minimum angle difference between them is only 20 degrees. Therefore, it is necessary to calculate the angle differences in both the forward and reverse directions simultaneously, and take the smaller value as the actual deviation. This calculation method ensures the accuracy of texture consistency assessment.

[0042] S105. Based on the texture consistency evaluation results, the high-frequency texture components are weighted. If the texture direction deviation angle is less than the preset angle threshold, the weight coefficient of that frame is increased. The high-frequency components of multiple frames are fused to obtain the high-frequency synthetic components with enhanced texture details.

[0043] Based on the texture direction deviation angle of each pixel in the texture consistency evaluation result, the weight value of the pixel is calculated by an exponential decay function. When the deviation angle is zero, the weight value is 1. As the deviation angle increases, the weight value decreases exponentially. When the deviation angle reaches the preset angle threshold, the weight value drops to 0.1. Weight calculation is performed on all pixels in the current frame to obtain the weight distribution matrix of the high-frequency texture components in the current frame.

[0044]

[0045] This formula represents the method for calculating the weight of a pixel, where w i θ represents the weight value of the i-th pixel. i θ represents the texture direction of that pixel. ref The reference texture direction is represented by the square of the difference between the two, which represents the square of the direction deviation angle. σ is a parameter controlling the decay rate. When the texture direction of a pixel is consistent with the reference direction, the weight is at most 1; as the deviation angle increases, the weight value decreases rapidly through an exponential decay function. For a continuous multi-frame sequence including the current frame, each frame generates its own weight distribution matrix using the same method, and obtains the high-frequency texture component data of these frames. For each pixel position, the high-frequency value (representing texture detail) and the corresponding weight value (representing the reliability or importance of the frame at that position) are extracted from multiple frames. The product of the high-frequency value and the weight of each frame is calculated and accumulated to obtain the total "weighted high-frequency value". At the same time, the weight values ​​of each frame are accumulated to obtain the sum of the weights. For each frame, the sum of the total "weighted high-frequency values" is divided by the sum of the weights to obtain the fused high-frequency value of that pixel position. Multi-frame weighted fusion calculations are repeatedly performed on all pixel locations in the image. The fused high-frequency values ​​at each location are arranged according to their original spatial positions to form a high-frequency synthesis component matrix for enhanced texture details. This matrix retains detail information with good texture direction consistency and suppresses noise components with large direction deviations.

[0046] For example, the application of the exponential decay function in weight calculation is based on an important principle: small deviations in texture direction should receive close to full weight, while the weight should decrease rapidly as the deviation increases.

[0047] Specifically, the function's form resembles the decay characteristics of a Gaussian distribution. When the deviation angle is 0 degrees, the exponent of the exponential function is 0, and the weight value remains 1. As the deviation angle gradually increases, the exponent becomes negative and its absolute value increases, causing the weight value to decrease exponentially. This non-linear decay method is more consistent with visual perception characteristics than linear decay because the human eye is not sensitive to small changes in texture direction but is very sensitive to large deviations.

[0048] In one possible implementation, the weight value is reduced to 0.1 instead of 0 to account for the robustness requirements in practical applications. Even if the texture orientation deviation of a pixel reaches a threshold, 10% of the weight is still retained to avoid completely losing the information of that pixel. This approach of retaining the minimum weight can prevent the complete loss of information due to individual outliers without significantly affecting the overall fusion result. The acquisition of consecutive multi-frame sequences involves the caching mechanism of the camera app.

[0049] It's important to note that during real-time processing, the camera continuously captures video frames and stores them in a circular buffer. When multi-frame fusion is required, several frames, including the current frame, are extracted from the buffer; typically, 5 to 7 frames are chosen as the fusion sequence. This choice of frame number balances the need for temporal continuity and computational efficiency. The weighted fusion calculation process reflects the principle of selective information preservation.

[0050] For example, for a given pixel location, if the texture direction deviation in the first frame is 5 degrees, the weight is 0.8; in the second frame, the deviation is 2 degrees, and the weight is 0.95; and in the third frame, the deviation is 15 degrees, and the weight is 0.3. During fusion, the second frame contributes the most to the final result due to its high weight, while the contribution of the third frame is significantly weakened. By accumulating the weighted values ​​of each frame and dividing by the total weight, normalization is achieved, ensuring that the fused pixel values ​​remain within a reasonable range. The spatial characteristics of the weight distribution matrix reflect the local consistency of the animal hair texture.

[0051] For example, in the animal's back region, fur typically grows longitudinally along the body, and the texture direction deviation between adjacent frames is small, thus the weight value in this region is generally high. However, at joints or body transitions, deformation caused by movement can lead to significant changes in texture direction, resulting in a corresponding decrease in weight value. Texture detail enhancement is achieved through selective fusion. Texture details in high-weight regions are enhanced after multi-frame fusion because consistent information from multiple frames is superimposed; while random noise or motion blur in low-weight regions is effectively suppressed because noise patterns in different frames are typically inconsistent and cancel each other out during weighted averaging. This adaptive fusion method based on texture consistency preserves stable texture features while eliminating unstable interference components.

[0052] S106. Use morphological gradient operators to perform edge detection on low-frequency contour components. Identify potential deformation areas by calculating the rate of curvature change of contour edges in contour breakage or blurry areas. If the rate of curvature change exceeds a preset deformation threshold, perform contour correction on the area to obtain low-frequency components with optimized contour preservation.

[0053] A morphological gradient operator is used to process the low-frequency contour components. The image is first dilated and then eroded using a structuring element to obtain the outer contour boundary image. Similarly, the same image is first eroded and then dilated to obtain the inner contour boundary image. The pixel value difference between the outer and inner boundary images is calculated to obtain the contour gradient image. Pixels with gradient values ​​below a preset continuity threshold are marked as contour breakpoints, and regions where gradient values ​​change drastically over short distances are marked as contour blur regions. A marker map containing the location information of all breakpoints and blur regions is output. Based on the locations of breakpoints and blur regions in the marker map, pixel coordinates are extracted along the contour direction at fixed intervals to form a sampling point sequence. For each sampling point in the sequence and its adjacent points, the local curvature value of that point is obtained by calculating the reciprocal of the radius of the arc formed by the three points. The curvature change value is obtained by subtracting the curvature values ​​of adjacent sampling points, and the curvature change value is divided by the sampling interval to obtain the curvature change rate value for each point. The numerical sequence of curvature change rate is judged and processed. If the curvature change rate of a certain contour exceeds the preset deformation threshold, the contour points with normal curvature change rate before and after the start and end positions of the contour segment are extracted as interpolation control points. A cubic spline interpolation algorithm is used to generate a smooth curve based on the coordinates of the control points. The pixel values ​​of the original broken or blurred segments are replaced with the pixel values ​​of the interpolated curve. After the contour correction is completed, the low-frequency component with contour preservation optimization is obtained.

[0054] For example, the working principle of the morphological gradient operator is based on the basic operations in mathematical morphology.

[0055] Specifically, dilation uses a structuring element to slide across the image, replacing the pixel value at each location with the maximum value within the area covered by the structuring element, thus expanding bright areas. Erosion, on the other hand, replaces pixel values ​​with the minimum value within the covered area, causing bright areas to shrink. The outer boundary obtained by dilation followed by erosion is slightly larger than the original contour, while the inner boundary obtained by erosion followed by dilation is slightly smaller than the original contour; the difference between the two precisely highlights the contour edge information.

[0056] In one possible implementation, the choice of structuring element directly affects the edge detection performance. For animal contour detection, 3×3 or 5×5 circular structuring elements are typically chosen because they exhibit the same response characteristics in all directions, enabling uniform detection of contour edges across various orientations. When a contour breaks, the gradient value at that location decreases significantly because the inner and outer boundaries tend to coincide at the break point; conversely, in blurred regions, the gradient value fluctuates drastically over short distances, reflecting the contour's instability. The extraction of the sampling point sequence involves a contour tracking algorithm.

[0057] It should be noted that, starting from a certain starting point on the contour, a fixed pixel distance is advanced along the tangent direction of the contour, and this position is recorded as the next sampling point. This fixed interval is usually set to 5 to 10 pixels, which ensures sufficient sampling density to capture contour details without causing excessive computational burden due to overly dense sampling. This equidistant sampling allows the continuous contour curve to be discretized into a processable sequence of points. The calculation of local curvature is based on the principle of circular arc approximation.

[0058] For example, for three consecutive sampling points, a circle passing through these three points can be uniquely determined. The radius of this circle reflects the degree of curvature of the profile at that point; the smaller the radius, the more severe the curvature. Therefore, curvature is defined as the reciprocal of the radius. When the profile is approximately straight, the radius of the fitted circle tends to infinity, and the curvature approaches zero; when the profile makes a sharp turn, the radius of the fitted circle is small, and the curvature value is large. The calculation of the rate of change of curvature reveals the local deformation characteristics of the profile.

[0059] For example, on a normal animal contour, the curvature of adjacent points typically transitions smoothly, with the rate of curvature change remaining at a low level. However, at locations of joint distortion or contour damage, curvature abruptly changes, leading to a sharp increase in the rate of curvature change. By setting a reasonable deformation threshold, these abnormal areas can be accurately identified. The application of cubic spline interpolation in contour restoration ensures the smoothness of the reconstructed contour. This method not only requires the interpolation curve to pass through all control points but also requires the curve's first and second derivatives to be continuous at the control points. This ensures that the reconstructed contour will not exhibit sharp turns or unnatural fluctuations. By selecting contour points with normal curvature at both ends of the broken region as control points, the interpolated curve can naturally connect the broken parts, restoring the integrity of the contour and ultimately obtaining a visually coherent optimized contour that conforms to the animal's body shape characteristics.

[0060] S107. Reconstruct the high-frequency synthetic component with enhanced texture details and the low-frequency component with optimized contour preservation, and perform multi-scale fusion processing on the reconstructed image to obtain the final synthetic image for nighttime animal observation.

[0061] The high-frequency synthetic components for enhanced texture details and the low-frequency components for contour preservation optimization are reconstructed using inverse wavelet transform. The high-frequency synthetic components are used as detail coefficients input to the high-frequency channel of the inverse transform, and the optimized low-frequency components are used as approximation coefficients input to the low-frequency channel of the inverse transform. Multi-level inverse wavelet reconstruction operations are performed to obtain a preliminary reconstructed image containing enhanced texture and optimized contours. A Laplacian pyramid is constructed on the preliminary reconstructed image. After smoothing the image with a Gaussian filter, a 2x downsampling is performed to generate the first low-resolution image. This process is repeated to generate a multi-layer image sequence with decreasing resolution to form a Gaussian pyramid. Each layer image is upsampled and subtracted from the original image of the previous layer to obtain a Laplacian layer image containing the detail information of that layer, thus forming the Laplacian pyramid structure. Based on the frequency characteristics of each layer of the Laplacian pyramid image, fusion weights are assigned. High-frequency layers are given a weight value greater than 0.8 to preserve texture details, mid-frequency layers are given a weight value between 0.5 and 0.8 to balance details and contours, and low-frequency layers are given a weight value between 0.3 and 0.5 to maintain basic shape. Each layer image is multiplied by its corresponding weight, and upsampling is performed layer by layer starting from the top of the pyramid and added to the weighted image of the next layer. The images are recursively reconstructed to the original resolution to obtain the final synthetic image of nighttime animal observation.

[0062] For example, inverse wavelet transform plays a crucial role in information integration during image reconstruction.

[0063] Specifically, this process is the exact opposite of the forward wavelet transform, recovering the complete image by recombining the separated high-frequency and low-frequency components. In practice, the high-frequency synthesized component contains hair texture information enhanced through multi-frame fusion; these detail coefficients are precisely placed in the corresponding frequency bands of the inverse transform. Simultaneously, the low-frequency component carries overall animal morphology information after contour restoration, serving as the basic framework for reconstruction. The inverse transform organically combines these two types of information in the frequency domain through progressive upsampling and filtering operations, ultimately generating a complete image in the spatial domain.

[0064] In one possible implementation, the construction process of the Laplacian pyramid embodies the idea of ​​multi-resolution analysis. The Gaussian filter acts as a low-pass filter, its core function being to remove high-frequency components before downsampling to prevent aliasing. After the original image is Gaussian smoothed and downsampled by a factor of 2, the resulting image size is halved while retaining the main low-frequency information. This process is executed recursively, downsampling on top of the previous layer each time, forming a gradually shrinking image sequence. The generation of Laplacian layer images reveals the distribution of details at different scales.

[0065] It's important to note that when a low-resolution image is upsampled back to its original size, a difference exists between the upsampled and original images due to the loss of high-frequency information during downsampling. This difference precisely represents the detail information at that scale. By calculating the difference between the original and upsampled images, the resulting Laplacian layer accurately captures texture and edge information within a specific frequency band. The weight allocation strategy is designed based on the frequency characteristics of the image content.

[0066] For example, in nighttime animal images, the high-frequency layer primarily contains the fine texture of fur, details crucial for animal species identification, and is therefore assigned a high weight of 0.8 or higher. The mid-frequency layer contains key animal features such as eye and ear outlines, offering both detail and structural integrity, with weights set between 0.5 and 0.8 for balance. The low-frequency layer represents the animal's overall outline and posture; while lacking detail, it provides basic morphological information, with weights controlled between 0.3 and 0.5 to ensure a moderate contribution. A recursive reconstruction process using multi-scale fusion achieves layer-by-layer integration of information. Starting from the top of the pyramid, this layer contains the lowest-resolution outline information, which is then upsampled to the size of the next layer. Upsampling typically employs bilinear interpolation to ensure a smooth image transition. The upsampled result is added to the weighted Laplacian image of the next layer, achieving the fusion of low-frequency information with the details of that layer. This process proceeds layer by layer downwards, adding finer details at each step until the original resolution is restored, ultimately generating a nighttime animal observation image that retains both clear outlines and rich textural detail.

[0067] S108. In the image quality assessment module of the camera APP, key animal feature points are extracted from the synthesized image. The reliability of animal feature recognition is verified by calculating the stability of the feature point descriptor, and the verification result of the reliability of animal feature recognition is obtained.

[0068] In the image quality assessment module of the camera app, a scale-space pyramid is constructed using the SIFT algorithm for synthesized images of animals observed at night. Local extrema are detected at different scale levels using the difference of Gaussian function as candidate feature points. After removing low-contrast and edge response points, stable key feature points are obtained. For each feature point, the gradient direction histogram of its neighborhood pixels is calculated to generate a rotation-invariant multidimensional feature descriptor vector. Multiple original frame images used in the synthesis image generation process are acquired, and the same feature points and descriptors are extracted from each original frame. The Euclidean distance between the synthesized image feature descriptor and the corresponding feature descriptors of each original frame is calculated. The reciprocal of the normalized distance value is taken as the stability score of the feature point. The average stability score of all successfully matched feature points is calculated to obtain the feature point stability index. The feature point stability index is compared with a preset reliability threshold. If the index is lower than the threshold, a weight adjustment map is generated based on the spatial distribution of feature points in the image and the stability score. The region where the low-stability feature points are located is mapped to the spatial position of the aforementioned texture fusion process. The texture detail fusion weight coefficient of the region is reduced, and the weight coefficient of the high-stability region is increased, forming the updated texture fusion weight allocation parameters.

[0069] For example, the scale-space pyramid construction process of the SIFT algorithm embodies the multi-scale perception mechanism of the biological visual system.

[0070] Specifically, by applying Gaussian blurring to the original image to varying degrees, the effect of the human eye observing objects at different distances is simulated. Each scale layer is convolved using a Gaussian kernel with a specific standard deviation, with the standard deviation increasing layer by layer, forming an image sequence from sharp to blurred. Differential operations between adjacent scale layers produce Gaussian difference images. This difference operation approximates the Laplacian operator and can effectively detect blob and corner features in the image.

[0071] In one possible implementation, local extrema detection needs to be performed in three-dimensional space. Each candidate point is compared not only with its eight neighboring pixels at the same scale layer, but also with the nine pixels at corresponding positions in the adjacent scale layers above and below, for a total of 26 comparison points. Only when the response value of a point is the maximum or minimum value among these 26 points is it identified as an extremum. This rigorous screening mechanism ensures the stability of feature points under scale changes. The feature descriptor construction process fully considers the rotation invariance requirement.

[0072] It should be noted that a neighborhood window of a certain size is selected around the feature point, and the gradient magnitude and direction of each pixel within the window are calculated. All gradient directions are normalized according to the principal direction of the feature point, ensuring the descriptor is unaffected by image rotation. The neighborhood is divided into multiple sub-regions, and a gradient direction histogram is calculated for each sub-region. Finally, the histograms of all sub-regions are concatenated to form a high-dimensional feature vector. The core of stability evaluation lies in verifying the reliability of feature matching.

[0073] For example, when an animal feature point can be stably detected and matched across multiple frames of images, it indicates that the feature has high repeatability. The calculation of Euclidean distance reflects the similarity between descriptors; the smaller the distance, the more similar the features. Normalization maps the distance values ​​of different feature points to a uniform scale, facilitating subsequent overall evaluation. Taking the reciprocal converts the distance into a score, implementing the evaluation logic of "smaller distance, higher score." The generation of the weighted mapping map establishes a correspondence between feature stability and image regions.

[0074] For example, highly stable feature points can usually be detected in prominent feature areas such as the eyes and nose of animals. These areas should maintain a high weight during texture fusion to highlight details. However, in areas susceptible to motion, such as the edges of fur, feature stability is lower, and appropriately reducing the fusion weight can reduce motion artifacts. This feature stability-based feedback adjustment mechanism achieves adaptive optimization of image quality. By feeding back the reliability information of local features into the global fusion parameters, the final synthesized animal image maintains overall naturalness while effectively enhancing the recognition features of key areas, thus improving the practical value of nighttime observation images.

[0075] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within this application. No reference numerals in the claims should be construed as limiting the scope of the claims.

Claims

1. A low-light environment imaging enhancement method based on multi-frame synthesis, characterized in that, The method includes: The pixel displacement of the animal between consecutive frames is calculated by optical flow estimation algorithm, and the motion vector distribution of each part of the animal's body in adjacent frames is analyzed to obtain the overall motion trend of the animal. Animal images are segmented into three processing units: head region, trunk region, and limbs region. Different registration parameters are used for the differences in motion characteristics of each region. The consistency of pixel displacement in each region is calculated, and the registration transformation matrix of each region is determined. Geometric correction is performed on each region by applying a registration transformation matrix, and the pixel displacement changes at the region boundaries are processed to generate corrected regional image data. After wavelet transform decomposition and correction, the regional image data is divided into high-frequency texture components and low-frequency contour components. The consistency of animal hair texture direction in the high-frequency component is detected, and the continuity of animal body shape contour in the low-frequency contour component is analyzed. Weights are assigned based on texture direction consistency, and high-frequency components from multiple frames are fused to generate texture-enhanced high-frequency synthetic components. Edge detection is performed on low-frequency contour components to repair broken or blurred areas of the contour and generate low-frequency components with optimized contour preservation. High-frequency and low-frequency components are reconstructed using inverse wavelet transform, and the reconstructed image is processed using a multi-scale fusion algorithm to generate a synthetic image for nighttime animal observation.

2. The low-light environment imaging enhancement method based on multi-frame synthesis according to claim 1, characterized in that, The process involves calculating the animal pixel displacement between consecutive frames using an optical flow estimation algorithm, analyzing the motion vector distribution of various parts of the animal's body in adjacent frames, and obtaining the overall motion trend of the animal, including: The displacement is calculated pixel by pixel for adjacent video frames to generate a motion vector field. Cluster analysis is used to divide the pixels into stable regions, moderate motion regions, and violent motion regions according to the motion amplitude. Connected components are extracted from the stable regions to detect circular structures, and the head region is marked to obtain the head center coordinates and contour boundaries. Based on the head center coordinates, connected components of adjacent moderate motion regions are identified as the trunk region. The mean and standard deviation of the motion vectors in the trunk region are calculated to obtain the trunk motion direction and deformation degree. Violent motion regions connected to the trunk region are extracted as limb regions. The time series of motion vectors in the limb regions is analyzed to obtain the overall movement trend of the animal.

3. The low-light environment imaging enhancement method based on multi-frame synthesis according to claim 1, characterized in that, The process involves segmenting the animal image into three processing units: head region, trunk region, and limb region. Different registration parameters are used for the differences in motion characteristics of each region. The consistency of pixel displacement within each region is calculated, and the registration transformation matrix for each region is determined, including: The algorithm selects the pixel with the smallest movement amplitude as the seed point for the head and expands it to generate the head region boundary. It expands from the center of the pixel cluster area with medium movement amplitude to generate the trunk region boundary. It expands from the edge of the pixel cluster area with violent movement amplitude to generate the limb region boundary. It extracts the pixel motion vectors in each processing unit, calculates the variance in the horizontal and vertical directions, and assigns different transformation types according to the variance. It calculates the transformation parameters using the least squares method, generates the predicted position of the pixel in each processing unit, calculates the displacement consistency coefficient, iteratively optimizes the transformation parameters, and outputs the registration transformation matrix of the head, trunk and limb processing units.

4. The low-light environment imaging enhancement method based on multi-frame synthesis according to claim 1, characterized in that, The process of applying a registration transformation matrix to each region for geometric correction, processing pixel displacement changes at region boundaries, and generating corrected region-specific image data includes: A registration transformation matrix is ​​applied to the pixels in each region to calculate the floating-point coordinates of the target image. Bilinear interpolation is used to determine the pixel value, and the pixel offset change rate is calculated. Discontinuous points at the boundaries are marked. A circular window is established for the discontinuous points at the boundaries. The distance from each position in the window to the center is calculated and normalized to the interval between 0 and 1 as a mixing weight. The pixel value of the current region is weighted and averaged with the pixel values ​​of the adjacent regions according to the weight to obtain the feathered transition pixel value. The pixel value of the discontinuous points at the boundaries is replaced, and the corrected regional image data is output.

5. The low-light environment imaging enhancement method based on multi-frame synthesis according to claim 1, characterized in that, The image data, after wavelet transform decomposition and correction, is divided into high-frequency texture components and low-frequency contour components. The consistency of animal hair texture direction in the high-frequency components is detected, and the continuity of animal body contour in the low-frequency contour components is analyzed, including: After wavelet transform decomposition and correction of the regional image data, a high-frequency texture component matrix and a low-frequency contour component matrix are generated. A multi-directional filter is applied to the high-frequency texture component matrix to generate a texture direction map. The texture direction difference between adjacent frames is calculated, and the proportion of consistent points and the distribution of angle differences are statistically analyzed. Contour lines are extracted from the low-frequency contour component matrix, the connection strength and local curvature are calculated, and contour break points and blurred points are marked.

6. The low-light environment imaging enhancement method based on multi-frame synthesis according to claim 1, characterized in that, The detection of the consistency of animal hair texture direction in the high-frequency components includes: A multi-directional filter bank is applied to the high-frequency texture component matrix to calculate the filter response value of each pixel. The direction with the maximum response is selected to generate a texture direction map. The texture direction maps of adjacent frames are compared to calculate the direction difference of the pixels. Texture consistent points are marked, and the proportion of consistent points and the distribution range of angle difference are statistically analyzed to generate texture consistency evaluation results.

7. The low-light environment imaging enhancement method based on multi-frame synthesis according to claim 1, characterized in that, The step of edge detection of low-frequency contour components, repairing broken or blurred areas of the contour, and generating low-frequency components with optimized contour preservation includes: Morphological gradient operators are applied to low-frequency contour components to generate contour gradient images, marking breakpoints and blurred regions; the rate of curvature change of contour edges in broken or blurred regions is calculated; for regions where the rate of curvature change exceeds a threshold, an interpolation algorithm is used to generate smooth curves, replacing pixel values ​​of broken or blurred segments, and generating low-frequency components with optimized contour preservation.

8. The low-light environment imaging enhancement method based on multi-frame synthesis according to claim 1, characterized in that, The process of reconstructing high-frequency and low-frequency components using inverse wavelet transform, and then processing the reconstructed image using a multi-scale fusion algorithm to generate a synthetic image for nighttime animal observation includes: The high-frequency and low-frequency components are reconstructed by inverse wavelet transform to generate a preliminary reconstructed image. A Laplacian pyramid is constructed, and weights are assigned to the high-frequency, mid-frequency, and low-frequency layers. The layers are upsampled and weighted summed to reconstruct the original resolution, generating a synthetic image of animal observation at night.

Citation Information

Cited By

  • Online calibration method and system for hydrological flow data

    CN121346947A