Three-dimensional DIC method based on sub-pixel cost body and uncertainty

By introducing subpixel cost bodies and uncertainties into the three-dimensional deformation measurement technology, combined with the iterative update method of recurrent neural networks, the problems of mismatch and calculation error in the existing technology are solved, and high-precision and efficient three-dimensional deformation measurement are achieved.

CN120147525APending Publication Date: 2025-06-13CHONGQING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510217867.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing three-dimensional deformation measurement technology has mismatch and calculation errors when dealing with complex scenes and high-precision measurements, and it is difficult to achieve subpixel-level accuracy measurements.

Method used

Using a three-dimensional DIC method based on subpixel cost volume and uncertainty, images are collected through binocular cameras, feature extraction is performed using a network with shared weights, combined with whole pixel and subpixel-level measurement modules, a circular neural network is used to iteratively update displacements, and the matching strategy is dynamically adjusted to improve measurement accuracy.

Benefits of technology

This method can adaptively adjust the matching strategy, reduce mismatch and calculation errors, improve calculation efficiency and measurement accuracy, and realize subpixel-level three-dimensional digital image-related measurements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147525A_ABST
    Figure CN120147525A_ABST
Patent Text Reader

Abstract

The invention requests to protect a three-dimensional DIC method based on a sub-pixel cost body and uncertainty. According to the method, a neural network based on a sub-pixel cost body and uncertainty is constructed, the sampling range in the warping process is adjusted through uncertainty estimation based on the variance of the cost body, preliminary stereoscopic vision image matching and time sequence image displacement measurement are output through iteration of a recurrent neural network, interpolation is performed on an input image, and the warping degree is improved. And a sub-pixel cost body is constructed by using the amplified image and is used for calculating high-precision displacement. A binocular camera obtains an image and then carries out distortion removal and stereo correction, a speckle region segmentation result is used as a global feature to improve measurement precision, stereo vision image matching and time sequence image displacement measurement are output through iteration of a recurrent neural network, and then a high-precision three-dimensional digital image correlation method is realized according to a binocular stereo vision theory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence and optical measurement, and particularly relates to a three-dimensional DIC method based on sub-pixel cost volume and uncertainty. Background Technique

[0002] With the rapid progress of science and technology, three-dimensional deformation measurement technology has been widely applied in many fields such as industrial manufacturing, structural health monitoring, and material property testing. The core task of three-dimensional deformation measurement technology is to perform high-precision spatial reconstruction and displacement field analysis of the object surface deformation, aiming to obtain the deformation information of the object during the force application or environmental change process. These information are crucial for material property evaluation, structural safety analysis, and product quality control, and can provide important data support for engineering design and safety assurance.

[0003] Currently, three-dimensional deformation measurement methods are mainly divided into two categories: contact type and non-contact type. Contact measurement methods (such as extensometers and strain gauges) have high measurement accuracy and can accurately obtain the deformation of the object under different conditions. However, these methods require physical contact with the measurement object, which will have a certain impact on the measurement results in some cases. Especially in the measurement of flexible materials or surface-sensitive materials, their applications are limited. In contrast, non-contact measurement methods achieve non-contact measurement of the target through technologies such as optical imaging or laser scanning, can avoid physical interference, and can achieve large-area full-field measurement, so they have become a current research hotspot.

[0004] In contrast, non-contact measurement methods use technologies such as optical imaging or laser scanning to measure the target object. These methods do not rely on physical contact and avoid any interference to the object. Therefore, non-contact methods can achieve large-area full-field measurement, especially suitable for the measurement of some complex or sensitive materials. As an important branch of non-contact methods, three-dimensional deformation measurement technology based on stereo digital images has received extensive attention and use in practical applications due to its advantages such as high precision, non-contact, and full-field measurement. This technology usually uses binocular cameras to collect the deformed images of the target object, calculates the disparity using stereo matching methods, and combines with the temporal matching algorithm to extract the two-dimensional in-plane displacement information in the image sequence, thereby realizing the three-dimensional deformation measurement of the object. It shows good adaptability and stability in complex environments and dynamic deformation scenarios.

[0005] Due to the significant differences between stereo vision image matching and temporal image displacement measurement in terms of image transformation types and geometric constraint conditions, there are also great differences in the network structure and loss function design of existing deep learning models for these two matching tasks. Temporal image displacement measurement usually focuses on texture or optical flow changes in image sequences, while stereo vision image matching focuses on the geometric consistency of parallax estimation. Therefore, constructing a unified deep learning model that can simultaneously adapt to stereo vision image matching and temporal image displacement measurement is an important challenge in the field of stereo DIC technology. The unified matching model for speckle images plays a crucial role in the matching accuracy and robustness of a three-dimensional deformation measurement system. Its design not only requires effective integration of the feature expressions of the two matching tasks but also needs to achieve compatibility and efficiency in the network structure, which has important research value and practical application potential.

[0006] After retrieval, the application publication number is CN118918245A, a three-dimensional digital image correlation method based on deep learning that is fully automatic and outputs pixel by pixel. Based on traditional three-dimensional reconstruction and displacement field calculation using binocular stereo vision, a deep learning displacement calculation framework containing two sub-models is proposed to replace the digital image correlation method based on image sub-regions, realizing stereo matching and temporal matching of images. For this deep learning framework, the large displacement estimation sub-model first calculates a rough displacement field and warps and deforms the image with this displacement field to eliminate the main displacement components. Subsequently, the high-precision small displacement estimation sub-model further extracts the smaller residual deformation between the reference image and the warped image. Finally, the large and small displacements are subtracted to obtain the final refined high-precision large-range displacement field. The present invention utilizes the advantages of the deep learning displacement calculation framework, such as no need for complex manual parameter selection and the ability to output a pixel-by-pixel deformation field, to achieve fully automatic and pixel-by-pixel three-dimensional digital image correlation measurement.

[0007] Object deformation is a continuous process, and the intermediate images during the deformation process are not fully utilized, and the measurement accuracy of three-dimensional digital image correlation is at the integer pixel level. The present invention realizes the full utilization of the acquired data by using the result of current sub-pixel displacement measurement as the initial value of integer pixel displacement measurement for the image at the next moment, and introduces a sub-pixel cost volume to achieve sub-pixel-level three-dimensional digital image correlation measurement accuracy. Summary of the Invention

[0008] The present invention aims to solve the above problems of the prior art. A three-dimensional DIC method based on sub-pixel cost volume and uncertainty is proposed. The technical solution of the present invention is as follows:

[0009] A three-dimensional DIC method based on sub-pixel cost volume and uncertainty, which includes the following steps:

[0010] Step 1: Collect photos of the speckle deformation process on the object surface through a binocular camera, and perform operations of removing distortion, stereo rectification, and segmenting the speckle area on the collected images;

[0011] Step 2: Extract features through a network with shared weights to obtain a multi-scale pyramid feature map;

[0012] Step 3: Input the multi-scale pyramid feature map into the whole-pixel measurement module for whole-pixel displacement measurement. The whole-pixel measurement module includes a whole-pixel cost volume, uncertainty, and a recurrent neural network, and the measurement result is used as the initial value for subsequent sub-pixel measurement;

[0013] Step 4: Input the whole-pixel displacement measurement result and the left and right images into the sub-pixel level measurement module to calculate the sub-pixel level displacement. The sub-pixel level measurement module includes a whole-pixel cost volume, uncertainty, and a recurrent neural network, and the measurement result is used as the initial value of the pixel-level displacement of the two frames of images at the next moment.

[0014] Further, Step 1 specifically includes:

[0015] Step 11: Calibrate the binocular camera, and the calibration result is used for removing distortion, stereo rectification of the left and right images, and three-dimensional reconstruction of the object surface;

[0016] Step 12: Spray random spray speckles on the material surface, and use the camera to continuously collect images to record the deformation process of the material under external force;

[0017] Step 13: Remove distortion from the collected images;

[0018] Step 14: Perform stereo rectification on the left and right images collected at the same moment;

[0019] Step 15: Segment the speckle area from the collected images.

[0020] Further, Step 2 extracts features through a network with shared weights to obtain a multi-scale pyramid feature map, specifically including:

[0021] Step 21: The feature extraction network with shared weights includes three convolutional layers and three skip connections. Use the three convolutional layers to gradually downsample the input image to extract multi-scale features;

[0022] Step 22: Use the three skip connections to fuse global and local features to generate a pyramid-style multi-scale feature map.

[0023] Further, Step 3 specifically includes:

[0024] Step 31: The initial parallax is 0. Based on the previous moment's parallax map and uncertainty, transform the right image features;

[0025]

[0026] Among them represents the characteristics of the transformed right figure. K is the sampling point area centered on pixel p, and F R is the characteristic of the right view, and d n-1 (p + k) represents the disparity corresponding to the position (p + k), and w k (p) is the weight of the Kth point, and o(p, k) is the uncertainty-expanded sampling range;

[0027] Step 32: Calculate the pixel-level cost volume between the left figure feature and the transformed right feature map;

[0028]

[0029] V n (p) represents the cost volume at position p. F L (p) represents the feature vector of the left view at pixel p. R represents the search range of the current pixel in a specific direction, and r represents the offset;

[0030] Step 33: Calculate the uncertainty according to the variance of the pixel-level cost volume;

[0031]

[0032] U n represents the uncertainty mapping of the current nth iteration. V n is the cost volume, is the bilinear sampler, is the average value of V n , σ(·) is the sigmod function, and o represents the calculated offset;

[0033] Step 34: Iteratively update the displacement according to the pixel-level cost volume and the recurrent neural network, and repeat the above operations.

[0034] Furthermore, the specific steps of step 4 include the following steps:

[0035] Step 41: Obtain a high-resolution image by performing cubic spline interpolation on the original input image;

[0036] Step 42: Transform the right figure feature based on the previous moment's disparity map and uncertainty;

[0037] Step 43: Calculate the sub-pixel cost volume between the left figure feature and the transformed right figure feature by calculating the correlation within the range of 11 * 11 around the corresponding pixel points;

[0038] Step 44: Calculate the uncertainty according to the variance of the cost volume;

[0039] In step 45, taking the rough displacement measurement result as the initial value, combining with the sub-pixel cost volume, iteratively update the displacement through a recurrent neural network, repeat the above operations, and output high-precision stereo vision image matching and temporal image displacement measurement.

[0040] In step 46, the measurement result at this stage is used as the initial value for the rough displacement measurement of the subsequent two frames.

[0041] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the deep learning three-dimensional digital image correlation method based on sub-pixel cost volume and uncertainty as described in any one of the above.

[0042] A non-transitory computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the deep learning three-dimensional digital image correlation method based on sub-pixel cost volume and uncertainty as described in any one of the above.

[0043] The advantages and beneficial effects of the present invention are as follows:

[0044] A three-dimensional DIC method based on sub-pixel cost volume and uncertainty of the present invention, compared with the three-dimensional displacement measurement method of traditional stereo DIC, can adaptively adjust the matching strategy when dealing with complex scenes, reduce false matching and calculation errors, and improve the calculation efficiency. Compared with the existing stereo DIC methods based on deep learning, the method proposed by the present invention can simultaneously perform stereo vision image matching and temporal image displacement measurement. The deformation process of the object is continuous, and the intermediate images during the object deformation process are fully utilized. The present invention introduces uncertainty and sub-pixel cost volume, estimates the uncertainty using the variance of the cost volume, and thus dynamically adjusts the matching strategy in the region with higher uncertainty. By interpolation, a high-resolution image is obtained, a sub-pixel cost volume is constructed, and combined with a recurrent neural network, sub-pixel-level accuracy prediction can be achieved, significantly improving the measurement accuracy. Description of the Drawings

[0045] Figure 1 It is a flowchart of high-precision measurement of the material displacement field using the present invention in the preferred embodiment provided by the present invention;

[0046] Figure 2 It is a structural diagram of a three-dimensional DIC method based on sub-pixel cost volume and uncertainty shown in an exemplary embodiment. Detailed Embodiments

[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.

[0048] The technical solution of the present invention to solve the above technical problems is as follows:

[0049] The flowchart of the implementation of a three-dimensional DIC method based on sub-pixel cost volume and uncertainty of the present invention is as Figure 1 shown, and it is specifically divided into 4 steps:

[0050] Step 1. Use a binocular camera to collect photos of the speckle deformation process on the object surface, and perform operations such as removing distortion, stereo rectification, and segmenting the speckle area on the collected images.

[0051] Step 2. Perform feature extraction through a network with shared weights to obtain a multi-scale pyramid feature map.

[0052] Step 3. Input the feature map into the whole-pixel measurement module to measure the whole-pixel displacement, and use the measurement result as the initial value for subsequent sub-pixel measurement.

[0053] Step 4. Input the whole-pixel displacement measurement result and the left and right images into the sub-pixel level measurement module to calculate the sub-pixel level displacement. Use the measurement result as the initial value of the pixel level displacement of the two frames of images at the next moment.

[0054] The structure diagram for realizing three-dimensional digital image correlation through this model is as Figure 2 shown.

[0055] As a possible implementation manner of this embodiment, the specific process of the step 1 of using a binocular camera to take images of the object deformation process and performing operations such as removing distortion, stereo rectification, and segmenting the speckle area on the obtained images is as follows:

[0056] Step 11. Calibrate the binocular camera, and the calibration result is used for removing distortion, stereo rectification of the left and right images, and three-dimensional reconstruction of the object surface.

[0057] Step 12. Spray random spray speckles on the material surface, and use the camera to continuously collect images to record the deformation process of the material under external force.

[0058] Step 13. Remove distortion from the collected images.

[0059] Step 14. Perform stereo rectification on the left and right images collected at the same moment.

[0060] Step 15. Segment the speckle area from the collected images.

[0061] As a possible implementation manner of this embodiment, the specific process of the step 2 of performing feature extraction operations on the images processed in step 1 to obtain a multi-scale pyramid feature map is as follows:

[0062] Step 21: Gradually downsample the input image using three convolutional layers to extract multi-scale features.

[0063] Step 22: Use three skip connections to fuse global and local features to generate a pyramidal multi-scale feature map.

[0064] As a possible implementation of this embodiment, the specific process of inputting the feature map into the integer-pixel measurement module for integer-pixel displacement measurement and using the measurement result as the initial value for subsequent sub-pixel measurement in step 3 is as follows: Step 31: The initial disparity is 0. Based on the previous-frame disparity map and uncertainty, transform the right-view features.

[0065]

[0066] where represents the transformed right-view features, K is the sampling point region centered on pixel p, F R is the feature of the right view, d n-1 (p + k) represents the disparity corresponding to the position (p + k), w k (p) is the weight of the Kth point, and o(p, k) is the uncertainty-expanded sampling range;

[0067] Step 32: Calculate the pixel-level cost volume between the left-view features and the transformed right-view feature map.

[0068]

[0069] V n (p) represents the cost volume at position p, F L (p) represents the feature vector of the left view at pixel p, R represents the search range of the current pixel in a specific direction, and r represents the offset;

[0070] Step 33: Calculate the uncertainty based on the variance of the pixel-level cost volume.

[0071]

[0072] U n represents the uncertainty map of the current nth iteration, V n is the cost volume, is the bilinear sampler, is the average value of V n , σ(·) is the sigmod function, and o represents the calculated offset;

[0073] Step 34: Iteratively update the displacement according to the pixel-level cost volume and the recurrent neural network, and repeat the above operations.

[0074] As a possible implementation of this embodiment, step 4 inputs the whole pixel displacement measurement result and the left and right images into the sub-pixel level measurement module to calculate the sub-pixel level displacement. Taking the measurement result as the initial value of the pixel level displacement of two frames of images at the next moment includes the following steps:

[0075] Step 41: Perform cubic spline interpolation on the original input image to obtain a high-resolution image.

[0076] Step 42: Transform the right image features based on the previous moment's disparity map and uncertainty.

[0077] Step 43: Calculate the sub-pixel cost volume of the left image features and the transformed right feature map by calculating the correlation within the range of 11*11 around the corresponding pixel points.

[0078] Step 44: Calculate the uncertainty according to the variance of the cost volume.

[0079] Step 45: Take the rough displacement measurement result as the initial value, combine it with the sub-pixel cost volume, and iteratively update the displacement through a recurrent neural network. Repeat the above operations to output high-precision stereo vision image matching and sequential image displacement measurement.

[0080] Step 46: The measurement result at this stage is used as the initial value for the subsequent two-frame rough displacement measurement.

[0081] Before applying the model to three-dimensional displacement measurement, a large number of labeled datasets are required to train the model. However, it is difficult to obtain a real dataset of left and right images with labels. Therefore, a synthetic dataset is used to train the model. A large number of simulated datasets are generated by using multiple stereo displacement dataset synthesis methods to obtain the left and right images I Rt and I Lt at time t, the left and right images I R(t+1) and I L(t+1) at time t + 1, the disparity Disp t of the left and right images at time t, the disparity Disp t+1 of the left and right images at time t + 1, and the displacement fields u and v of I Lt and I L(t+1) in the x and y directions. And the total loss calculation formula during the model training process is as follows:

[0082]

[0083] where γ is set to 0.8, represents the prediction result after being processed by the sampler S.

[0084] The system, device, module or unit illustrated in the above embodiments can be specifically implemented by a computer chip or entity, or by a product with certain functions.

[0085] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0086] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0087] The above embodiments should be understood to be only used to illustrate the present invention and not to limit the protection scope of the present invention. After reading the contents of the present invention, technicians can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. A three-dimensional DIC method based on sub-pixel cost volume and uncertainty, characterized in that: The following steps are involved: Step 1: Use a binocular camera to collect photos of the speckle deformation process on the surface of the object, and perform operations such as removing distortion, stereo correction, and segmenting the speckle area on the collected images; Step 2: Extract features through a weight-sharing network to obtain a multi-scale pyramid feature map; Step 3: Input the multi-scale pyramid feature map into the integer pixel measurement module for integer pixel displacement measurement. The integer pixel measurement module includes an integer pixel cost volume, uncertainty and a recurrent neural network. The measurement result is used as the initial value of the subsequent sub-pixel measurement. Step 4: Input the whole pixel displacement measurement results and the left and right images into the sub-pixel measurement module to calculate the sub-pixel displacement. The sub-pixel measurement module includes a sub-pixel cost volume, uncertainty and a recurrent neural network. The measurement results are used as the initial value of the pixel displacement of the two frames of images at the next moment.

2. A three-dimensional DIC method based on sub-pixel cost volume and uncertainty according to claim 1, characterized in that: The step 1 specifically includes: Step 11, calibrate the binocular camera, and the calibration result is used to remove distortion, stereo correction of left and right images, and three-dimensional reconstruction of the object surface; Step 12, spraying random spray speckles on the surface of the material, and using a camera to continuously capture images to record the process of deformation of the material under the action of external force; Step 13, removing distortion from the collected image; Step 14, performing stereo correction on the left and right images collected at the same time; Step 15: segment the acquired image into a speckle area.

3. The three-dimensional DIC method based on sub-pixel cost volume and uncertainty according to claim 2, characterized in that: The step 2 extracts features through a weight-sharing network to obtain a multi-scale pyramid feature map, which specifically includes: Step 21, the weight-sharing feature extraction network includes three convolutional layers and three skip connections, and the three convolutional layers are used to gradually downsample the input image to extract multi-scale features; In step 22, three skip connections are used to fuse global and local features and generate a pyramid-like multi-scale feature map.

4. The three-dimensional DIC method based on sub-pixel cost volume and uncertainty according to claim 3, characterized in that: The step 3 specifically includes: Step 31, the initial disparity is 0, based on the disparity map and uncertainty at the previous moment, the right image features are transformed; in represents the feature vector of the transformed right view at pixel p, K is the sampling point area centered on pixel p, and F R is the characteristic of the right view, d n-1 (p+k) represents the disparity corresponding to the position (p+k), w k (p) is the weight of the Kth point, and o(p,k) is the expanded sampling range of uncertainty; Step 32, calculating the pixel-level cost volume between the left image feature and the transformed right feature image; V n (p) represents the cost volume at position p, F L (p) represents the feature vector of the left view at pixel p, R represents the search range of the current pixel in a specific direction, and r represents the offset; Step 33, calculating the uncertainty according to the variance of the pixel-level cost volume; U n represents the uncertainty map of the current nth iteration, V n For the price body, is a bilinear sampler, Yes V n The average value of , σ(·) is the sigmoid function, and o represents the calculated offset; Step 34, iteratively update the displacement according to the pixel-level cost volume and the recurrent neural network, and repeat the above operation.

5. The three-dimensional DIC method based on sub-pixel cost volume and uncertainty according to claim 4, characterized in that: The step 4 specifically comprises the following steps: Step 41, performing cubic spline interpolation on the original input image to obtain a high-resolution image; Step 42, based on the disparity map and uncertainty at the previous moment, transform the features of the right image; Step 43, calculating the sub-pixel cost volume of the left image feature and the transformed right feature image by calculating the correlation of the 11*11 range around the corresponding pixel point; Step 44, calculating the uncertainty based on the variance of the cost volume; Step 45, taking the coarse displacement measurement result as the initial value, combining it with the sub-pixel cost volume, iteratively updating the displacement through a recurrent neural network, repeating the above operation, and outputting high-precision stereo vision image matching and time-series image displacement measurement; Step 46: The measurement result of this stage is used as the initial value of the subsequent two frames of coarse displacement measurement.

6. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the deep learning three-dimensional digital image correlation method based on sub-pixel cost volume and uncertainty as claimed in any one of claims 1 to 4 is implemented.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the deep learning three-dimensional digital image correlation method based on sub-pixel cost volume and uncertainty as described in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Full-automatic pixel-by-pixel output three-dimensional digital image correlation method based on deep learning

    CN118918245A