Three-dimensional imaging method, device, equipment and storage medium of dynamic object

By combining two frames of fringe patterns with a convolutional neural network and a sinusoidal fringe projection measurement system, the problems of long processing time and insufficient accuracy in 3D imaging of dynamic objects are solved, and efficient and high-precision 3D imaging is achieved.

CN114820921BActive Publication Date: 2025-12-16SHENZHEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210270991.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-18
Publication Date
2025-12-16
Estimated Expiration
2042-03-18

AI Technical Summary

Technical Problem

Existing 3D imaging methods are time-consuming in the measurement of dynamic objects, which cannot meet the requirements of high-speed changing scenes, and their accuracy is insufficient.

Method used

By employing two frames of fringe patterns combined with a convolutional neural network and a sinusoidal fringe projection measurement system, the real and imaginary pixel distribution maps are obtained through an image feature extraction model. Phase calculation and unfolding are then performed, and the three-dimensional coordinate values ​​are calculated by combining binocular matching to achieve high-precision three-dimensional imaging.

Benefits of technology

It enables high-precision 3D imaging of dynamic objects with only two frames of fringe patterns, improving processing speed and accuracy, and meeting the measurement needs of dynamic objects in high-speed changing scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114820921B_ABST
    Figure CN114820921B_ABST
Patent Text Reader

Abstract

The application provides a three-dimensional imaging method, device and equipment of a dynamic object and a storage medium. The method comprises the following steps: projecting stripes on the dynamic object according to two preset stripe frequencies and shooting a stripe image; inputting the stripe image into a preset image feature extraction model to perform feature extraction processing, and obtaining a pixel point real distribution map and a pixel point imaginary distribution map of the dynamic object; performing phase calculation on the pixel points of the dynamic object according to the pixel point real distribution map and the pixel point imaginary distribution map, obtaining a folded phase map of the dynamic object, and performing phase unfolding processing on the folded phase map to obtain an absolute phase map of the dynamic object; performing binocular matching and depth calculation on all pixel points of the dynamic object according to the absolute phase map, obtaining three-dimensional coordinate values corresponding to each pixel point, and constructing a three-dimensional image of the dynamic object according to the three-dimensional coordinate values corresponding to each pixel point. The method has a small number of stripe images required in the implementation process and high three-dimensional imaging precision of the dynamic object.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of optical three-dimensional measurement, and in particular to a three-dimensional imaging method and device for a dynamic object, equipment and a storage medium. BACKGROUND

[0002] Fringe projection profilometry (FPP) is widely used in the field of three-dimensional measurement of objects due to its active and non-contact measurement characteristics. The fringe analysis method for static objects has been deeply researched and widely applied. With the development of high-speed projection and camera technology, three-dimensional surface measurement of dynamic objects has also become a research hotspot. In the existing three-dimensional imaging methods, in order to ensure high-precision three-dimensional measurement of the object, a phase shift method combined with time phase unwrapping is usually used. However, this method requires projection of multiple frames of fringe patterns with different phase shift amounts, and phase unwrapping also requires encoding of multiple fringe frequencies, which is time-consuming in the processing process. Although the accuracy is high, it cannot meet the requirements of high-speed change scenes of dynamic objects. SUMMARY

[0003] Therefore, the embodiments of the present application provide a three-dimensional imaging method and device for a dynamic object, equipment and a storage medium, which can realize high-precision dynamic three-dimensional imaging of the object by projecting two frames of fringe patterns, so that the number of fringe patterns required in the implementation process of three-dimensional imaging is small and the precision of dynamic three-dimensional imaging of the object is high.

[0004] A first aspect of the embodiments of the present application provides a three-dimensional imaging method for a dynamic object, comprising:

[0005] projecting fringe patterns on the dynamic object according to two preset fringe frequencies and capturing fringe patterns;

[0006] inputting the fringe patterns into a preset image feature extraction model for feature extraction processing to obtain a pixel point real distribution map and a pixel point imaginary distribution map of the dynamic object respectively;

[0007] performing phase calculation on the pixel points of the dynamic object according to the pixel point real distribution map and the pixel point imaginary distribution map to obtain a folded phase map of the dynamic object, and performing phase unwrapping processing on the folded phase map to obtain an absolute phase map of the dynamic object;

[0008] performing binocular matching and depth calculation on all pixel points of the dynamic object according to the absolute phase map to obtain three-dimensional coordinate values corresponding to each pixel point, and constructing a three-dimensional image of the dynamic object according to the three-dimensional coordinate values corresponding to each pixel point.

[0009] With reference to the first aspect, in a first possible implementation manner of the first aspect, before the step of inputting the fringe pattern into a preset image feature extraction model to perform feature extraction processing and obtaining a pixel point real part distribution map and a pixel point imaginary part distribution map of the dynamic object respectively, the method further includes:

[0010] determining a denoising threshold of the fringe pattern according to a background intensity distribution in the fringe pattern, and performing denoising processing on the fringe pattern according to the denoising threshold.

[0011] With reference to the first aspect, in a second possible implementation manner of the first aspect, before the step of inputting the fringe pattern into a preset image feature extraction model to perform feature extraction processing and obtaining a pixel point real part distribution map and a pixel point imaginary part distribution map of the dynamic object respectively, the method further includes:

[0012] collecting object fringe pattern samples at a plurality of different frequencies using a pre-built sinusoidal fringe projection measurement system;

[0013] calculating real part data values and imaginary part data values of all pixel points representing the object in the object fringe pattern samples using a phase shift method;

[0014] training a preset convolutional neural network model to a convergent state by taking the object fringe pattern samples as model inputs and taking the real part data values and the imaginary part data values of all pixel points representing the object in the object fringe pattern samples as model outputs, to generate an image feature extraction model.

[0015] With reference to the second possible implementation manner of the first aspect, in a third possible implementation manner of the first aspect, in the step of training the preset convolutional neural network model to a convergent state by taking the object fringe pattern samples as model inputs and taking the real part data values and the imaginary part data values of all pixel points representing the object in the object fringe pattern samples as model outputs to generate an image feature extraction model, the method includes:

[0016] randomly allocating all collected object fringe pattern samples according to a preset allocation ratio to obtain a training sample set and a verification sample set;

[0017] training the preset convolutional neural network model by taking a first object fringe pattern sample allocated to the training sample set as a model input and taking real part data values and imaginary part data values of all pixel points representing the object in the first object fringe pattern sample as model outputs, to obtain a trained convolutional neural network model;

[0018] inputting the second object fringe pattern sample into the trained convolutional neural network model for feature extraction processing to obtain first real part data values and first imaginary part data values of all pixel points of the object output by the trained convolutional neural network model;

[0019] comparing the first real part data values and the first imaginary part data values of all pixel points of the object output by the trained convolutional neural network model with real part data values and imaginary part data values of all pixel points of the object calculated by the phase shift method from the second object fringe pattern sample to obtain a similarity value;

[0020] comparing the similarity value with a similarity value of a previous iteration training, if a growth amplitude of the similarity value is less than a preset threshold, determining that the trained convolutional neural network model has been trained to a convergence state, and generating the trained convolutional neural network model as an image feature extraction model.

[0021] In a fourth possible implementation manner of the first aspect, in the third possible implementation manner of the first aspect, after the step of comparing the similarity value with a preset similarity threshold, if the similarity value is greater than the preset threshold, determining that the trained convolutional neural network model has been trained to a convergence state, and generating the trained convolutional neural network model as an image feature extraction model, the method further includes:

[0022] constructing a first loss function of the image feature extraction model according to the real part data values and the imaginary part data values of all pixel points of the object output by the image feature extraction model and the real part data values and the imaginary part data values of all pixel points of the object calculated by the phase shift method from the object fringe pattern sample;

[0023] obtaining a modulation distribution of the object fringe pattern sample to construct a second loss function of the image feature extraction model;

[0024] integrating the first loss function and the second loss function by weighting to generate a total loss function of the image feature extraction model, and performing model optimization training on the image feature extraction model by using the total loss function.

[0025] In a fifth possible implementation manner of the first aspect, in the step of training the preset convolutional neural network model to a convergent state by inputting the object fringe pattern sample as a model input and inputting the real part data value and the pixel point imaginary part data value of all pixel points representing the object in the object fringe pattern sample as a model output, the preset convolutional neural network model comprises four U-Net convolutional layer structures, and each U-Net convolutional layer structure is configured with a residual dense network block, and the residual dense network block is used to implement a feature extraction processing operation.

[0026] In a sixth possible implementation manner of the first aspect, the residual dense network block comprises a plurality of convolutional layers, and each convolutional layer is provided with a corresponding dilation rate.

[0027] A second aspect of the embodiment of the present application provides a three-dimensional imaging device of a dynamic object, the three-dimensional imaging device of the dynamic object comprising:

[0028] an image shooting module configured to project a fringe on the dynamic object according to a preset fringe frequency and shoot a fringe pattern;

[0029] a feature extraction module configured to input the fringe pattern into a preset image feature extraction model to perform a feature extraction processing, and obtain a pixel point real part distribution map and a pixel point imaginary part distribution map of the dynamic object respectively;

[0030] a phase map acquisition module configured to perform a phase calculation on pixel points of the dynamic object according to the pixel point real part distribution map and the pixel point imaginary part distribution map, obtain a folded phase map of the dynamic object, and perform a phase unwrapping processing on the folded phase map to obtain an absolute phase map of the dynamic object;

[0031] a three-dimensional imaging module configured to perform a binocular matching and a depth calculation on all pixel points of the dynamic object according to the absolute phase map, obtain a three-dimensional coordinate value corresponding to each pixel point, and construct a three-dimensional image of the dynamic object according to the three-dimensional coordinate value corresponding to each pixel point.

[0032] A third aspect of the embodiment of the present application provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the electronic device, and the processor implements each step of the three-dimensional imaging method of the dynamic object provided in the first aspect when executing the computer program.

[0033] A fourth aspect of the embodiment of the present application provides a computer readable storage medium, which stores a computer program, and each step of the three-dimensional imaging method of the dynamic object provided in the first aspect is implemented when the computer program is executed by a processor.

[0034] The three-dimensional imaging method, device, electronic equipment and storage medium for a dynamic object provided by the embodiments of the present application have the following beneficial effects:

[0035] The present application obtains the fringe pattern of the dynamic object according to the preset two fringe frequencies, inputs the obtained fringe pattern into the image feature extraction model pre-trained by the convolutional neural network for feature extraction processing, and respectively obtains the pixel point real distribution map and the pixel point virtual distribution map of the dynamic object by using neural network learning. According to the pixel point real distribution map and the pixel point virtual distribution map, the phase calculation of the pixel points of the dynamic object is performed, the folded phase map of the dynamic object is obtained, and the phase unwrapping processing of the folded phase map is performed to obtain the absolute phase map of the dynamic object. According to the absolute phase map, the binocular matching and depth calculation processing of all pixel points of the dynamic object are performed, the three-dimensional coordinate values corresponding to each pixel point are obtained, and the three-dimensional image of the dynamic object is constructed by using the three-dimensional coordinate values corresponding to each pixel point. The method improves the analysis accuracy of the single-frame fringe spectrum by using neural network training, simultaneously performs geometric constraint and preset of the virtual plane position in combination with the device structure of the sinusoidal fringe projection measurement system, realizes double-frequency phase unwrapping, and thus realizes the three-dimensional imaging of the dynamic object with high precision by using only two fringe patterns. BRIEF DESCRIPTION OF DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0037] Figure 1 The implementation flowchart of the three-dimensional imaging method for a dynamic object provided by the embodiments of the present application is provided.

[0038] Figure 2 The first implementation flowchart of training the image feature extraction model in the three-dimensional imaging method for a dynamic object provided by the embodiments of the present application is provided.

[0039] Figure 3 The second implementation flowchart of training the image feature extraction model in the three-dimensional imaging method for a dynamic object provided by the embodiments of the present application is provided.

[0040] Figure 4 The implementation flowchart of model optimization training of the image feature extraction model in the three-dimensional imaging method for a dynamic object provided by the embodiments of the present application is provided.

[0041] Figure 5A convolutional neural network structure diagram of an image feature extraction model in a three-dimensional imaging method of a dynamic object provided by an embodiment of the present application;

[0042] Figure 6 A basic structure block diagram of a three-dimensional imaging device of a dynamic object provided by an embodiment of the present application;

[0043] Figure 7 A basic structure block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0044] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0045] Please refer to Figure 1 , Figure 1 An implementation flowchart of a three-dimensional imaging method of a dynamic object provided by an embodiment of the present application. The details are as follows:

[0046] S11: Projecting stripes on the dynamic object according to two preset stripe frequencies and shooting stripe images.

[0047] In this embodiment, a pre-built sinusoidal stripe projection measurement system with a double camera and a single projector is used to project stripes on the dynamic object and shoot images, so as to obtain the stripe images of the dynamic object. Specifically, in the pre-built sinusoidal stripe projection measurement system, in combination with the system structure of the double camera and the single projector and the subsequent phase demodulation workflow required for three-dimensional imaging of the dynamic object, a plurality of selectable frequencies of sinusoidal projection stripes are pre-set. The sinusoidal stripe projection measurement system can generate stripe sequences with different pixel periods through different frequencies and download them to the projector. The stripe projection action of the projector and the shooting action of the camera on different scenes are triggered synchronously. In this embodiment, when three-dimensional imaging of the dynamic object is performed, the user can use the two frequencies pre-set by the sinusoidal stripe projection measurement system to project stripes on the dynamic object and shoot images, so as to obtain two frames of stripe images of the dynamic object.

[0048] S12: Inputting the stripe images into a pre-set image feature extraction model for feature extraction processing, and obtaining a pixel point real distribution image and a pixel point imaginary distribution image of the dynamic object, respectively.

[0049] In this embodiment, the preset image feature extraction model is a convolutional neural network model constructed using a U-Net network structure and residual dense network blocks. This convolutional neural network model is trained using stripe patterns of objects captured at a large number of different frequencies as training samples until it converges, in order to obtain the extraction of single-frame stripes. Figure 1 The ability to determine the real and imaginary parts of pixels in the spectrum of a given order means that the convolutional neural network model trained to convergence is used as the image feature extraction model to perform feature extraction processing. In this embodiment, after obtaining the stripe pattern of a dynamic object, the stripe pattern is input into the image feature extraction model for feature extraction processing, and single-frame stripes are extracted using the image feature extraction model. Figure 1 The real part of the pixel in the spectrum is extracted from the stripe pattern of the dynamic object to obtain the real part distribution map of the pixel in the dynamic object, and the single-frame stripes are extracted using this image feature extraction model. Figure 1 The channels of the imaginary part of the pixel in the first-order spectrum are extracted from the stripe pattern of the dynamic object to obtain the imaginary part distribution map of the pixel of the dynamic object.

[0050] S13: Based on the real part distribution map and the imaginary part distribution map of the pixels, perform phase calculation on the pixels of the dynamic object to obtain the folded phase map of the dynamic object, and perform phase unpacking processing on the folded phase map to obtain the absolute phase map of the dynamic object.

[0051] In this embodiment, the dynamic object is represented by a combination of several pixels in the stripe pattern. The real part distribution map of the dynamic object's pixels contains the real part data values ​​representing all pixels of the dynamic object, and the imaginary part distribution map of the dynamic object contains the imaginary part data values ​​representing all pixels of the dynamic object. For each pixel, the phase is calculated using the arctangent operation method based on the real and imaginary part data values ​​to obtain the folded phase of the pixel. Since the arctangent operation folds the phase values ​​of all pixels of the dynamic object between (-π, π), a folded phase map of the dynamic object can be obtained by plotting the folded phases of all pixels of the dynamic object onto a single image. Specifically, the formula for calculating the folded phase of a pixel using the arctangent operation method is as follows:

[0052]

[0053] Where θ represents the folded phase value of the pixel, a represents the real part data value of the pixel, and b represents the imaginary part data value of the pixel.

[0054] In the embodiment, after the folded phase map is obtained, the geometric constraint relationship and the virtual plane position of the view are determined through the double-camera and single-projector device structure in the sinusoidal fringe projection measurement system, and then the unfolded processing of the folded phase map is performed according to the geometric constraint relationship and the virtual plane position, so as to obtain the absolute phase map of the dynamic object. For example, in the embodiment, since the phase values of all pixel points of the dynamic object in the folded phase map are folded between (-π, π), the phenomenon of jumping of some pixel points in the folded phase map may occur, and in the embodiment, a measurement plane with a known depth is defined, and according to the geometric constraint relationship, the pixel points of the plane imaged on the camera can find a mapping region on the projector sensor plane, and the absolute phase value of the virtual plane is obtained. The phase map of the virtual object plane can be used for the absolute phase unfolding of the low-frequency fringe folded phase map, and the fringe order of the current pixel point in the folded phase map is determined through the relationship between the current pixel point and the pixel point at the corresponding position on the absolute phase map of the virtual plane, the folded phase value of the current pixel point is increased by the product of the fringe order and 2π, and in this way, all pixel points in the folded phase map are traversed, so that the unfolding processing of the low-frequency folded phase map is realized. Specifically, the calculation formula for determining the fringe order is,

[0055]

[0056] wherein, Φ min is the minimum phase map of the virtual plane, φ is the folded phase map, and ceil[·] is the upward rounding operation.

[0057] For the high-frequency fringe map, the fringe order of the current pixel point in the folded phase map is determined through the relationship between the current pixel point and the pixel point at the corresponding position on the low-frequency absolute phase map, the folded phase value of the current pixel point is increased by the product of the fringe order and 2π, and in this way, all pixel points in the folded phase map are traversed, so that the unfolding processing of the low-frequency folded phase map is realized. Specifically, the calculation formula for determining the fringe order is:

[0058]

[0059] wherein, t1, t2 represent the fringe periods of the two fringe sequences, φ w and φ u represent the high-frequency folded phase and the low-frequency unfolded phase respectively, and round[·] represents the nearest rounding operation. In this way, the unfolding processing of the folded phase map is realized, and a corresponding absolute phase map is obtained.

[0060] S14: According to the absolute phase map, all pixel points of the dynamic object are binocularly matched and depth calculated to obtain three-dimensional coordinate values corresponding to each pixel point, so as to construct a three-dimensional image of the dynamic object according to the three-dimensional coordinate values corresponding to each pixel point.

[0061] In this embodiment, the absolute phase map contains absolute phase data values of all pixel points of the dynamic object. According to the absolute phase map, the absolute phase value corresponding to each pixel point can be obtained. For example, for all pixel points in the absolute phase map, the disparity value and the depth value of each pixel point that is successfully matched are calculated through binocular matching, so as to determine the three-dimensional positional relationship between each pixel point. Further, a coordinate system of a three-dimensional space is constructed in combination with the calibration parameters of the sinusoidal fringe projection measurement system when the fringe pattern of the dynamic object is acquired. The calibration parameters include the calibration parameters of the binocular camera and the calibration parameters of the projector. Based on the coordinate system of the three-dimensional space, the three-dimensional coordinate value corresponding to each pixel point is obtained. The three-dimensional coordinate values corresponding to all pixel points are plotted in the three-dimensional space to form a three-dimensional point cloud. The three-dimensional image of the dynamic object can be obtained from the three-dimensional point cloud. For example, in this embodiment, the process of binocular matching is based on the pinhole imaging model. The projection center points of the two cameras in the binocular camera and the projection points of the object points on the imaging planes of the two cameras are obtained. After one of the cameras is calibrated, a ray equation is established according to the projection point of the object point on the imaging plane of the camera and the projection center point of the camera. Another ray equation is established according to the projection point of the object point on the imaging plane of the other camera and the projection center point of the other camera. The two ray equations correspond to two straight lines. The intersection point of the two straight lines is obtained by jointly solving the two ray equations. The three-dimensional coordinate value of the intersection point is set as the three-dimensional coordinate value of the object point.

[0062] As can be seen from the above, the three-dimensional imaging method of the dynamic object provided in this embodiment first acquires the fringe pattern of the dynamic object according to the two preset fringe frequencies, and inputs the acquired fringe pattern into the image feature extraction model pre-trained by the convolutional neural network for feature extraction processing. The pixel point real distribution map and the pixel point virtual distribution map of the dynamic object are respectively obtained by using neural network learning. Then, the phase of the pixel points of the dynamic object is calculated according to the pixel point real distribution map and the pixel point virtual distribution map, the folded phase map of the dynamic object is obtained, and the phase unwrapping processing is performed on the folded phase map to obtain the absolute phase map of the dynamic object. Finally, all pixel points of the dynamic object are binocularly matched and depth calculated according to the absolute phase map, the three-dimensional coordinate value corresponding to each pixel point is obtained, and the three-dimensional image of the dynamic object is constructed by using the three-dimensional coordinate value corresponding to each pixel point. This method improves the analysis accuracy of single-frame fringe spectrum by using neural network training, and realizes double-frequency phase unwrapping by combining the geometric constraint and the preset virtual plane position of the device structure of the sinusoidal fringe projection measurement system, so as to realize high-precision three-dimensional imaging of the dynamic object by using only two fringe patterns.

[0063] In some embodiments of the present application, the background intensity distribution can be represented as the mean value of the pixel color values presented by all background pixels in the fringe pattern. In this embodiment, the calculation formula of the background intensity distribution can be set as:

[0064]

[0065] wherein I n is the intensity distribution graph of the deformed fringe pattern with the same frequency and different phase shifts, and N is the phase shift step.

[0066] In this embodiment, the de-noising threshold for de-noising the fringe pattern can be determined according to the background intensity distribution in the fringe pattern, for example, the de-noising threshold can be set as the mean value of the pixel color values presented by all background pixels in the fringe pattern, for removing the background pixels in the fringe pattern. In this embodiment, after obtaining the fringe pattern of the dynamic object, the pixel color values of all pixels in the fringe pattern are obtained by scanning the fringe pattern, and based on the obvious jump difference between the pixel color values reflected by the target object and the pixel color values reflected by the background, the pixels representing the background and the pixels representing the target object can be distinguished from the fringe pattern by the difference between the pixel color values of each pixel. Further, the fringe pattern can be de-noised according to the de-noising threshold, and the pixels in the fringe pattern that meet the de-noising threshold condition can be removed, so that the pixels representing the background are removed as noise, so that only the pixels representing the dynamic object are retained in the fringe pattern. It can be understood that in this embodiment, the pixel color value reflected by the target object is generally within a color value interval and has a jump change difference with the pixel color value of the background, so that the pixel can be determined as a noise point in the fringe pattern. Therefore, the de-noising threshold can also be set based on the color value interval of the pixel color value reflected by the target object obtained by scanning, for removing the noise points in the pixel point area representing the dynamic object, for example, removing the overexposed pixel points in the fringe pattern.

[0067] In some embodiments of the present application, please refer to Figure 2 , Figure 2 The first implementation flowchart of training the image feature extraction model in the three-dimensional imaging method of the dynamic object provided by the embodiments of the present application is as follows:

[0068] S21: Collect the fringe pattern samples of the object at several different frequencies by using the pre-built sinusoidal fringe projection measurement system.

[0069] In this embodiment, in the pre-built sinusoidal fringe projection measurement system, several different frequencies of sinusoidal projection fringes are set, for example, assuming that the horizontal resolution of the projector device in the sinusoidal fringe projection measurement system is 1280 pixels, and the fringe frequencies required for the subsequent phase demodulation process are 16 and 64, that is, the fringe pixel periods are 80 and 20 respectively. In order to ensure the accuracy of the phase demodulation, the number of pixel periods of the coded fringe in the phase demodulation process should be an integer multiple of the number of phase shift steps, and the number of phase shift steps can be selected as 10 steps or 20 steps considering the common factor of the number of phase shift steps. At this time, when setting several different frequencies of sinusoidal projection fringes in the pre-built sinusoidal fringe projection measurement system, 10, 20, 30, 40, 50, 60, 70, 80 and the like can be set as the pixel periods of the fringes, and eight different frequencies can be obtained accordingly. In this embodiment, by using the pre-built sinusoidal fringe projection measurement system to project and shoot a plurality of different dynamic objects according to each frequency built-in, a large number of object fringe patterns under several different frequencies can be collected, and all the collected object fringe patterns are used as object fringe pattern samples.

[0070] S22: The real part data value and the imaginary part data value of all the pixel points representing the object in the object fringe pattern sample are calculated by using the phase shift method.

[0071] In this embodiment, after obtaining the object fringe pattern sample, the real part data value and the imaginary part data value of all the pixel points representing the object in the object fringe pattern sample are calculated by using the phase shift method for each object fringe pattern sample. In this embodiment, the number of phase shift steps has a suppression effect on harmonic error, and the N+2 step phase shift method is used to calculate the suppression of N order harmonic error, so as to calculate the real part data value and the imaginary part data value of all the pixel points representing the object in the object fringe pattern sample. Specifically, the collected object fringe pattern is calculated by using the ten-step phase shift method, and the calculation formula can be set as:

[0072]

[0073]

[0074] Wherein, Im represents the real part data value, Re represents the imaginary part data value, B represents the modulation distribution of the object fringe pattern, φ represents the phase of the object fringe pattern, and δ n represents the phase shift amount.

[0075] S23: The object fringe pattern sample is used as the model input, and the real part data value and the imaginary part data value of all the pixel points representing the object in the object fringe pattern sample are used as the model output to train the pre-set convolutional neural network model to the convergence state, so as to generate an image feature extraction model.

[0076] In this embodiment, the object fringe pattern samples are input as a model, and the real part data values and the pixel point imaginary part data values of all the pixels representing the object in the object fringe pattern samples calculated by the phase shift method are output as a model until the convolutional neural network model is trained to a convergent state. The convolutional neural network model trained to the convergent state is generated as an image feature extraction model.

[0077] In some embodiments of the present application, referring to Figure 3 , Figure 3 The second implementation flowchart of training an image feature extraction model in the three-dimensional imaging method of a dynamic object provided by the embodiments of the present application is provided. Details are as follows:

[0078] S31: All the collected object fringe pattern samples are randomly distributed according to a preset distribution ratio to obtain a training sample set and a verification sample set;

[0079] S32: For a first object fringe pattern sample distributed to the training sample set, the first object fringe pattern sample is input as a model, and the real part data values and the imaginary part data values of all the pixels representing the object in the first object fringe pattern sample are output as a model to train the preset convolutional neural network model to obtain a trained convolutional neural network model;

[0080] S33: For a second object fringe pattern sample distributed to the verification sample set, the second object fringe pattern sample is input into the trained convolutional neural network model for feature extraction processing to obtain first real part data values and first imaginary part data values of all the pixels representing the object output by the trained convolutional neural network model;

[0081] S34: The first real part data values and the first imaginary part data values of all the pixels representing the object output by the trained convolutional neural network model are compared with the real part data values and the imaginary part data values of all the pixels representing the object calculated by the phase shift method from the second object fringe pattern sample to obtain a similarity value;

[0082] S35: The similarity value is compared with the similarity value of the last iteration training. If the growth rate of the similarity is less than a preset threshold, it is judged that the trained convolutional neural network model has been trained to a convergent state, and the trained convolutional neural network model is generated as an image feature extraction model.

[0083] In this embodiment, all collected object fringe pattern samples can be randomly allocated to obtain the training sample set and the verification sample set of the convolutional neural network model according to an allocation ratio of 8:2. For example, the object fringe pattern samples allocated to the training sample set are first object fringe pattern samples, and the object fringe pattern samples allocated to the verification sample set are second object fringe pattern samples. At this time, the first object fringe pattern samples allocated to the training sample set can be used as model input, and the real part data values and the imaginary part data values of all pixel points representing the object in the first object fringe pattern samples calculated in advance by the phase shift method can be used as model output to train the convolutional neural network model, so as to obtain a trained convolutional neural network model. After obtaining the trained convolutional neural network model, the second object fringe pattern samples allocated to the verification sample set are used to verify the convergence of the model. Specifically, the second object fringe pattern samples in the verification sample set are input into the trained convolutional neural network model for feature extraction processing, and the first real part data values and the first imaginary part data values of all pixel points representing the object output by the trained convolutional neural network model are obtained. The first real part data values and the first imaginary part data values output by the trained convolutional neural network model are used as the prediction results of the model. Then, the real part data values and the imaginary part data values of all pixel points representing the object calculated in advance by the phase shift method in the second object fringe pattern samples are used as the expected results. The first real part data values and the first imaginary part data values of all pixel points representing the object output by the trained convolutional neural network model are compared with the real part data values and the imaginary part data values of all pixel points representing the object calculated by the phase shift method in the second object fringe pattern samples to obtain the similarity value between the prediction results and the expected results. Finally, the similarity value obtained by comparison is compared with the similarity value obtained in the last iteration training. If the similarity growth rate is greater than a preset threshold, it is indicated that the prediction results are close to the expected results, the trained convolutional neural network model is determined to have been trained to a convergence state, and the trained convolutional neural network model is generated as an image feature extraction model.

[0084] In some embodiments of the present application, please refer to Figure 4 , Figure 4 An implementation flowchart of the model optimization training of the image feature extraction model in the dynamic object three-dimensional imaging method provided by the embodiments of the present application is provided. Details are as follows:

[0085] S41: According to the real part data values and the imaginary part data values of all pixel points representing the object output by the image feature extraction model and the real part data values and the imaginary part data values of all pixel points representing the object calculated by the phase shift method in the object fringe pattern sample, a first loss function of the image feature extraction model is constructed.

[0086] In this embodiment, the calculation formula of the first loss function is set as:

[0087] L1=||Im output -Im gt ||1+||Re output -Re gt ||1

[0088] Wherein, Re output represents the real part data value of all pixel points representing the object output by the trained convolutional neural network model, Im output represents the imaginary part data value of all pixel points representing the object output by the trained convolutional neural network model, Re gt represents the real part data value of all pixel points representing the object calculated by the phase shift method of the object fringe pattern sample, Im gt represents the imaginary part data value of all pixel points representing the object calculated by the phase shift method of the object fringe pattern sample.

[0089] S42: Obtain the modulation distribution of the object in the object fringe pattern sample, and construct a second loss function of the image feature extraction model.

[0090] In this embodiment, the calculation formula of the second loss function is set as:

[0091]

[0092] Wherein, 5 represents 5 intermediate layers selected from the neural network of the image feature extraction model, λ i represents the weight parameter of the i-th intermediate layer, φ i represents the output of the i-th intermediate layer, represents the modulation distribution of the object calculated according to the real part data value and the imaginary part data value of all pixel points representing the object output by the image feature extraction model, and B represents the modulation distribution of the object calculated according to the real part data value and the imaginary part data value of all pixel points representing the object calculated by the phase shift method, wherein the calculation formula for calculating the modulation distribution of the object is

[0093] S43: The first loss function and the second loss function are weighted and integrated to generate a total loss function of the image feature extraction model, and the image feature extraction model is trained by using the total loss function.

[0094] In this embodiment, the total loss function is represented as:

[0095] Loss=L1+λL feat

[0096] Here, λ represents the weight parameter.

[0097] In this embodiment, after obtaining the total loss function, the image feature extraction model is optimized and trained by minimizing the loss calculated from the total loss function.

[0098] Please refer to the following: Figure 5 , Figure 5 This diagram illustrates the convolutional neural network structure of the image feature extraction model in the 3D imaging method for dynamic objects provided in this application embodiment. Figure 5 As shown, the preset convolutional neural network model has four U-Net convolutional layer structures. The feature processing of each U-Net convolutional layer structure is implemented using a residual dense network block (RDB). Different U-Net network structure layers are used to extract features at different scales in the stripe pattern.

[0099] In some embodiments of this application, the residual dense network block contains five consecutive convolutional layers. A corresponding dilation rate can be set for each convolutional layer in the residual dense network block. Dilation convolution is performed on each convolutional layer in the residual dense network block using this dilation rate, expanding the receptive field of each U-Net convolutional layer structure during feature processing, making the features extracted by each U-Net convolutional layer structure more accurate and effective. Since the convolutional kernels of dilated convolutions are discontinuous, the dilation rates of consecutively stacked dilated convolutions cannot have a common divisor greater than 1. To avoid grid effects, the dilation rates of each convolutional layer in the residual dense network block are set sequentially to 1, 2, 3, 2, 1.

[0100] It is understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0101] In some embodiments of this application, please refer to Figure 6 , Figure 6 This is a basic structural block diagram of a three-dimensional imaging device for dynamic objects provided in an embodiment of this application. In this embodiment, the device includes units used to perform the steps in the above-described method embodiments. Please refer to the relevant descriptions in the above-described method embodiments for details. For ease of explanation, only the parts relevant to this embodiment are shown. Figure 6As shown, the three-dimensional imaging device of a dynamic object comprises: an image shooting module 61, a feature extraction module 62, a phase map obtaining module 63, and a three-dimensional imaging module 64. The image shooting module 61 is configured to project and shoot a stripe pattern on a dynamic object according to a preset two stripe frequencies. The feature extraction module 62 is configured to input the stripe pattern into a preset image feature extraction model to perform feature extraction processing, and obtain a pixel point real part distribution map and a pixel point imaginary part distribution map of the dynamic object respectively. The phase map obtaining module 63 is configured to perform phase calculation on pixel points of the dynamic object according to the pixel point real part distribution map and the pixel point imaginary part distribution map, obtain a folded phase map of the dynamic object, and perform phase unfolding processing on the folded phase map to obtain an absolute phase map of the dynamic object. The three-dimensional imaging module 64 is configured to perform binocular matching and depth calculation on all pixel points of the dynamic object according to the absolute phase map, obtain three-dimensional coordinate values corresponding to each pixel point, and construct a three-dimensional image of the dynamic object according to the three-dimensional coordinate values corresponding to each pixel point.

[0102] It should be understood that the above three-dimensional imaging device of a dynamic object corresponds to the above three-dimensional imaging method of a dynamic object one by one, which will not be described here.

[0103] In some embodiments of the present application, please refer to Figure 7 , Figure 7 a basic structure block diagram of an electronic device provided by the embodiments of the present application. As shown in the figure, Figure 7 the electronic device 7 of the embodiments comprises a processor 71, a memory 72, and a computer program 73 stored in the memory 72 and executable on the processor 71, such as a program of a three-dimensional imaging method of a dynamic object. The processor 71 implements the steps in each embodiment of the above-mentioned three-dimensional imaging method of a dynamic object when executing the computer program 73. Alternatively, the processor 71 implements the functions of each module in the corresponding embodiment of the above-mentioned three-dimensional imaging device of a dynamic object when executing the computer program 73. For details, please refer to the related description in the embodiments, which will not be described here.

[0104] For example, the computer program 73 can be divided into one or more modules (units) stored in the memory 72 and executed by the processor 71 to complete the present application. The one or more modules can be a series of computer program instruction segments capable of completing a certain function, which are used to describe the execution process of the computer program 73 in the electronic device 7. For example, the computer program 73 can be divided into an image shooting module, a feature extraction module, a phase map obtaining module, and a three-dimensional imaging module, and the specific functions of each module are as described above.

[0105] The electronic device can include, but is not limited to, a processor 71, a memory 72. Those skilled in the art can understand that Figure 7 The electronic device 7 is only an example and does not constitute a limitation on the electronic device 7, and can include more or fewer components than illustrated, or combine certain components, or different components, for example, the electronic device can also include an input / output device, a network access device, a bus, etc.

[0106] The processor 71 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0107] The memory 72 can be an internal storage unit of the electronic device 7, such as a hard disk or a memory of the electronic device 7. The memory 72 can also be an external storage device of the electronic device 7, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 72 can include both the internal storage unit and the external storage device of the electronic device 7. The memory 72 is used to store the computer program and other programs and data required by the electronic device. The memory 72 can also be used to temporarily store data that has been output or will be output.

[0108] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units, since the same concept as the method embodiments of the present application, the specific functions and the technical effects brought about, specific can be seen from the method embodiments, this will not be repeated here.

[0109] The embodiments of the present application also provide a computer readable storage medium, the computer readable storage medium stores a computer program, the computer program is executed by the processor to realize the steps in each of the above method embodiments. In this embodiment, the computer readable storage medium can be non-volatile, and can also be volatile.

[0110] The embodiment of the present application provides a computer program product, when the computer program product runs on the mobile terminal, causes the mobile terminal to execute the steps in the above-mentioned various method embodiments.

[0111] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the system can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0112] The integrated module / unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program to instruct related hardware, and the computer program can be stored in a computer-readable storage medium. When the processor executes the computer program, the steps of the various method embodiments can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms, etc. The computer-readable medium can include any entity or device, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. that can carry the computer program code. It should be noted that the contents included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0113] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0114] The above-described embodiments are only used to illustrate the technical solutions of the present application, but not limit them; although the present application is described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method of three-dimensional imaging of a dynamic object, characterized in that, The method comprises the following steps: projecting and shooting a dynamic object with two preset fringe frequencies; inputting the fringe pattern into a preset image feature extraction model for feature extraction processing to obtain a pixel point real part distribution map and a pixel point imaginary part distribution map of the dynamic object respectively; calculating the phase of the pixel points of the dynamic object according to the pixel point real part distribution map and the pixel point imaginary part distribution map to obtain a folded phase map of the dynamic object, and performing phase unwrapping processing on the folded phase map to obtain an absolute phase map of the dynamic object; performing binocular matching and depth calculation on all pixel points of the dynamic object according to the absolute phase map to obtain three-dimensional coordinate values corresponding to each pixel point, and constructing a three-dimensional image of the dynamic object according to the three-dimensional coordinate values corresponding to each pixel point; Before the step of inputting the fringe pattern into a preset image feature extraction model for feature extraction processing to obtain a pixel point real part distribution map and a pixel point imaginary part distribution map of the dynamic object respectively, the method further comprises the following steps: collecting object fringe pattern samples at several different frequencies using a pre-built sinusoidal fringe projection measurement system; calculating the real part data values and the imaginary part data values of all pixel points of the object in the object fringe pattern samples using the phase shift method; training a preset convolutional neural network model to a convergent state by taking the object fringe pattern samples as model input and taking the real part data values and the imaginary part data values of all pixel points of the object in the object fringe pattern samples as model output to generate an image feature extraction model; In the step of training a preset convolutional neural network model to a convergent state by taking the object fringe pattern samples as model input and taking the real part data values and the imaginary part data values of all pixel points of the object in the object fringe pattern samples as model output to generate an image feature extraction model, the step comprises the following steps: randomly allocating all collected object fringe pattern samples according to a preset allocation ratio to obtain a training sample set and a verification sample set; training the preset convolutional neural network model by taking a first object fringe pattern sample allocated to the training sample set as model input and taking the real part data values and the imaginary part data values of all pixel points of the object in the first object fringe pattern sample as model output to obtain a trained convolutional neural network model; inputting a second object fringe pattern sample allocated to the verification sample set into the trained convolutional neural network model for feature extraction processing to obtain first real part data values and first imaginary part data values of all pixel points of the object output by the trained convolutional neural network model; comparing the first real part data values and the first imaginary part data values of all pixel points of the object output by the trained convolutional neural network model with the real part data values and the imaginary part data values of all pixel points of the object calculated by the phase shift method from the second object fringe pattern sample to obtain a similarity value. The similarity value is compared with a similarity value of a previous round of iterative training, and if a growth amplitude of the similarity is less than a preset threshold, it is determined that the trained convolutional neural network model has been trained to a convergence state, and the trained convolutional neural network model is generated as an image feature extraction model.

2. The method of three-dimensional imaging of dynamic objects according to claim 1, characterized in that, Before the step of inputting the fringe pattern into a preset image feature extraction model for feature extraction processing to obtain a pixel point real part distribution map and a pixel point imaginary part distribution map of the dynamic object respectively, the method further comprises: A de-noising threshold of the fringe pattern is determined according to a background intensity distribution in the fringe pattern, and de-noising processing is performed on the fringe pattern according to the de-noising threshold.

3. The method of three-dimensional imaging of dynamic objects according to claim 1, characterized in that, After the step of comparing the similarity value with a similarity value of a previous round of iterative training, and if a growth amplitude of the similarity is less than a preset threshold, it is determined that the trained convolutional neural network model has been trained to a convergence state, and the trained convolutional neural network model is generated as an image feature extraction model, the method further comprises: A first loss function of the image feature extraction model is constructed according to real part data values and imaginary part data values of all pixel points of an object output by the image feature extraction model and real part data values and imaginary part data values of all pixel points of the object fringe pattern sample obtained by phase shift method; A modulation distribution of the object fringe pattern sample is obtained, and a second loss function of the image feature extraction model is constructed; The first loss function and the second loss function are weighted and integrated to generate a total loss function of the image feature extraction model, and the total loss function is used for model optimization training of the image feature extraction model.

4. The method of three-dimensional imaging of dynamic objects according to claim 1, characterized in that, In the step of training a preset convolutional neural network model to a convergence state by taking the object fringe pattern sample as model input and taking real part data values and imaginary part data values of all pixel points of an object in the object fringe pattern sample as model output to generate an image feature extraction model, the preset convolutional neural network model comprises four U-Net convolutional layer structures, and each U-Net convolutional layer structure is configured with a residual dense network block.

5. The method of three-dimensional imaging of dynamic objects according to claim 4, characterized in that, The residual dense network block comprises a plurality of convolutional layers, and each convolutional layer is provided with a corresponding expansion rate.

6. A device for three-dimensional imaging of a dynamic object, characterized in that The three-dimensional imaging device of the dynamic object comprises: An image capturing module configured to project a fringe pattern on the dynamic object at a preset fringe frequency and capture the fringe pattern; A feature extraction module configured to input the fringe pattern into a preset image feature extraction model for feature extraction processing to obtain a pixel point real part distribution map and a pixel point imaginary part distribution map of the dynamic object; A phase map acquisition module configured to perform phase calculation on pixel points of the dynamic object based on the pixel point real part distribution map and the pixel point imaginary part distribution map, obtain a folded phase map of the dynamic object, and perform phase unwrapping processing on the folded phase map to obtain an absolute phase map of the dynamic object; The three-dimensional imaging module is configured to perform binocular matching and depth calculation on all pixels of the dynamic object according to the absolute phase map, to obtain three-dimensional coordinate values corresponding to each pixel, and to construct a three-dimensional image of the dynamic object according to the three-dimensional coordinate values corresponding to each pixel. Before the step of inputting the fringe pattern into a preset image feature extraction model for feature extraction processing to obtain a real part distribution map and an imaginary part distribution map of the pixels of the dynamic object, the method further comprises the steps of: collecting object fringe pattern samples at a plurality of different frequencies using a pre-built sinusoidal fringe projection measurement system; calculating real part data values and imaginary part data values of all pixels representing the object in the object fringe pattern samples using a phase shift method; training a preset convolutional neural network model to a convergent state using the object fringe pattern samples as model input and the real part data values and the imaginary part data values of all pixels representing the object in the object fringe pattern samples as model output, to generate an image feature extraction model; In the step of training the preset convolutional neural network model to a convergent state using the object fringe pattern samples as model input and the real part data values and the imaginary part data values of all pixels representing the object in the object fringe pattern samples as model output, to generate an image feature extraction model, the method further comprises the steps of: randomly allocating all collected object fringe pattern samples according to a preset allocation ratio to obtain a training sample set and a verification sample set; training the preset convolutional neural network model using a first object fringe pattern sample allocated to the training sample set as model input and the real part data values and the imaginary part data values of all pixels representing the object in the first object fringe pattern sample as model output, to obtain a trained convolutional neural network model; inputting a second object fringe pattern sample allocated to the verification sample set into the trained convolutional neural network model for feature extraction processing, to obtain first real part data values and first imaginary part data values of all pixels representing the object output by the trained convolutional neural network model; comparing the first real part data values and the first imaginary part data values of all pixels representing the object output by the trained convolutional neural network model with real part data values and imaginary part data values of all pixels representing the object calculated by the second object fringe pattern sample using a phase shift method, to obtain a similarity value; comparing the similarity value with a similarity value obtained in a previous iteration, and if a growth rate of the similarity value is less than a preset threshold, determining that the trained convolutional neural network model has been trained to a convergent state, and generating the trained convolutional neural network model as an image feature extraction model.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the method of any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, the computer-readable storage medium comprising: The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Binocular double-frequency complementary three-dimensional surface type measurement method based on fringe projection

    CN113551617A