Three-dimensional reconstruction method, device and equipment of moving object and readable storage medium
By acquiring the multi-frame striped image sequence of moving objects and using deep learning networks for pixel-level motion tracking and phase calculation, the problem of the inability to reconstruct the three-dimensional morphology of moving objects in the traditional PSP method is solved, and high-precision three-dimensional morphology reconstruction is achieved.
Patent Information
- Application Number
- CN202510973199.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-05
AI Technical Summary
The prior art cannot reconstruct the three-dimensional morphology of rigid objects moving in any direction with high precision. The traditional PSP method requires that the object be stationary and cannot be effectively applied to dynamic measurements.
By obtaining the multi-frame striped image sequence of moving objects, selecting reference image frames, using deep learning networks to perform pixel-level motion tracking, constructing a phase calculation equation system, and performing phase expansion and height mapping to obtain three-dimensional morphological data.
It significantly reduces the reconstruction error caused by motion interference, realizes high-quality three-dimensional morphological reconstruction in dynamic scenarios, and provides guarantees for the accurate measurement of high-speed motion targets and industrial detection.
Smart Images

Figure CN120599151A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, apparatus, device and readable storage medium for three-dimensional reconstruction of a moving object. Background Art
[0002] Non-contact 3D surface measurement technology based on structured light has been widely used in fields such as intelligent manufacturing, biomedicine, and virtual reality. Phase shifting profilometry (PSP) has become one of the most popular 3D measurement techniques due to its high point cloud accuracy and data density. PSP projects a series of phase-shifted fringe patterns onto the surface of the object being measured, captures the corresponding deformed fringe images, and uses phase calculation methods to obtain surface height information.
[0003] However, the traditional PSP method requires the object to remain stationary during the projection and acquisition of multiple fringe patterns. If the object moves, errors may be introduced or even cause reconstruction failure. This requirement for a static object limits the application of PSP in dynamic measurements. To solve this problem, researchers have proposed a variety of methods, including improving hardware performance to suppress the influence of motion (such as using digital binary defocusing technology) and developing various motion error estimation algorithms (such as linear least squares fitting, iterative estimation of average phase shift error, real-time compensation methods, etc.). In addition, deep learning methods have also been applied to motion error compensation in recent years, achieving error elimination by establishing a nonlinear mapping relationship between phase and motion.
[0004] Although the above methods have alleviated the motion error problem in PSP systems to some extent, they still have the following key shortcomings: they are unable to reconstruct the 3D shape of rigid objects moving in arbitrary directions with high accuracy. Summary of the Invention
[0005] In view of this, the embodiments of the present application provide a three-dimensional reconstruction method, device, equipment and readable storage medium for a moving object, which can effectively solve the problem that the existing technology cannot reconstruct the three-dimensional shape of a rigid object moving in any direction with high precision.
[0006] In a first aspect, an embodiment of the present application provides a method for three-dimensional reconstruction of a moving object, comprising: Acquire a multi-frame fringe image sequence of the moving object under test at consecutive moments; Selecting a frame from the multi-frame stripe image sequence as a reference image frame, and constructing an image set including the reference image frame and adjacent image frames; Inputting the image set into a deep learning network model, performing pixel-level motion tracking processing based on the reference image frame, and obtaining corresponding motion position information; Constructing a phase calculation equation group based on the motion position information and the intensity information of the corresponding pixel points in the reference image frame, and solving the phase calculation equation group to obtain the relative phase value of the pixel points in the reference image frame; Performing phase unwrapping processing on the relative phase value to obtain absolute phase information of the pixel in the reference image frame; Based on the absolute phase information and system calibration parameters, phase height mapping processing is performed to obtain three-dimensional shape data of the moving object under test at the corresponding moment of the reference image frame.
[0007] In some embodiments, acquiring a multi-frame fringe image sequence of the detected moving object at consecutive moments includes: Controlling the fringe projection device to project a preset number of fringe patterns onto the surface of a moving object at a preset spatial frequency based on a cyclic projection strategy within a preset time period; Synchronously triggering an image acquisition device to acquire an image when each fringe pattern is projected onto the surface of the moving object, thereby obtaining a corresponding deformed fringe image; The deformed fringe images are arranged in time sequence to generate the multi-frame fringe image sequence.
[0008] In some embodiments, selecting a frame from the multi-frame stripe image sequence as a reference image frame and constructing an image set including the reference image frame and adjacent image frames includes: According to the time sequence of the multi-frame fringe image sequence, an image frame at an intermediate moment is selected as the reference image frame; wherein the reference image frame is selected based on the fringe image corresponding to the first preset frequency; Based on the position of the reference image frame in the multi-frame stripe image sequence, at least two image frames before and after the reference image frame are extracted to generate an image set including the reference image frame and adjacent image frames.
[0009] In some embodiments, inputting the image set into a deep learning network model, performing pixel-level motion tracking processing based on the reference image frame, and obtaining corresponding motion position information includes: performing size normalization and downsampling processing on the reference image frame and the adjacent image frame respectively, and extracting corresponding feature representations; Constructing a concatenated context image feature map by concatenating the feature representations; Obtaining a preset query point and position code in the reference image frame; The context image feature map and the query point and position encoding are input into a deep learning network model based on the COTR model to perform encoding and decoding processing based on the Transformer structure to obtain the corresponding motion position information of the query point and position encoding in the adjacent image frames.
[0010] In some embodiments, constructing a phase calculation equation group based on the motion position information and the intensity information of the corresponding pixel points in the reference image frame, and solving the phase calculation equation group to obtain the relative phase value of the pixel points in the reference image frame includes: Determining, based on the motion position information, a corresponding pixel position of each pixel in the reference image frame in the adjacent image frame, and constructing a fringe intensity sequence containing intensity information based on intensity information at the corresponding pixel positions; Based on the fringe intensity sequence and the motion position information, a phase calculation equation group for phase solution is constructed, and a numerical solution process is performed on the phase calculation equation group to obtain the relative phase value of each pixel point in the reference image frame.
[0011] In some embodiments, performing phase unwrapping processing on the relative phase value to obtain absolute phase information of the pixel in the reference image frame includes: Based on a preset multi-frequency heterodyne method, the relative phase values of the pixels in the reference image frame are jointly expanded and calculated to determine the absolute phase values of the pixels in the reference image frame.
[0012] In some embodiments, performing phase height mapping processing based on the absolute phase information and system calibration parameters to obtain three-dimensional shape data of the moving object under test at a time corresponding to the reference image frame includes: Based on the absolute phase value and the preset system calibration parameters, the depth coordinate of the pixel point in the camera coordinate system is calculated through a preset phase-to-height mapping relationship; Based on the pixel coordinates of the pixel points in the reference image frame, using a geometric mapping relationship, calculating the three-dimensional spatial coordinates of each pixel point in the camera coordinate system; Based on the three-dimensional spatial coordinates of all the pixel points, the three-dimensional shape data of the moving object at the corresponding moment of the reference image frame is reconstructed.
[0013] In a second aspect, an embodiment of the present application provides a three-dimensional reconstruction device for a moving object, comprising: An image acquisition module is used to acquire a multi-frame fringe image sequence of the moving object under test at consecutive moments; An image set construction module is used to select a frame from the multi-frame stripe image sequence as a reference image frame and construct an image set including the reference image frame and adjacent image frames; a position information acquisition module, configured to input the image set into a deep learning network model, perform pixel-level motion tracking processing based on the reference image frame, and obtain corresponding motion position information; a model construction module, configured to construct a phase calculation equation group based on the motion position information and the intensity information of the corresponding pixel points in the reference image frame, and solve the phase calculation equation group to obtain the relative phase values of the pixel points in the reference image frame; a fusion processing module, configured to perform phase unwrapping processing on the relative phase value to obtain absolute phase information of pixels in the reference image frame; A mapping module is used to perform phase height mapping processing based on the absolute phase information and system calibration parameters to obtain three-dimensional shape data of the moving object at a corresponding moment in the reference image frame.
[0014] In a third aspect, an embodiment of the present application provides a computer device, comprising a processor and a memory, wherein the memory stores a computer program, and the processor is configured to execute the computer program to implement the three-dimensional reconstruction method of a moving object according to the first aspect.
[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, and when the computer program is executed on a processor, the three-dimensional reconstruction method of a moving object according to the first aspect is implemented.
[0016] The embodiments of the present application have the following beneficial effects: first, multiple frames of stripe images projected on the surface of a moving object at continuous moments are collected, and a frame is selected from the image sequence as a reference image frame, and its adjacent frames are extracted to form an image set; then, the image set is input into a pre-trained deep learning network model, and the pixel-level position is tracked based on the reference frame to obtain the motion position information of each pixel in the reference frame in the adjacent frames; then, the motion position information and the intensity information of the corresponding pixels in the reference frame are used to jointly construct a phase calculation equation group, and the relative phase value is obtained by solving the phase calculation equation group; then, the relative phase value is phase-unwrapped to obtain the absolute phase information of the pixel point in the reference image frame; finally, the absolute phase is combined with the system calibration parameters, and the phase height mapping process is performed according to the preset phase height mapping relationship to obtain the three-dimensional shape data of the measured moving object at the corresponding moment of the reference image frame. The method of the present application significantly reduces the reconstruction error caused by motion interference, and can also output high-quality three-dimensional shape in dynamic scenes, providing a strong guarantee for the accurate measurement of high-speed moving targets and industrial inspection. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 A flowchart of a method for three-dimensional reconstruction of a moving object according to an embodiment of the present application is shown; Figure 2 A flowchart showing the workflow of the model in the method for 3D reconstruction of a moving object according to an embodiment of the present application is shown; Figure 3 A timing diagram illustrating a fringe pattern cyclic projection strategy used in a method for 3D reconstruction of a moving object according to an embodiment of the present application is shown; Figure 4 A schematic diagram of a geometric object "Qin Nu" used in a 3D reconstruction experiment in a 3D reconstruction method for a moving object according to an embodiment of the present application is shown; Figure 5 A schematic diagram showing a comparison of tracking errors between the network tracking method and the optical flow method in the 3D reconstruction method of a moving object according to an embodiment of the present application is shown; Figure 6 A schematic diagram of a multi-target dynamic scene and corresponding motion tracking and phase error distribution in a method for 3D reconstruction of a moving object according to an embodiment of the present application is shown; Figure 7 A schematic diagram showing comparison results of different 3D reconstruction methods in the 3D reconstruction method for a moving object according to an embodiment of the present application is shown; Figure 8 A structural schematic diagram of a method for three-dimensional reconstruction of a moving object according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0019] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments.
[0020] The components of the embodiments of the present application generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but rather merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.
[0021] Hereinafter, the terms "including", "having" and their cognates used in various embodiments of the present application are intended only to indicate specific features, numbers, steps, operations, elements, components or combinations of the aforementioned items, and should not be understood as excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the aforementioned items or adding the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the aforementioned items. In addition, the terms "first", "second", "third" and the like are only used to distinguish descriptions and should not be understood as indicating or implying relative importance.
[0022] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art to which the various embodiments of the present application belong. The terms (such as those defined in generally used dictionaries) will be interpreted as having the same meaning as in the context of the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning unless clearly defined in the various embodiments of the present application.
[0023] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.
[0024] Taking into account the problem that the existing technology cannot reconstruct the three-dimensional morphology of rigid objects moving in any direction with high precision, a three-dimensional reconstruction method of a moving object is proposed. By continuously projecting multi-frequency phase-shifted stripes and synchronously acquiring multiple frames of deformed images, the image set is input into a deep learning network for pixel-level motion tracking to extract motion position information and combine the intensity information to construct a phase calculation equation group, and then perform phase unwrapping and phase height mapping to obtain dynamic three-dimensional morphology data. Through the above means, the method of the present application can effectively eliminate motion interference and achieve high-precision, real-time three-dimensional morphology reconstruction of targets moving in any direction.
[0025] The following describes the three-dimensional reconstruction method of the moving object with reference to some specific embodiments.
[0026] Figure 1 A flow chart of a method for 3D reconstruction of a moving object according to an embodiment of the present application is shown. Exemplarily, the method for 3D reconstruction of a moving object includes the following steps: Step S100: acquiring a multi-frame fringe image sequence of the moving object under test at consecutive moments.
[0027] A fringe image sequence refers to a series of modulated fringe images captured by an image acquisition device after a fringe projection device continuously projects a specific fringe pattern onto the surface of a moving object at multiple times. To exemplify this, while the moving object is within visual range, the image acquisition device continuously records grayscale variations on its surface due to the projected fringe patterns. These images are then organized in a time sequence based on the capture times, forming a multi-frame fringe image sequence. This sequence contains not only information about the spatial motion of the target object, but also the modulation characteristics of the fringe pattern induced by its surface topography.
[0028] In an optional embodiment, step S100 includes the following sub-steps: Step S101 : within a preset time period, controlling a fringe projection device to project a preset number of fringe patterns onto a surface of a moving object at a preset spatial frequency based on a cyclic projection strategy.
[0029] The cyclic projection strategy refers to repeatedly projecting fringe patterns of different frequencies in a fixed order within a certain time window to ensure that the images at each frequency are evenly distributed in the time dimension and adapt to the dynamic characteristics of the moving object. The preset spatial frequencies include, for example (Unit: pixel), the preset number can be selected as 5, 7 or other phase shift numbers. Exemplarily, the fringe projection device is controlled to project the fringe projections with a frequency of Five stripe patterns of frequency are projected in sequence. The five stripe patterns are circulated in this way to form a time sequence arrangement of stripe patterns of different frequencies.
[0030] Step S102 : synchronously triggering an image acquisition device to acquire an image when each fringe pattern is projected onto the surface of the moving object, thereby obtaining a corresponding deformed fringe image.
[0031] The image acquisition device refers to an imaging module used to receive and record the modulated fringe pattern on an object's surface. It typically includes a high-speed camera and trigger control circuitry, enabling precise timing of each fringe pattern's image capture. A deformed fringe image is the resulting image of the fringe pattern geometrically distorted when it encounters the object's surface topography. Exemplarily, through a synchronized triggering mechanism with the projection device, the image acquisition device automatically exposes and reads the fringe pattern at the same moment it is projected onto the object's surface, thereby capturing the corresponding fringe deformation state. Each trigger records a single frame, and continuous triggering forms a time-series image set.
[0032] Step S103 , arranging the deformed fringe images in chronological order to generate a multi-frame fringe image sequence.
[0033] Among them, the image sequence is sorted by time to ensure the subsequent phase error estimation and temporal feature extraction. The image sequence contains not only spatial coding information, but also implicitly contains temporal change information. Specifically, the continuously acquired image frames are numbered according to their acquisition time stamps, for example , and store them in chronological order to construct a multi-frame stripe image sequence ,in Represents the deformed stripe image of the i-th frame.
[0034] Step S200 , selecting a frame from a multi-frame stripe image sequence as a reference image frame, and constructing an image set including the reference image frame and adjacent image frames.
[0035] Among them, the reference image frame refers to the target frame determined as the current three-dimensional reconstruction processing in the entire multi-frame stripe image sequence, and serves as the basic image frame for subsequent motion tracking, phase estimation and reconstruction reasoning. The image set is an image set composed of several adjacent image frames selected forward and backward with the reference image frame as the center, which is used to construct temporal context information and support motion estimation and phase fitting tasks based on deep learning. Specifically, the construction of the image set usually follows a symmetric window strategy. For example, the first frequency is used as a reference, and two frames are selected before and after with the reference image frame as the center, eventually forming a 5-frame image set. The image set also contains the intensity change characteristics and stripe distortion characteristics of the reference image frame at different time positions.
[0036] In an optional embodiment, step S200 includes the following sub-steps: Step S201 : selecting an image frame at an intermediate moment as a reference image frame according to the time sequence of a multi-frame fringe image sequence.
[0037] The first preset frequency refers to the first fringe frequency used to construct the reference image frame during the fringe pattern projection process, which is usually the lowest frequency (for example: ), used to improve reconstruction stability and motion compensation accuracy. The selection of intermediate frames ensures that the reference image frame is centered within the entire image sequence, facilitating the construction of a symmetrical temporal image set. For example, if a multi-frame streak image sequence contains 15 frames, the image projected at the first preset frequency can be selected as the reference image frame at the eighth frame. This avoids the potential for incomplete information in edge frames, thereby improving overall tracking robustness.
[0038] Step S202 : based on the position of the reference image frame in the multi-frame stripe image sequence, extract at least two image frames before and after the reference image frame to generate an image set including the reference image frame and adjacent image frames.
[0039] Among them, adjacent image frames refer to the forward and backward image frames that are close to the reference image frame in time, which together constitute the temporal input in the motion estimation and phase compensation model. By selecting adjacent frames, the modeling capability of the dynamic changes of the object surface can be enhanced, which is especially suitable for the reconstruction of the surface structure of objects with continuous micro-motion or periodic motion. For example, if the reference image frame is the 8th frame in the sequence, the 6th and 7th frames can be selected forward, and the 9th and 10th frames can be selected backward to form an image set together with the reference image frame. The image set is used as the input of the deep learning network model in the subsequent steps to complete the dynamic feature alignment and pixel-level motion estimation between images.
[0040] In step S300 , the image set is input into the deep learning network model, pixel-level motion tracking processing based on the reference image frame is performed, and corresponding motion position information is obtained.
[0041] Among them, pixel-level motion tracking refers to the precise identification of the corresponding position of each pixel in the image set between different frames, which is used to quantify the motion changes of the surface points of the object; the deep learning network model refers to a neural network based on the COTR (Coordinates Transformer) architecture, which realizes high-precision point correspondence calculation by sharing the CNN and Transformer encoding and decoding structures. For example, Figure 2 As shown in the figure, for a pair of stripe images to be matched before and after motion, the size normalization and downsampling are first performed to extract their respective high-dimensional feature representations, and then the context image feature map is formed through splicing operations; then, these features are It is input into the Transformer encoder based on the COTR architecture for information fusion between multiple frames. At the same time, the system defines a set of query points h(x,y) on the reference image frame, and combines them with the position encoding information p as input to the Transformer decoder. The decoder calculates the corresponding positions of the query points in adjacent image frames through the self-attention and cross-attention mechanisms to achieve cross-frame pixel-level matching. Finally, the matching function D is used to output the exact position of each query point in the target image frame. , thereby obtaining the motion displacement information of the surface points of the object.
[0042] In an optional embodiment, step S300 includes the following sub-steps: In step S301 , the reference image frame and the adjacent image frames are resized and downsampled to extract corresponding feature representations.
[0043] Among them, size normalization and downsampling refer to scaling and downsampling the original 256×256 pixel image to a 16×16 resolution feature map to reduce computational complexity; feature representation refers to the high-dimensional spatial features extracted by sharing CNN, which contains rich local texture and structural information of the stripe pattern.
[0044] Step S302 : Concatenate the feature representations to construct a concatenated context image feature map.
[0045] The splicing operation involves stacking the feature maps of the reference image frame and each adjacent image frame at the same channel dimension to form a contextual feature set across time frames. The contextual image feature map is a set of feature expressions concatenated in the time dimension, providing clues about temporal associations across frames. Specifically, the downsampled feature maps corresponding to each frame are concatenated channel-wise into a multi-channel tensor. For example, if the feature dimensions of each frame are C×H×W, the dimensions of the feature map after splicing five frames are (5×C)×H×W.
[0046] Step S303: Obtain the preset query point and position code in the reference image frame.
[0047] The query point and position encoding refers to a predefined set of pixel coordinates on the reference image frame, combined with a position encoding function P to generate a high-dimensional position representation associated with each query point. Specifically, N points are uniformly sampled from the reference image frame as query points, and each point is assigned a two-dimensional position vector. For example, the position encoding is generated using sine and cosine encoding, which is then used as part of the input to the subsequent Transformer.
[0048] In step S304, the context image feature map and the query point and position codes are input into a deep learning network model based on the COTR model to perform encoding and decoding processing based on the Transformer structure to obtain the corresponding motion position information of the query point and position codes in adjacent image frames.
[0049] Among them, the COTR model is a sparse point-level registration network based on the Transformer architecture, which can achieve high-precision motion tracking of query points in image sequences; the Transformer structure consists of an encoder and a decoder. The encoder fuses information of contextual image features, and the decoder predicts the position based on the query point and its position encoding.
[0050] Specifically, the concatenated context image feature map is first input into the encoder part to extract the multi-frame joint feature expression; then the query point position code is input into the decoder, which interacts with the encoder output to predict the pixel position coordinates of each query point in the adjacent frame. The output coordinate value is the motion position information, that is, the position of the object in the rest of the fringe image. For example, given a set of fringe images before and after the object moves , cropped and resized to 256 × 256 images, and converted into downsampled feature maps of size 16 × 16 × 256 with a shared CNN backbone. Then, the representations of the two corresponding images are concatenated side by side to form a feature map of size 16 × 32 × 256, and combined with the positional encoding P of the coordinate function to produce the context feature map C.
[0051] The context feature map is obtained by the following formula:
[0052] in: Represents the following feature map; Represents the stripe image before the object moves; Represents the stripe image after the object moves; express Positional encoding; For example, for a set of stripe images before and after the object moves, we first pass a shared CNN backbone Feature extraction and downsampling are performed to obtain a feature map with a size of 16×32 (height × width) and 256 channels. Based on the spatial dimensions of this feature map, a normalized two-dimensional grid coordinate map is generated. The grid is normalized to [0, 1) in the width (X) direction and [0, 2) in the height (Y) direction (i.e., the normalized coordinate range in the Y direction is twice that in the X direction). A positional encoding method (such as sinusoidal encoding or learnable encoding) is applied to this grid coordinate to generate a positional encoded feature P with the same spatial dimensions as the feature map (16×32) and 256 channels. Finally, the positional encoded feature map is fused with the feature map extracted by the CNN backbone (e.g., element-wise addition or channel concatenation) to generate the final contextual feature map C.
[0053] The corresponding motion position information of the query point and position code in adjacent image frames is obtained by the following formula:
[0054] in: Represents the query point and the corresponding motion position information of the position code in the adjacent image frames; Represents the overall process function of COTR; represents the fully connected layer of the shared CNN backbone; is the decoder of Transformer; Represents the position code associated with the query point x; is the encoder of Transformer; Represents the following feature map; In other words, the context feature map is fed into the Transformer encoder generate , the position code P of the coordinate function is used to generate the position code related to the query point x .connect and , using the Transformer decoder Explain the results. Finally, through the fully connected layer Process the output of the Transformer decoder to obtain an estimate of the corresponding point x.
[0055] In an optional embodiment, in order to obtain the motion changes of the object in the entire image, the pixel query points on the entire object can be input at one time, and the Transformer decoder and decoder It can be converted into multiple coordinates. In another optional implementation, the COTR network can be trained in a supervised manner, using the mean square error (MSE) as a loss term to measure the difference between the predicted point of the object's motion and the actual point after movement. Furthermore, by inverting the image before and after the movement, the predicted point is used as the input point, and the error between the new predicted point and the initial input point is calculated, which is recorded as the recurrent mean square error (REMSE). Finally, the following loss function equation is obtained and used to train the deep learning network model:
[0056]
[0057]
[0058] in, Represents the loss function of model training; Represents the loss term, which is used to measure the difference between the predicted point of the object's motion and the point after the actual movement. The error between the new predicted point and the initial input point is calculated; Indicates the error corresponding to the cycle consistency principle; Represents the mean square error (MSE) of each set of query images to the target image prediction points; After representing the image before and after the flipping motion, the predicted point is used as the input point to obtain the new recurrent mean square error (REMSE).
[0059] Step S400: constructing a phase calculation equation group based on the motion position information and the intensity information of the corresponding pixel points in the reference image frame, and solving the phase calculation equation group to obtain the relative phase value of the pixel points in the reference image frame.
[0060] Intensity information refers to the grayscale values of the fringe images recorded at the same pixel location in the reference image frame and adjacent image frames. The phase calculation equations incorporate motion-induced displacement information and combine the intensity values of each pixel in the five phase-shifted fringe patterns to construct a mathematical expression for error estimation. This expression reflects the pattern of pixel intensity variation with phase in the fringe pattern. Specifically, solving the phase calculation equations yields the relative phase value affected by motion.
[0061] In an optional embodiment, step S400 includes the following sub-steps: Step S401 : determining the corresponding pixel position of each pixel point in the reference image frame in the adjacent image frame based on the motion position information, and constructing a fringe intensity sequence according to the intensity information at the corresponding pixel position.
[0062] Among them, adjacent image frames refer to the several frames of fringe images that are adjacent to the reference image frame in the time dimension, which are used to obtain the pixel positions before and after the movement; the fringe intensity sequence refers to the set of pixel coordinates of the same physical point in different frame images, which reflects the intensity change process of the pixel as the object moves. For example, .
[0063] Step S402 : constructing a phase calculation equation group for phase solution based on the fringe intensity sequence, performing numerical solution processing on the phase calculation equation group, and obtaining the relative phase value of each pixel point in the reference image frame.
[0064] Among them, the phase calculation equation group is an analytical equation constructed based on the multi-frame intensity input in phase shift profilometry, which is used to extract accurate phase information under the condition of object motion. For example, assuming that the acquired image sequence is divided into continuous short time windows, the phase change caused by the object's motion in the height direction within the same time window can be considered constant. Indicates the phase error. Specifically, for the five fringe images captured continuously , by the object at point The phase error introduced by the height direction movement at , you can use and Then, the intensity of the fringe pattern can be expressed by the following equation: . in, It is recorded as a certain point in the fringe image; A is the background intensity, B is the modulation degree; is the phase value on the reference plane; It is the phase change caused by the object's height modulation; Represents a function that describes the motion information of an object; where, Indicates that the third frame is the reference frame; Through the simultaneous equations, we can get two The nonlinear equation is:
[0065] Given that Very small, satisfying and , the above formula is further simplified to:
[0066] in, and are the coefficients related to the known parameters, omitting the higher-order minor terms , will obtain Value Substitution , and the final phase is obtained:
[0067] in, Indicates the phase value; Representation and Error Related functions; Step S500 , performing phase unwrapping processing on the relative phase value to obtain absolute phase information of the pixel in the reference image frame.
[0068] Among them, phase unwrapping processing refers to eliminating the discontinuity caused by the phase value jumping within the periodic range, thereby restoring a spatially continuous and unambiguous phase distribution; absolute phase information refers to the complete and continuous phase value, which can subsequently be used to map to the height information of the object surface.
[0069] In an optional embodiment, step S500 includes the following sub-steps: Based on a preset multi-frequency heterodyne method, the relative phase values of the pixels in the reference image frame are jointly expanded and calculated to determine the absolute phase values of the pixels in the reference image frame.
[0070] Among them, phase unwrapping refers to the process of eliminating phase ambiguity by solving the relative phase values obtained under multiple frequency conditions when they are periodic and non-unique, so as to restore them to absolute phase values with global continuity. Specifically, the relative phase value changes continuously, but due to The function calculation range is , the calculated phase value is wrapped into arrive In the process, every time it exceeds π, it will return to . Similar to the characteristics of a periodic function, this makes the phase seen in the image segmented. That is, the relative phase is only valid within one period, and the relative phase is the phase change within this period, and does not reflect the index information of the entire stripe. Due to its periodicity, multiple points may have the same wrapped phase value, resulting in phase ambiguity. In order to obtain a globally continuous true phase value, the following formula is used to remove the discontinuity of the wrapped phase value:
[0071] in, Represents the phase value after expansion; Indicates the initial calculated package phase value; represents the stripe order; Represents the phase period constant.
[0072] By determining the fringe order at the pixel position, the wrapped phase can be unwrap. This embodiment uses a multi-frequency heterodyne unwrapping method. Through the heterodyne principle, the phase difference of different frequencies can provide a wider range of phase information. As long as a specific fringe frequency is reasonably selected, the phase distribution within a single cycle can be obtained. The fringe order corresponding to other frequencies can be inversely solved through the global continuous phase information. Ultimately, it is possible to rely on layer-by-layer unwrapping to expand the phase range from low frequency to high frequency, thereby calculating a globally consistent absolute phase.
[0073] Step S600 , performing phase height mapping processing based on the absolute phase information and the system calibration parameters to obtain the three-dimensional shape data of the moving object under test at the corresponding moment of the reference image frame.
[0074] Absolute phase information refers to numerical data that directly reflects the phase distribution on an object's surface after phase unwrapping. System calibration parameters refer to the parameter set obtained by pre-calibrating the geometric relationship between the projector and camera, including information such as the camera's focal length, principal point coordinates, and projector projection angle. Phase height mapping processing combines the acquired absolute phase values with the system calibration parameters to convert phase differences into depth information in physical space, thereby providing the spatial position basis for 3D reconstruction.
[0075] In an optional embodiment, step S600 includes the following sub-steps: Step S601 : Based on the absolute phase value and preset system calibration parameters, the depth coordinate of the pixel point in the camera coordinate system is calculated through a preset phase-to-height mapping relationship.
[0076] The phase-to-height mapping relationship is a function that establishes a numerical correspondence between phase values and corresponding depths based on the projection inclination and camera viewing angle measured during system calibration. The depth coordinate refers to the physical distance of a pixel point on the object's surface along the camera's line of sight. As an example, the absolute phase value of each pixel in the reference image frame is input into the mapping model. Combined with the system calibration parameters, the depth coordinate value of that pixel point relative to the camera is calculated through a table lookup or interpolation calculation, completing the corresponding phase-to-depth conversion.
[0077] Step S602 : Based on the pixel coordinates of the pixel points in the reference image frame, the three-dimensional space coordinates of each pixel point in the camera coordinate system are calculated using a geometric mapping relationship.
[0078] Pixel coordinates refer to the position index of a pixel in the reference image frame on the two-dimensional image plane, while the geometric mapping relationship refers to the mathematical relationship that links the two-dimensional coordinates on the image plane to the three-dimensional coordinates in the camera coordinate system, which is usually determined by the camera's intrinsic parameters and distortion correction parameters. This process is responsible for expanding the two-dimensional depth information into three-dimensional spatial coordinates to achieve a true mapping of spatial positions. Specifically, using the internal parameters such as the focal length and principal point coordinates obtained by camera calibration, a pixel point (u, v) and its corresponding depth value Z are brought into the geometric mapping model. The three-dimensional coordinates (X, Y, Z) of the pixel point in the camera coordinate system are calculated using the back-projection algorithm, and the two-dimensional to three-dimensional coordinate conversion is completed point by point.
[0079] Step S603 : reconstructing the three-dimensional shape data of the moving object at the corresponding moment of the reference image frame based on the three-dimensional spatial coordinates of all the pixel points.
[0080] Here, the three-dimensional spatial coordinates refer to the (X, Y, Z) values of each pixel in the camera coordinate system; the three-dimensional shape data refers to the point cloud composed of the set of all pixel coordinates, which is used to represent the complete three-dimensional outline of the object surface at that moment. This process is responsible for fusing the three-dimensional coordinates of independent pixels into an overall point cloud, presenting the continuous shape of the object surface. Exemplarily, all pixels and three-dimensional coordinates in the reference image frame are collected, and these coordinates are merged and organized according to their spatial position relationship to form a 3D model dataset containing a set of points, and finally outputting the complete three-dimensional shape data representing the object's appearance at that moment.
[0081] In an optional embodiment, the method of the present application can be further understood through the following examples: like Figure 3 As shown in the figure, it shows that within a continuous time window, the projector projects phase-shifted fringe patterns of multiple frequencies in a preset order and repeats the process after a complete cycle to achieve synchronized image acquisition of moving objects: Step 1: Four-order fringe patterns with three different frequencies , and according to Figure 3 The fringe projection strategy described in the method is arranged in this way. The projector projects the fringe pattern onto the surface of the moving object under test in a circular manner, and the camera synchronously records the deformed fringe pattern after modulation by the object. Step 2: Taking the middle fringe pattern as the benchmark, the COTR model is used to obtain the object motion information in each fringe pattern. Step 3: Substitute the motion information into the phase calculation equation group to obtain the phases of the three frequencies respectively. Step 4: Use the phase diagrams of any two frequencies to assist in solving the absolute phase information of the third frequency. Finally, the correct three-dimensional morphology information of the object is obtained through phase-height mapping. Step 5: Move the fringe selection window to the right to skip the three stripes, and repeat the operations of Step 2 to Step 4.
[0082] In an optional embodiment, as Figure 4 As shown, the method of the present application can be further understood through the following examples: Select a geometric object with a complex surface - Qin Nu. Figure 4 As shown in (a), the linear stage is placed in front of the camera optical axis and at a 40° angle to the camera optical axis. The object is then moved toward the camera at a speed of 1 mm to simulate the object's motion in the xy and z directions. The continuously acquired image frames are numbered according to their acquisition time stamps, for example , and store them in chronological order to construct a multi-frame stripe image sequence ,in Represents the i-th frame of deformed stripe image. Taking a set of stripe images I3 and I4 under continuous motion as an example, Figure 4 (b) shows the motion vector field estimated by the COTR network in the horizontal and vertical directions. To further verify its tracking accuracy, the traditional optical flow method is used to track and compare the pure object images at the corresponding positions. Figure 4 (c) shows the tracking results of the COTR network on the stripe image and the tracking error of the optical flow method on the pure object image. From the error distribution point of view, the pixel error is mainly concentrated in The COTR network can effectively eliminate the stripe interference and still achieve accurate object position tracking in images with periodic stripe background.
[0083] Afterwards, if Figure 5 As shown in Figure 2, a variety of algorithms were used to reconstruct the moving object, including the traditional PSP algorithm, Guo's method, Lu's method and the algorithm proposed in this application. Figure 5As shown in (a), the comparison between the traditional optical flow tracking path based on pure object images and the pixel-level tracking path based on the COTR network of this application is shown: local A, local B, and reconstruction results. Experiments show that significant reconstruction errors are introduced during the movement of objects: the standard PSP method has obvious stripe artifacts and contour distortions, especially in the boundary areas of objects (local A, local B). Although Guo's method reduces artifacts to a certain extent, it is still difficult to accurately reconstruct edge details. Lu's method effectively improves the matching problem in the boundary area by aligning the stripe images of moving objects. However, due to the failure to compensate for the additional phase offset caused by the Z-direction motion, obvious ripple artifacts still remain on the surface of the object. In contrast, the method proposed in this application combines motion compensation with pixel-level tracking mechanisms, which can more accurately restore the true contour of the object and effectively suppress the geometric distortion caused by motion. The cross-sectional contour curve was extracted at pixel row y=260, and the corresponding depth change is shown as follows. Figure 5 As shown in (b) in the figure, it can be seen from the contour diagram that the curve Proposed reconstructed by the method of the present application is highly consistent with the true value truth, while other methods have deviations and fluctuations to varying degrees. In addition, as Figure 5 (c) shows the phase error caused by motion in the Z direction.
[0084] For the movement of an object in any direction, the motion trajectories of different pixels and the phase errors caused by them may be different. In order to comprehensively evaluate the adaptability and robustness of the method proposed in this application in various dynamic scenes, two sets of representative complex motion experiments were designed, such as Figure 6 As shown, the details are as follows: The first set of experiments is a multi-target scene with different motion characteristics, such as Figure 6 As shown in (a), it includes a stationary astronaut statue and a "Marseille" statue rotating around the Y axis with an angular velocity of 0.4 rad / s. This experiment is used to evaluate the reconstruction performance of the algorithm under different motion states of multiple targets. The second group of experimental objects is a fan model with a smooth surface and weak texture, which rotates slowly around the X axis at an angular velocity of 0.02 rad / s, aiming to study the performance of this method under smooth surface conditions. Figure 6 As shown in (b), it shows the COTR network based on Pixel-level motion tracking results between time points; e.g. Figure 6 (c) shows the phase error caused by motion in the Z direction.
[0085] The reconstruction results of PSP, Guo, Lu, the proposed method and the 12-step phase shift method at rest are shown in Figure 7As shown in the experimental results, the standard PSP method will produce serious stripe artifacts and blurred boundaries under strong dynamic interference, especially in the edge area of the object where there is obvious discontinuity, such as Figure 7 As shown in (a) and (f) in the figure, the Guo method suppresses some artifacts to a certain extent by indirectly compensating for motion errors, such as Figure 7 However, due to the failure to achieve accurate boundary alignment, there are still large errors in the edge area of the object. The Lu method successfully reconstructs the overall shape of the object, as shown in (b) and (g). Figure 7 However, due to the lack of error compensation in the z direction, periodic ripple interference is still retained in the reconstruction result. In contrast, the algorithm proposed in this application also effectively eliminates the artifact interference caused by different motion states, such as Figure 7 As shown in (d) (i). The reconstruction result of the 12-step phase shift method at rest is as follows Figure 7 The experimental results verify the adaptability and stability of this method in complex dynamic scenes with multiple objects and multiple directions.
[0086] Figure 8 A schematic diagram of the structure of a 3D reconstruction apparatus for a moving object according to an embodiment of the present application is shown. Exemplarily, the 3D reconstruction apparatus 100 for a moving object includes: The image acquisition module 110 is used to acquire a multi-frame fringe image sequence of the moving object under test at consecutive moments; An image set construction module 120 is configured to select a frame from the multi-frame stripe image sequence as a reference image frame and construct an image set including the reference image frame and adjacent image frames; a position information acquisition module 130 for inputting the image set into a deep learning network model, performing pixel-level motion tracking processing based on the reference image frame, and obtaining corresponding motion position information; a model building module 140 for building a phase calculation equation group based on the motion position information and the intensity information of the corresponding pixel points in the reference image frame, and solving the phase calculation equation group to obtain the relative phase values of the pixel points in the reference image frame; A fusion processing module 150 is configured to perform phase unwrapping processing on the relative phase value to obtain absolute phase information of the pixel in the reference image frame; The mapping module 160 is configured to perform phase height mapping processing based on the absolute phase information and system calibration parameters to obtain three-dimensional shape data of the moving object at a corresponding moment in the reference image frame.
[0087] It can be understood that the apparatus of this embodiment corresponds to the method of the above embodiment, and the options in the above embodiment are also applicable to this embodiment, so they will not be described again here.
[0088] The present application also provides a computer device. Exemplarily, the computer device includes a processor and a memory, wherein the memory stores a computer program, and the processor runs the computer program to enable the computer device to execute the functions of each module in the above method or the above apparatus.
[0089] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0090] The memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM). The memory is used to store computer programs, and the processor can execute the computer programs accordingly after receiving an execution instruction.
[0091] This application also provides a computer-readable storage medium for storing the computer program used in the aforementioned computer device. For example, the computer-readable storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0092] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in an alternative implementation, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the structure diagram and / or flowchart, and the combination of boxes in the structure diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0093] In addition, the functional modules or units in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0094] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a smart phone, personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.
[0095] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A method for three-dimensional reconstruction of a moving object, characterized in that: The method comprises: Acquire a multi-frame fringe image sequence of the moving object under test at consecutive moments; Selecting a frame from the multi-frame stripe image sequence as a reference image frame, and constructing an image set including the reference image frame and adjacent image frames; Inputting the image set into a deep learning network model, performing pixel-level motion tracking processing based on the reference image frame, and obtaining corresponding motion position information; Constructing a phase calculation equation group based on the motion position information and the intensity information of the corresponding pixel points in the reference image frame, and solving the phase calculation equation group to obtain the relative phase value of the pixel points in the reference image frame; Performing phase unwrapping processing on the relative phase value to obtain absolute phase information of the pixel in the reference image frame; Based on the absolute phase information and system calibration parameters, phase height mapping processing is performed to obtain three-dimensional shape data of the moving object under test at the corresponding moment of the reference image frame.
2. The method for 3D reconstruction of a moving object according to claim 1, wherein: The step of obtaining a sequence of multiple fringe images of the moving object under test at consecutive moments includes: Controlling the fringe projection device to project a preset number of fringe patterns onto the surface of a moving object at a preset spatial frequency based on a cyclic projection strategy within a preset time period; Synchronously triggering an image acquisition device to acquire an image when each fringe pattern is projected onto the surface of the moving object, thereby obtaining a corresponding deformed fringe image; The deformed fringe images are arranged in time sequence to generate the multi-frame fringe image sequence.
3. The method for 3D reconstruction of a moving object according to claim 1, wherein: The step of selecting a frame from the multi-frame stripe image sequence as a reference image frame and constructing an image set including the reference image frame and adjacent image frames includes: According to the time sequence of the multi-frame fringe image sequence, an image frame at an intermediate moment is selected as the reference image frame; wherein the reference image frame is selected based on the fringe image corresponding to the first preset frequency; Based on the position of the reference image frame in the multi-frame stripe image sequence, at least two image frames before and after the reference image frame are extracted to generate an image set including the reference image frame and adjacent image frames.
4. The method for 3D reconstruction of a moving object according to claim 3, wherein: The step of inputting the image set into a deep learning network model, performing pixel-level motion tracking processing based on the reference image frame, and obtaining corresponding motion position information includes: performing size normalization and downsampling processing on the reference image frame and the adjacent image frame respectively, and extracting corresponding feature representations; Constructing a concatenated context image feature map by concatenating the feature representations; Obtaining a preset query point and position code in the reference image frame; The context image feature map and the query point and position encoding are input into a deep learning network model based on the COTR model to perform encoding and decoding processing based on the Transformer structure to obtain the corresponding motion position information of the query point and position encoding in the adjacent image frames.
5. The method for 3D reconstruction of a moving object according to claim 1, wherein: The step of constructing a phase calculation equation group based on the motion position information and the intensity information of the corresponding pixel points in the reference image frame, and solving the phase calculation equation group to obtain the relative phase value of the pixel points in the reference image frame includes: Determining, based on the motion position information, the corresponding pixel position of each pixel in the reference image frame in the adjacent image frame, and constructing a fringe intensity sequence according to intensity information at the corresponding pixel positions; A phase calculation equation group for phase solution is constructed based on the fringe intensity sequence, and a numerical solution process is performed on the phase calculation equation group to obtain a relative phase value of each pixel point in the reference image frame.
6. The method for 3D reconstruction of a moving object according to claim 1, wherein: The performing phase unwrapping processing on the relative phase value to obtain absolute phase information of the pixel in the reference image frame includes: Based on a preset multi-frequency heterodyne method, the relative phase values of the pixels in the reference image frame are jointly expanded and calculated to determine the absolute phase values of the pixels in the reference image frame.
7. The method for 3D reconstruction of a moving object according to claim 6, wherein: The performing of phase height mapping processing based on the absolute phase information and the system calibration parameters to obtain three-dimensional shape data of the moving object under test at a time corresponding to the reference image frame includes: Based on the absolute phase value and the preset system calibration parameters, the depth coordinate of the pixel point in the camera coordinate system is calculated through a preset phase-to-height mapping relationship; Based on the pixel coordinates of the pixel points in the reference image frame, using a geometric mapping relationship, calculating the three-dimensional spatial coordinates of each pixel point in the camera coordinate system; Based on the three-dimensional spatial coordinates of all the pixel points, the three-dimensional shape data of the moving object at the corresponding moment of the reference image frame is reconstructed.
8. A three-dimensional reconstruction device for a moving object, characterized in that: include: An image acquisition module is used to acquire a multi-frame fringe image sequence of the moving object under test at consecutive moments; An image set construction module is used to select a frame from the multi-frame stripe image sequence as a reference image frame and construct an image set including the reference image frame and adjacent image frames; a position information acquisition module, configured to input the image set into a deep learning network model, perform pixel-level motion tracking processing based on the reference image frame, and obtain corresponding motion position information; a model construction module, configured to construct a phase calculation equation group based on the motion position information and the intensity information of the corresponding pixel points in the reference image frame, and solve the phase calculation equation group to obtain the relative phase values of the pixel points in the reference image frame; a fusion processing module, configured to perform phase unwrapping processing on the relative phase value to obtain absolute phase information of pixels in the reference image frame; A mapping module is used to perform phase height mapping processing based on the absolute phase information and system calibration parameters to obtain three-dimensional shape data of the moving object at a corresponding moment in the reference image frame.
9. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores a computer program, and the processor is configured to execute the computer program to implement the three-dimensional reconstruction method for a moving object according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The device stores a computer program, which, when executed on a processor, implements the three-dimensional reconstruction method for a moving object according to any one of claims 1 to 7.
Citation Information
Cited By
Environment prediction method, electronic equipment and storage medium
CN121391959A
Dynamic scene high-precision three-dimensional reconstruction method based on inter-frame geometric constraint and motion tracking
CN122023461A