Photogrammetry method, photogrammetry device and electronic equipment

By using a motion prediction model to predict camera pose information in a monocular scanning method and combining it with initial 3D coordinates, monocular tracking photogrammetry without coded points is achieved, solving the problems of high cost and complex operation in existing technologies and realizing efficient and accurate 3D reconstruction.

CN122130045APending Publication Date: 2026-06-02SCANTECH (HANGZHOU) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SCANTECH (HANGZHOU) CO LTD
Filing Date
2026-02-04
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing monocular scanning methods rely on coded points for high-precision point identification and positioning, resulting in high costs and complex operation procedures.

Method used

By acquiring image sequences of the target object, the camera pose information is predicted using a motion prediction model. Combined with the initial 3D coordinates, monocular tracking photogrammetry without coding points is achieved. The motion prediction model includes the motion dynamics model of the monocular camera, a temporal neural network, and an inertial sensor model for prediction.

Benefits of technology

It reduces scanning costs, simplifies the operation process, improves measurement efficiency and accuracy, and enables the reconstruction of 3D models without coded points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122130045A_ABST
    Figure CN122130045A_ABST
Patent Text Reader

Abstract

This application discloses a photogrammetry method, photogrammetry device, and electronic device, belonging to the field of visual measurement technology. The method includes: acquiring an image sequence of a target object, the image sequence comprising multiple image frames captured by a monocular camera during its movement relative to the target object; determining the initial three-dimensional coordinates of the target object, and using a motion prediction model to predict the camera pose information corresponding to a second image frame based on the camera pose information corresponding to the first image frame; wherein the first image frame precedes the second image frame in the image sequence, and the motion prediction model is used to predict the motion of the monocular camera; and determining the three-dimensional coordinate information of the target object based on the camera pose information corresponding to the second image frame and the initial three-dimensional coordinates. This method achieves monocular tracking photogrammetry without relying on encoded points by using predicted camera pose information and initial three-dimensional coordinates, thus reducing scanning costs and simplifying the operation process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of visual measurement technology, and in particular relates to a photogrammetry method, photogrammetry device and electronic equipment. Background Technology

[0002] Monocular scanning involves capturing a series of two-dimensional images of an object from different perspectives using a single camera. A three-dimensional model of the object or scene is then gradually reconstructed using feature point matching and Structure from Motion (SfM) techniques. Currently, monocular scanning in practical applications primarily relies on coded points to achieve high-precision point identification and localization. However, the cost of using coded points is high, and they typically need to be recycled and reused, complicating the scanning process. Summary of the Invention

[0003] This invention aims to at least solve one of the technical problems existing in the prior art. To this end, this invention proposes a photogrammetry method, a photogrammetry device, and an electronic device that can achieve monocular tracking photogrammetry without the use of coded points, thereby reducing scanning costs and simplifying the operation process.

[0004] In a first aspect, this application provides a photogrammetry method, the method comprising: Acquire an image sequence of a target object, the image sequence comprising multiple image frames acquired by a monocular camera during its movement relative to the target object; Given the initial three-dimensional coordinates of the target object, a motion prediction model is used to predict the camera pose information corresponding to the second image frame based on the camera pose information corresponding to the first image frame; wherein, in the image sequence, the first image frame is located before the second image frame, and the motion prediction model is used to predict the motion of the monocular camera; Based on the camera pose information and the initial three-dimensional coordinates corresponding to the second image frame, the three-dimensional coordinate information of the target object is determined.

[0005] According to the photogrammetry method of this application, a target object is scanned by a moving monocular camera to obtain an image sequence of the target object. The sequence is then initialized to establish the initial three-dimensional coordinates of the target object. After successful initialization, the camera motion of the current frame is predicted using the motion information of historical frames. The camera pose information of the second image frame is predicted using a motion prediction model. Based on the camera pose information of the second image frame and the initial three-dimensional coordinates, three-dimensional points are reconstructed from two-dimensional points. The coordinate point matching relationship between frames is established, realizing monocular tracking photogrammetry that does not rely on coded points. This reduces scanning costs and simplifies the operation process.

[0006] According to one embodiment of this application, the motion prediction model includes at least one of a first prediction model, a second prediction model, and a third prediction model; The first prediction model is constructed based on the motion dynamics model of the monocular camera; the second prediction model is constructed based on a temporal neural network; and the third prediction model is based on inertial measurement data collected by an inertial sensor, which is installed on the monocular camera.

[0007] According to one embodiment of this application, the first prediction model includes a basic motion module and a constraint correction module, wherein the basic motion module includes multiple camera motion models; The basic motion module is used to predict the target motion model corresponding to the monocular camera based on the camera pose information corresponding to the first image frame, wherein the target motion model is one of the plurality of camera motion models; The constraint correction module is used to verify the pose prediction results output by the basic motion module.

[0008] According to one embodiment of this application, the second prediction model includes a cascaded convolutional neural network module and a long short-term memory network module; the convolutional neural network module is used to extract spatial features of the first image frame, and the long short-term memory network module is used to predict the camera pose information corresponding to the second image frame based on the spatial features and the camera pose information corresponding to the first image frame.

[0009] According to one embodiment of this application, the third prediction module is used to determine the camera pose information corresponding to the second image frame based on the inertial measurement data corresponding to the second image frame and the camera pose information of the first image frame.

[0010] According to one embodiment of this application, determining the three-dimensional coordinate information of the target object based on the camera pose information corresponding to the second image frame and the initial three-dimensional coordinates includes: Based on the camera pose information corresponding to the second image frame, the camera imaging plane corresponding to the monocular camera when acquiring the second image frame is determined; The initial three-dimensional coordinates are projected onto the camera imaging plane and matched with the two-dimensional coordinates of the second image frame to obtain the three-dimensional coordinate information of the target object.

[0011] According to one embodiment of this application, the step of projecting the initial three-dimensional coordinates onto the camera imaging plane and matching them with the two-dimensional coordinates of the second image frame to obtain the three-dimensional coordinate information of the target object includes: The initial three-dimensional coordinates are projected onto the camera imaging plane and matched with the two-dimensional coordinates of the second image frame to obtain a coordinate matching relationship; Based on the coordinate matching relationship, bundle adjustment is used to optimize the camera pose information and the three-dimensional coordinate information.

[0012] According to one embodiment of this application, the environment in which the target object is located is further provided with a scale, and the method further includes: At least one scale image frame is acquired, wherein the scale image frame is an image frame captured by the monocular camera during the movement relative to the target object; Based on the scale image frame, the scale factor corresponding to the three-dimensional coordinate information is determined.

[0013] Secondly, this application provides a photogrammetry apparatus, which includes: The first processing module is used to acquire an image sequence of a target object, the image sequence including multiple image frames acquired by a monocular camera during its movement relative to the target object; The second processing module is used to predict the camera pose information corresponding to the second image frame based on the camera pose information corresponding to the first image frame, after determining the initial three-dimensional coordinates of the target object; wherein, in the image sequence, the first image frame is located before the second image frame, and the motion prediction model is used to predict the motion of the monocular camera. The third processing module is used to determine the three-dimensional coordinate information of the target object based on the camera pose information corresponding to the second image frame and the initial three-dimensional coordinates.

[0014] Thirdly, this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the photogrammetry method as described in the first aspect above.

[0015] Fourthly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the photogrammetry method as described in the first aspect above.

[0016] Fifthly, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the photogrammetry method as described in the first aspect.

[0017] In a sixth aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the photogrammetry method as described in the first aspect above.

[0018] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0019] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is one of the schematic flowcharts of the photogrammetry method provided in the embodiments of this application; Figure 2 This is a second schematic flowchart of the photogrammetry method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the photogrammetric device provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0021] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0022] The photogrammetry method, photogrammetry device, electronic device, and readable storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0023] The photogrammetry method provided in this application embodiment can be executed by an electronic device or a functional module or entity in an electronic device that can implement the photogrammetry method. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras, and wearable devices. The photogrammetry method provided in this application embodiment will be described below using an electronic device as the execution subject as an example.

[0024] like Figure 1 As shown, the photogrammetry method includes steps 110, 120 and 130.

[0025] Step 110: Obtain the image sequence of the target object.

[0026] The image sequence includes multiple image frames acquired by a monocular camera as it moves relative to the target object.

[0027] It is understandable that the target object is the marker object to be reconstructed in photogrammetry. The shape and size of the target object are determined according to the specific measurement requirements. For example, the target object can be a circular object.

[0028] In this embodiment, the target object is placed in a certain environment and is stationary. By controlling the monocular camera to move relative to the target object, image frames of the target object are captured from different angles to obtain an image sequence of the target object.

[0029] In practice, a monocular camera can be fixed to a mechanical device such as a robotic arm or sliding mechanism. By controlling the mechanical device, the monocular camera can be made to move along a preset trajectory.

[0030] It should be noted that the environment in which the target object is placed does not require pre-setting coding points, nor does it require setting auxiliary measurement features such as textures or patterns on the surface of the target object.

[0031] Step 120: Given the initial three-dimensional coordinates of the target object, use a motion prediction model to predict the camera pose information corresponding to the second image frame based on the camera pose information corresponding to the first image frame.

[0032] In this image sequence, the first image frame precedes the second image frame, and the motion prediction model is used to predict the motion of the monocular camera.

[0033] In this embodiment, during the process of scanning the target object with a moving monocular camera, initialization is performed, and the initial three-dimensional coordinates of the target object are established based on the image frames in the image sequence. The initial three-dimensional coordinates represent the first batch of three-dimensional coordinates obtained by reconstructing the target object.

[0034] In practice, a pair of image frames with sufficient parallax can be acquired first. By matching the two-dimensional feature points of the two image frames, the relative pose of the monocular camera between the two viewpoints can be recovered. Combined with triangulation, the three-dimensional coordinates corresponding to the matching points can be calculated to obtain the initial three-dimensional coordinates of the target object.

[0035] In this step, after successful initialization, the camera pose information corresponding to the second image frame is predicted using the camera pose information corresponding to the first image frame through a motion prediction model.

[0036] It should be noted that the first and second image frames are image frames acquired after successful initialization. That is, in the image sequence, the image frame that establishes the initial three-dimensional coordinates is located before the first and second image frames, and the first image frame is located before the second image frame. In other words, after successful initialization, the camera motion information of the historical frames is used to predict the camera motion of the current frame.

[0037] It is understandable that the second image frame can be the current frame, representing the image information captured at the moment.

[0038] In actual execution, the first image frame can be one or more frames preceding the second image frame.

[0039] For example, the image sequence of the target object includes 4 image frames. The initial three-dimensional coordinates of the target object are established through the first and second image frames. The camera pose information corresponding to the fourth image frame can be predicted using a motion prediction model based on the camera pose information corresponding to the third image frame.

[0040] For example, the image sequence of the target object includes 6 image frames. The initial three-dimensional coordinates of the target object are established through the first and second image frames. Based on the camera pose information corresponding to the third, fourth and fifth image frames, the motion prediction model can be used to predict the camera pose information corresponding to the sixth image frame.

[0041] Step 130: Determine the three-dimensional coordinate information of the target object based on the camera pose information and initial three-dimensional coordinates corresponding to the second image frame.

[0042] In this embodiment, based on the camera pose information corresponding to the second image frame predicted by the motion prediction model, the initial three-dimensional coordinates of the target object are transformed to the camera coordinate system corresponding to the second image frame, thereby determining the three-dimensional spatial position of the target object under the viewpoint corresponding to the second image frame and obtaining the three-dimensional coordinate information of the target object in the second image frame.

[0043] In some embodiments, determining the three-dimensional coordinate information of the target object based on the camera pose information and initial three-dimensional coordinates corresponding to the second image frame may include: Based on the camera pose information corresponding to the second image frame, determine the camera imaging plane corresponding to the monocular camera when acquiring the second image frame; The initial three-dimensional coordinates are projected onto the camera's imaging plane and matched with the two-dimensional coordinates of the second image frame to obtain the three-dimensional coordinate information of the target object.

[0044] In this embodiment, based on the camera pose information corresponding to the second image frame and combined with the calibration data such as the intrinsic parameters of the monocular camera, the camera imaging plane corresponding to the monocular camera when acquiring the second image frame is determined. The initial three-dimensional coordinates are projected onto the camera imaging plane corresponding to the second image frame and matched with the two-dimensional coordinates of the second image frame. The three-dimensional coordinate information describing the position and orientation of the target object in the camera coordinate system corresponding to the second image frame is calculated.

[0045] In actual execution, the initial three-dimensional coordinates are projected onto the camera imaging plane and matched with the two-dimensional coordinates of the second image frame to establish a coordinate matching relationship between the three-dimensional coordinates and the two-dimensional coordinates (i.e., 3D-2D corresponding point pairs). This coordinate matching relationship satisfies the camera projection equation. By solving the projection equation using perspective n-point algorithms such as the efficient perspective n-point method (EPnP) and direct linear transformation (DLT), the three-dimensional coordinate information describing the position and orientation of the target object in the camera coordinate system corresponding to the second image frame can be obtained.

[0046] In some embodiments, projecting the initial three-dimensional coordinates onto the camera imaging plane and matching them with the two-dimensional coordinates of the second image frame to obtain the three-dimensional coordinate information of the target object may include: The initial three-dimensional coordinates are projected onto the camera imaging plane and matched with the two-dimensional coordinates of the second image frame to obtain the coordinate matching relationship; Based on coordinate matching relationships, bundle adjustment is used to optimize camera pose information and 3D coordinate information.

[0047] It is understandable that the camera pose information corresponding to the second image frame is predicted by the motion prediction model. There is a deviation between the predicted pose and the actual pose. There is also a deviation between the camera imaging plane determined based on the camera pose information and the actual imaging plane. Optimization algorithms such as bundle adjustment can be used to minimize the reprojection error and optimize the solution of the precise pose of the monocular camera and the three-dimensional coordinates of the target object.

[0048] The objective of bundle adjustment is to minimize the reprojection error, which refers to the difference between the predicted two-dimensional coordinates obtained by projecting the initial three-dimensional coordinates onto the camera image plane according to the camera pose and the actual two-dimensional coordinates in the corresponding image frame.

[0049] In actual execution, the target object is measured by a monocular camera. After the initial three-dimensional coordinates of the target object are successfully obtained through initialization, the camera motion information of the current frame is predicted using the camera motion information of the historical frames. The camera pose information of the second image frame is predicted by the motion prediction model, and then the camera imaging plane corresponding to the monocular camera when acquiring the second image frame is determined. The initial three-dimensional coordinates are projected onto the camera imaging plane to establish a 3D-2D coordinate matching relationship. The camera pose and three-dimensional coordinate information are simultaneously optimized by bundle adjustment.

[0050] Understandably, after a monocular camera accumulates a certain number of image frames, the three-dimensional coordinate information of the target object tends to stabilize, and the camera pose of the monocular camera in the historical frames also tends to stabilize. Under such a stable state, it is possible to match the two-dimensional feature points that have not been reconstructed before, to reconstruct them in three dimensions, and to integrate multiple image frames for joint optimization, thereby improving the integrity and geometric accuracy of the three-dimensional model of the target object and obtaining more accurate photogrammetric results.

[0051] It should be noted that the initial three-dimensional coordinates and three-dimensional coordinate information can be dimensionless normalized values ​​or have actual physical units. A ruler of known length can be placed in the environment where the target object is located to restore the size information of the target object.

[0052] In some embodiments, the environment in which the target object is located is further provided with a scale, and the photogrammetry method may further include: Acquire at least one ruler image frame; Based on the scale image frame, determine the scale factor corresponding to the three-dimensional coordinate information.

[0053] Among them, the scale image frame is the image frame acquired by the monocular camera during its movement relative to the target object.

[0054] In this embodiment, during the scanning process of the monocular camera, when a scale image frame containing a scale is acquired, the scale factor corresponding to the three-dimensional coordinate information can be calculated based on the actual size of the scale in the scale image frame, and the dimensionless normalized value can be converted into size information with actual physical units (such as millimeters).

[0055] It is understandable that the scale and the target object are relatively stationary. The scale and the target object can coexist in the scale image frame. During the scanning process, the scale can be photographed at any time. When the scale image frame is acquired, the scale factor can be calculated based on the scale image frame to restore the three-dimensional coordinate information of the target object to the true physical scale.

[0056] In related technologies, monocular photogrammetry relies on coded points for positioning, which is costly and requires recycling.

[0057] According to the photogrammetry method provided in the embodiments of this application, a target object is scanned by a moving monocular camera to obtain an image sequence of the target object. The sequence is then initialized to establish the initial three-dimensional coordinates of the target object. After successful initialization, the camera motion of the current frame is predicted using the motion information of historical frames. The camera pose information of the second image frame is predicted using a motion prediction model. Based on the camera pose information of the second image frame and the initial three-dimensional coordinates, three-dimensional points are reconstructed from two-dimensional points. The coordinate point matching relationship between frames is established, realizing monocular tracking photogrammetry that does not rely on coded points. This reduces scanning costs and simplifies the operation process.

[0058] In some embodiments, the motion prediction model includes at least one of a first prediction model, a second prediction model, and a third prediction model.

[0059] In this embodiment, one or more of the first prediction model, the second prediction model, and the third prediction model can be selected to predict the pose of the monocular camera. By fusing the prediction results of multiple models, the prediction accuracy of camera pose information can be improved.

[0060] The first prediction model is constructed based on the motion dynamics model of a monocular camera.

[0061] In this embodiment, the first prediction model is constructed based on the motion dynamics model of a monocular camera, which is suitable for scenarios where the motion trajectory of the monocular camera is relatively stable. The first prediction model has high computational efficiency and strong real-time performance.

[0062] In some embodiments, the first prediction model includes a basic motion module and a constraint correction module.

[0063] In this embodiment, the basic motion module includes multiple camera motion models. The basic motion module is used to predict the target motion model corresponding to the monocular camera based on the camera pose information corresponding to the first image frame. The target motion model is one of the multiple camera motion models. The constraint correction module is used to verify the pose prediction results output by the basic motion module.

[0064] It should be noted that the camera motion model is a mathematical description of the changes in the camera's motion state based on the principles of motion dynamics.

[0065] In practice, multiple camera motion models can be pre-set in the basic motion module, such as uniform motion models and uniformly variable speed motion models (uniform acceleration or uniform deceleration). Based on the motion characteristics of the monocular camera, a target motion model is selected from multiple camera motion models. For example, the target motion model for a uniform scanning scene is a uniform motion model, and the target motion model for a uniformly variable speed scanning scene is a uniformly variable speed motion model. The motion of the monocular camera is then predicted using the target motion model.

[0066] In this embodiment, the constraint correction module has built-in constraint condition judgment logic (such as critical value) to perform a reasonableness check on the pose prediction result output by the basic motion module. When the check passes, the first prediction model outputs the predicted camera pose information; when the check fails, the first prediction model re-predicts.

[0067] The prediction process of the first prediction model will be described in detail below.

[0068] The input parameters of the first prediction model may include the camera pose information of the first image frame of the first N frames (N≥3, preferably N=5), the sampling time interval Δt between adjacent frames (calculated by the camera acquisition frame rate, such as Δt=1 / 30s when the frame rate is 30fps), and camera motion constraints (such as preset maximum rotation angle ±15° / frame, maximum translation distance ±0.5m / frame).

[0069] In this embodiment, the camera pose information of the image frame is represented by a rotation matrix R and a translation vector T, such as (R1,T1), (R2,T2), ..., (R... n ,T n (R1,T1) represents the camera pose information of the first image frame. n ,T n ) represents the camera pose information of the nth image frame.

[0070] The first prediction model adopts a dual-module structure of a basic motion module and a constraint correction module. The basic motion module pre-sets a uniform motion model and a uniformly variable motion model, and adaptively selects the motion state in the first two frames. The constraint correction module has built-in threshold judgment logic to verify the rationality of the pose prediction results output by the basic motion module.

[0071] The output parameters of the first prediction model may include camera pose information from the second image frame (R). pred ,T pred ), prediction confidence (a value between 0 and 1, calculated based on the prediction error fluctuation of the previous N frames; the smaller the fluctuation, the higher the confidence), and model adaptation flag (marking whether the current selected model is a uniform motion model or a uniformly variable motion model).

[0072] It should be noted that the first prediction model does not require a large number of samples for training. In the initialization stage, the model parameters can be calibrated using the actual pose data of the first 3 frames (which can be determined based on the measurement data of the inertial sensor). For example, the velocity parameters of the uniform motion model or the acceleration or deceleration parameters of the uniformly accelerated motion model can be determined by fitting the motion trajectory of the first 3 frames using the least squares method. This is suitable for scenarios where the motion trajectory of a monocular camera is relatively stable, and it has high computational efficiency and strong real-time performance.

[0073] In actual operation, during the scanning process of the monocular camera on the target object, the parameters of the first prediction model can be updated every 10 frames to ensure that the first prediction model matches the actual motion state of the monocular camera and guarantee prediction accuracy.

[0074] The second prediction model is based on a temporal neural network.

[0075] In this embodiment, the second prediction model is a neural network architecture based on temporal feature learning, which is suitable for scenarios with complex motion trajectories of monocular cameras (such as variable acceleration and turning), and has high prediction accuracy and strong robustness.

[0076] In some embodiments, the second prediction model includes a cascaded convolutional neural network module and a long short-term memory network module; the convolutional neural network module is used to extract spatial features of the first image frame, and the long short-term memory network module is used to predict the camera pose information corresponding to the second image frame based on the spatial features and the camera pose information corresponding to the first image frame.

[0077] In this embodiment, the second prediction model adopts a CNN-LSTM hybrid network structure, which combines the features of convolutional neural networks (CNN) and long short-term memory networks (LSTM). The convolutional neural network module extracts the spatial features of the first image frame, and the long short-term memory network module learns the temporal dependencies through gating units to predict the camera pose information corresponding to the second image frame.

[0078] The prediction process of the second prediction model will be described in detail below.

[0079] The input parameters of the second prediction model are divided into two categories. One category is image feature input, which includes the target object detection results in the first image frame of the first N frames. The detection results include the two-dimensional coordinates and scale-invariant feature transform (SIFT) feature vectors corresponding to the target object. The other category is time-series data input, which includes the camera pose information of the first image frame of the first N frames, the sampling time interval Δt between adjacent frames, and the raw inertial measurement data (if an inertial sensor is provided).

[0080] In practice, the input data can first be normalized to map parameters such as coordinates and pose to the [0,1] interval before being input into the second prediction model for prediction.

[0081] The second prediction model adopts a CNN-LSTM hybrid network structure. The convolutional neural network module has a 6-layer structure, including 3 convolutional layers (3×3 kernel size, stride 1) and 3 pooling layers (max pooling, pooling kernel 2×2), which are used to extract the spatial features of the target object in the first N frames. The long short-term memory network module has a 3-layer structure, with the input dimension being the feature vector dimension of the CNN output (128 dimensions). Temporal dependencies are learned through gating units, and the output layer is a fully connected layer with a linear activation function.

[0082] The output parameters of the second prediction model may include camera pose information from the second image frame (R). pred ,T pred ), pose prediction error range (e.g., rotation angle error ±0.5°, translation distance error ±2cm), feature matching guidance information (marking 2D point regions with high probability of matching in the current frame).

[0083] Understandably, the prediction accuracy of the second prediction model is related to model training.

[0084] In practice, a large training dataset can be used for training, such as a training dataset containing more than 100,000 sets of multi-scene data, with each set containing 20 consecutive image frames, corresponding camera pose information, and inertial measurement data.

[0085] The second prediction model can adopt a training strategy of pre-training plus fine-tuning. First, pre-training is completed on a public photogrammetry dataset, and then fine-tuning is performed using custom scene data. The loss function can be designed as a combination of pose error loss and feature matching loss. The pose error loss can adopt Euclidean distance loss, and the feature matching loss can adopt cross-entropy loss.

[0086] During model training, data augmentation techniques such as image rotation, scaling, and noise addition can be used to enrich the training data, simulate actual shooting interference, and improve the model's prediction accuracy.

[0087] The third prediction model is based on inertial measurement data collected by an inertial sensor (IMU) located on a monocular camera.

[0088] In this embodiment, the angular velocity and acceleration of the monocular camera are measured by an inertial sensor to achieve pose prediction of the monocular camera. The third prediction model can be used in combination with other models (such as the first prediction module and the second prediction model) to improve stability in dynamic scenes.

[0089] In some embodiments, the third prediction module is used to determine the camera pose information corresponding to the second image frame based on the inertial measurement data corresponding to the second image frame and the camera pose information of the first image frame.

[0090] The prediction process of the third prediction model will be described in detail below.

[0091] The input parameters of the third prediction model may include inertial measurement data acquired in real time by the IMU (such as three-axis angular velocity and three-axis acceleration), camera pose information of the previous frame, extrinsic parameter matrix of the monocular camera, and zero-bias calibration parameters.

[0092] In this embodiment, the third prediction model can adopt an IMU integral plus Kalman filter structure. The rotation increment is calculated by integrating angular acceleration and the translation increment is calculated by integrating linear acceleration. The Kalman filter is divided into a prediction step (predicting pose based on IMU data) and an update step (correcting the prediction result with the pose data of the previous frame).

[0093] The output parameters of the third prediction model may include the camera pose information of the current frame, the IMU data reliability score (0-10 points, when the score is lower than 3 points, it triggers fusion with other models), and the IMU data abnormality alarm signal (output when the acceleration or angular velocity data exceeds the preset threshold).

[0094] It should be noted that the third prediction model does not require complex sample training. The core lies in the IMU calibration process, which can be achieved by combining static and dynamic calibration. During static calibration, the IMU is left stationary for 10 minutes to collect zero-bias data and establish a zero-bias model. During dynamic calibration, data is collected by moving the monocular camera at a constant speed to correct the extrinsic parameter matrices of the IMU and the camera, ensuring the consistency of the data coordinate system.

[0095] The following is a specific example.

[0096] like Figure 2 As shown, a target object and a scale with known true length values ​​are placed in the scene to be measured. The scale has special markings that can distinguish it from the target object, in order to recover the scale information of monocular photogrammetry.

[0097] Image sequences are acquired using monocular scanning. Each image frame contains 2D points. Initialization is performed based on the image frames to establish initial 3D coordinates, i.e., initial 3D points. If the number of reconstructed 3D points is greater than 0, it indicates that the initialization was successful.

[0098] After successful initialization, the camera motion of the current frame is predicted using the motion information of historical frames. The initial 3D points are projected onto the predicted camera imaging plane and matched with the actual 2D points of the current frame. If the matching fails, the current frame is discarded.

[0099] If a match is successful, the 3D point coordinates are optimized using bundle adjustment to accurately calculate the camera pose. After collecting a certain number of frames, the 3D point coordinates tend to stabilize, and the camera's historical frame poses also tend to stabilize. In this state, other unreconstructed 2D points are matched, 3D reconstruction is performed, and multi-frame optimization is carried out. The maximum number of frames is reached or the scan is completed.

[0100] The above process continues until the preset maximum number of frames is reached or the entire scanning task is completed.

[0101] In this embodiment, the initial three-dimensional coordinates of the target object are established through initialization. After successful initialization, the camera motion of the current frame is predicted by the motion information of the historical frames, and the camera pose information is predicted by the motion prediction model. Based on the predicted camera pose information and the initial three-dimensional coordinates, three-dimensional points are reconstructed from two-dimensional points, and the coordinate point matching relationship between frames is established. The three-dimensional coordinates are optimized by bundle adjustment, and the camera pose is accurately calculated. This achieves monocular tracking photogrammetry without relying on coded points, which reduces scanning costs and simplifies the operation process. At the same time, when a scale image frame is acquired, the scale factor can be calculated based on the scale image frame to restore the three-dimensional coordinate information of the target object to the true physical scale.

[0102] The photogrammetric method provided in this application can be executed by a photogrammetric device. This application uses a photogrammetric device executing the photogrammetric method as an example to illustrate the photogrammetric device provided in this application.

[0103] This application also provides a photogrammetric device.

[0104] like Figure 3 As shown, the photogrammetric device includes: The first processing module 310 is used to acquire an image sequence of the target object, the image sequence including multiple image frames acquired by the monocular camera during the movement relative to the target object. The second processing module 320 is used to predict the camera pose information corresponding to the second image frame based on the camera pose information corresponding to the first image frame, after determining the initial three-dimensional coordinates of the target object; wherein, in the image sequence, the first image frame is located before the second image frame, and the motion prediction model is used to predict the motion of the monocular camera. The third processing module 330 is used to determine the three-dimensional coordinate information of the target object based on the camera pose information and initial three-dimensional coordinates corresponding to the second image frame.

[0105] According to the photogrammetry apparatus provided in the embodiments of this application, a moving monocular camera scans a target object to obtain an image sequence of the target object, initializes it, establishes the initial three-dimensional coordinates of the target object, and after successful initialization, predicts the camera motion of the current frame through the motion information of historical frames, predicts the camera pose information of the second image frame through the motion prediction model, and reconstructs three-dimensional points from two-dimensional points based on the camera pose information of the second image frame and the initial three-dimensional coordinates, establishes the coordinate point matching relationship between frames, and realizes monocular tracking photogrammetry without relying on coded points, which reduces scanning costs and simplifies the operation process.

[0106] In some embodiments, the motion prediction model includes at least one of a first prediction model, a second prediction model, and a third prediction model; The first prediction model is constructed based on the motion dynamics model of the monocular camera; the second prediction model is constructed based on the temporal neural network; and the third prediction model is based on the inertial measurement data collected by the inertial sensor, which is set on the monocular camera.

[0107] In some embodiments, the first prediction model includes a basic motion module and a constraint correction module, wherein the basic motion module includes multiple camera motion models; The basic motion module is used to predict the target motion model corresponding to the monocular camera based on the camera pose information corresponding to the first image frame. The target motion model is one of multiple camera motion models. The constraint correction module is used to verify the pose prediction results output by the basic motion module.

[0108] In some embodiments, the second prediction model includes a cascaded convolutional neural network module and a long short-term memory network module; the convolutional neural network module is used to extract spatial features of the first image frame, and the long short-term memory network module is used to predict the camera pose information corresponding to the second image frame based on the spatial features and the camera pose information corresponding to the first image frame.

[0109] In some embodiments, the third prediction module is used to determine the camera pose information corresponding to the second image frame based on the inertial measurement data corresponding to the second image frame and the camera pose information of the first image frame.

[0110] In some embodiments, the third processing module 330 is configured to determine the three-dimensional coordinate information of the target object based on the camera pose information and initial three-dimensional coordinates corresponding to the second image frame, including: Based on the camera pose information corresponding to the second image frame, determine the camera imaging plane corresponding to the monocular camera when acquiring the second image frame; The initial three-dimensional coordinates are projected onto the camera's imaging plane and matched with the two-dimensional coordinates of the second image frame to obtain the three-dimensional coordinate information of the target object.

[0111] In some embodiments, the third processing module 330 is configured to project the initial three-dimensional coordinates onto the camera imaging plane and match them with the two-dimensional coordinates of the second image frame to obtain the three-dimensional coordinate information of the target object, including: The initial three-dimensional coordinates are projected onto the camera imaging plane and matched with the two-dimensional coordinates of the second image frame to obtain the coordinate matching relationship; Based on coordinate matching relationships, bundle adjustment is used to optimize camera pose information and 3D coordinate information.

[0112] In some embodiments, the environment in which the target object is located is further provided with a scale, and the third processing module 330 is further used for: Acquire at least one scale image frame, which is an image frame captured by a monocular camera during movement relative to the target object; Based on the scale image frame, determine the scale factor corresponding to the three-dimensional coordinate information.

[0113] The photogrammetric device in the embodiments of this application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip.

[0114] The photogrammetric apparatus provided in this application embodiment can realize all the processes implemented in the above-described photogrammetric method embodiment, and will not be repeated here to avoid repetition.

[0115] In some embodiments, such as Figure 4 As shown, this application embodiment also provides an electronic device 400, including a processor 401, a memory 402, and a computer program stored in the memory 402 and executable on the processor 401. When the program is executed by the processor 401, it implements the various processes of the above-described photogrammetry method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0116] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0117] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described photogrammetry method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0118] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0119] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described photogrammetry method.

[0120] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0121] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described photogrammetry method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0122] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0123] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0124] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0125] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0126] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0127] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. A photogrammetric method, characterized in that, include: Acquire an image sequence of a target object, the image sequence comprising multiple image frames acquired by a monocular camera during its movement relative to the target object; Given the initial three-dimensional coordinates of the target object, a motion prediction model is used to predict the camera pose information corresponding to the second image frame based on the camera pose information corresponding to the first image frame; wherein, in the image sequence, the first image frame is located before the second image frame, and the motion prediction model is used to predict the motion of the monocular camera; Based on the camera pose information and the initial three-dimensional coordinates corresponding to the second image frame, the three-dimensional coordinate information of the target object is determined.

2. The photogrammetric method according to claim 1, characterized in that, The motion prediction model includes at least one of a first prediction model, a second prediction model, and a third prediction model; The first prediction model is constructed based on the motion dynamics model of the monocular camera; the second prediction model is constructed based on a temporal neural network; and the third prediction model is based on inertial measurement data collected by an inertial sensor, which is installed on the monocular camera.

3. The photogrammetry method according to claim 2, characterized in that, The first prediction model includes a basic motion module and a constraint correction module, wherein the basic motion module includes multiple camera motion models; The basic motion module is used to predict the target motion model corresponding to the monocular camera based on the camera pose information corresponding to the first image frame, wherein the target motion model is one of the plurality of camera motion models; The constraint correction module is used to verify the pose prediction results output by the basic motion module.

4. The photogrammetric method according to claim 2, characterized in that, The second prediction model includes a cascaded convolutional neural network module and a long short-term memory network module; the convolutional neural network module is used to extract the spatial features of the first image frame, and the long short-term memory network module is used to predict the camera pose information corresponding to the second image frame based on the spatial features and the camera pose information corresponding to the first image frame.

5. The photogrammetry method according to claim 2, characterized in that, The third prediction module is used to determine the camera pose information corresponding to the second image frame based on the inertial measurement data corresponding to the second image frame and the camera pose information of the first image frame.

6. The photogrammetric method according to any one of claims 1-5, characterized in that, Determining the three-dimensional coordinate information of the target object based on the camera pose information corresponding to the second image frame and the initial three-dimensional coordinates includes: Based on the camera pose information corresponding to the second image frame, the camera imaging plane corresponding to the monocular camera when acquiring the second image frame is determined; The initial three-dimensional coordinates are projected onto the camera imaging plane and matched with the two-dimensional coordinates of the second image frame to obtain the three-dimensional coordinate information of the target object.

7. The photogrammetric method according to claim 6, characterized in that, The step of projecting the initial three-dimensional coordinates onto the camera imaging plane and matching them with the two-dimensional coordinates of the second image frame to obtain the three-dimensional coordinate information of the target object includes: The initial three-dimensional coordinates are projected onto the camera imaging plane and matched with the two-dimensional coordinates of the second image frame to obtain a coordinate matching relationship; Based on the coordinate matching relationship, bundle adjustment is used to optimize the camera pose information and the three-dimensional coordinate information.

8. The photogrammetric method according to any one of claims 1-5, characterized in that, The environment in which the target object is located is also equipped with a scale, and the method further includes: At least one scale image frame is acquired, wherein the scale image frame is an image frame captured by the monocular camera during the movement relative to the target object; Based on the scale image frame, the scale factor corresponding to the three-dimensional coordinate information is determined.

9. A photogrammetric apparatus, characterized in that, include: The first processing module is used to acquire an image sequence of a target object, the image sequence including multiple image frames acquired by a monocular camera during its movement relative to the target object; The second processing module is used to predict the camera pose information corresponding to the second image frame based on the camera pose information corresponding to the first image frame, after determining the initial three-dimensional coordinates of the target object; wherein, in the image sequence, the first image frame is located before the second image frame, and the motion prediction model is used to predict the motion of the monocular camera. The third processing module is used to determine the three-dimensional coordinate information of the target object based on the camera pose information corresponding to the second image frame and the initial three-dimensional coordinates.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the photogrammetry method as described in any one of claims 1-8.