Method and apparatus for dynamically mapping and generating digital twin scenarios based on standard band calibration

By setting standard belt correction image perspective at the construction site and using Procrustes analysis and bias correction algorithm, the problem of insufficient coordinate calculation accuracy and real-time performance in digital twin scenes on the construction site is solved, and accurate and accurate mapping of dynamic objects is achieved.

CN119918154BActive Publication Date: 2025-07-22XIAMEN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510405616.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-22
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

In the digital twin scenarios at the construction site, the accuracy and real-time performance of coordinate calculations are insufficient, making it difficult to adapt to changes in dynamic objects in complex environments. Especially under the influence of factors such as camera distortion and lighting changes, the detection and mapping of target objects are inaccurate.

Method used

By setting a standard image viewing angle, combining adjacent point deviation correction algorithms analyzed by deep learning and Procrustes, the camera distortion is corrected, perspective projection is used to calculate the three-dimensional coordinates of the target object, and the digital twin scene in the BIM model is updated in real time.

Benefits of technology

It improves the accuracy and stability of target object detection, meets the real-time mapping requirements of dynamic scenarios, and enhances scene adaptability and coordinate calculation accuracy in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918154B_ABST
    Figure CN119918154B_ABST
Patent Text Reader

Abstract

The method and device for dynamically mapping and generating a digital twin scenario based on standard belt correction provided by the present invention relate to the technical fields of data processing and digital twin scenario construction. The present invention collects a dynamic image data set at a construction site and processes it to box and classify each target object in the image; sets a standard belt for the image and adjusts it to correct the image perspective, frames all target objects in the image within the standard belt for target recognition; converts the two-dimensional pixel coordinates of the recognized target objects into three-dimensional coordinates in the actual scenario through perspective projection calculation; then corrects the calculated three-dimensional coordinates based on the adjacent point deviation correction algorithm of Procrustes analysis; inputs the corrected three-dimensional coordinates into a pre-established BIM model to realize real-time mapping and update of dynamic objects in the digital twin scenario. The present invention can improve the coordinate calculation accuracy during the dynamic mapping of the digital twin scenario and meet the requirements of real-time performance and complexity of the dynamic scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of data processing and digital twin scenario construction. Specifically, it relates to a method and device for dynamically mapping and generating a digital twin scenario based on standard band correction. Background Art

[0002] Digital twin technology is a technology that constructs a dynamic mapping model of a physical entity (such as a building, equipment, construction site, etc.) in a virtual space through digital means, and is widely used in fields such as architecture, manufacturing, and transportation.

[0003] The construction site in the field of architecture is a complex and dynamically changing environment, involving a large number of dynamic objects, such as building materials, construction machinery, construction workers, etc. In order to achieve digital management of the construction site, it is necessary to collect the position information of these dynamic objects in real time and map it to the digital twin model. In the prior art, surveillance cameras are usually used as the main data collection devices, and the two-dimensional pixel coordinates in the image are converted into three-dimensional coordinates in the actual scene through the perspective projection principle. However, this process involves complex mathematical calculations, has poor real-time performance in dynamic scenarios, and is easily affected by factors such as camera distortion, viewing angle limitations, lighting changes, and multi-target interference, resulting in insufficient accuracy and real-time performance of coordinate calculation.

[0004] In view of this, the applicant has specifically proposed this application after studying the existing technologies. Summary of the Invention

[0005] The present invention aims to provide a method and device for dynamically mapping and generating a digital twin scenario based on standard band correction to solve the disadvantages such as insufficient accuracy and real-time performance of coordinate calculation in the existing methods.

[0006] To solve the above technical problems, the present invention is realized through the following technical solutions:

[0007] A method for dynamically mapping and generating a digital twin scenario based on standard band correction includes:

[0008] S1, collecting a dynamic image dataset of the construction site and processing each image in the image dataset to box and classify each target object in the image;

[0009] S2, setting and adjusting the standard band of the image to correct the image viewing angle and frame all target objects in the image within the standard band for target recognition;

[0010] S3, converting the two-dimensional pixel coordinates of the recognized target objects into three-dimensional coordinates in the actual scene through perspective projection calculation;

[0011] S4. Use the adjacent point correction algorithm based on Procrustes analysis to correct the calculated three-dimensional coordinates.

[0012] S5. Input the corrected three-dimensional coordinates into the pre-built BIM model to achieve real-time mapping and updating of dynamic objects in the digital twin scenario.

[0013] Preferably, the target objects are safety factors existing at the construction site, including personnel, stacked materials, construction tools, construction items, ladders with a risk of falling, and construction machinery.

[0014] Preferably, the operations for processing each image include:

[0015] Filter and screen the image dataset to remove images that do not contain target objects or have incomplete target object captures in the scene, images with blurred target objects due to environmental impacts, and duplicate images.

[0016] Use deep learning technology to frame and label the category of each target object in the image data.

[0017] Preferably, the standard band is: based on the camera imaging principle, a position area within a specific range from the camera and with a weak impact on image distortion; and it is possible to move this area by simulating and adjusting the pitch angle of the camera with respect to the horizontal direction to dynamically capture and frame the target objects in the image, ensuring that the coordinates of the target objects at all points can be accurately calculated. Specifically:

[0018] Let the height of the camera be h, that is, the two-dimensional coordinate is (0, h); the reference plane is the plane where the target object is located, that is, the X-axis; the angle between the camera and the projection range of the standard band is γ. Calculate the angle γ according to the height h of the camera and the distance to the standard band. The formula is:

[0019] ;

[0020] Where, 、 are respectively the lower limit value and the upper limit value of the standard band on the X-axis;

[0021] If the optical axis of the camera is used as the basis for moving the standard band, when moving the standard band, according to the fixed angle γ of each camera, calculate the actual distance of each target object on the reference plane, so as to obtain the two-dimensional position of the target object. The expression is:

[0022] ;

[0023] Where, $x_i$ is the position of the target point $i$ on the $X$-axis, that is, the distance between the camera and the target projected on the reference plane; $i$ is the serial number of the target point; $n$ is the number of times the optical axis point moves; $\theta$ is the pitch angle between the camera and the horizontal direction.

[0024] Preferably, the said S3 is specifically:

[0025] The angle formed by the two edges of the maximum range between the camera as the vertex and the measured target through the lens is called the field of view angle FOV; if the target exceeds the range of the field of view angle FOV, it will not be captured by the camera. The included angle of the field of view angle is , that is, the diagonal field of view angle;

[0026] According to the imaging object diameter of the diagonal line of the rectangular photosensitive surface, the diagonal field of view angle is used to calculate the object distance $s$. Since the field of view angle of the camera is fixed and known, and the range that can be photographed is also related to the focal length of the lens, the formula for the object distance $s$ that the camera can photograph is:

[0027] ;

[0028] where is the imaging object diameter of the visible range of the lens, $c$ is the point of the target object image and also the intersection point of the optical axis and the reference plane;

[0029] According to the fact that the two-dimensional image captured by the camera is the projection of the three-dimensional space on the plane, based on the camera pose and the coordinates of the target object in the image, that is, the target box, the position of the target object in the real world is obtained;

[0030] Let the three-dimensional coordinates of the camera be the point $H(0, h, 0)$, and the pitch angle between $H$ and the horizontal direction is $\theta$; the reference plane is the plane where the target object is located and is set on the same plane as the center of the target object. The straight line $s$ is the object distance that the camera can photograph, that is, the optical axis. Let the point $c$ ( , 0, 0) be the point of the target object and also the intersection point of the optical axis and the reference plane. The coordinate of point $c$ and the trigonometric function relationship of angle $\theta$ are:

[0031] ;

[0032] ;

[0033] According to the perspective projection principle, the image captured by the camera is perpendicular to the optical axis and its projection on the reference plane is a trapezoid. Therefore, the midpoint of the target object image is set as point $c$; then, for any two-dimensional coordinate point $Q$ ( , ) in the target object image, its value is obtained through measurement transformation, and the corresponding three-dimensional coordinates of $Q$ in the three-dimensional coordinate system are , and the formula is:

[0034] ;

[0035] Thus, the equation of the straight line l passing through the point and H is obtained:

[0036] ;

[0037] Then, the coordinates of the intersection point P of the straight line l and the reference plane are ( , , ), and the three-dimensional coordinate formula of point P is:

[0038] ;

[0039] Thus, the two-dimensional coordinate point Q is mapped to the world coordinate point P corresponding in the digital twin scenario;

[0040] Convert the two-dimensional coordinates of each target in the image into three-dimensional coordinates; then perform perspective projection transformation of the three-dimensional coordinates by the camera, so as to obtain the world coordinates corresponding to all the target objects mapped to the digital twin scenario.

[0041] Preferably, the adjacent point position rectification algorithm based on Procrustes analysis performs rectification by finding the optimal rotation matrix through singular value decomposition SVD. Specifically:

[0042] Take the calculated three-dimensional coordinates of the target object as the set of points to be rectified, and select a point to be rectified from the set of points to be rectified, and select the three-dimensional coordinates of the actual point position of a corresponding target object as the reference point;

[0043] Calculate and 's centroid, and translate the centroid of to the centroid of to make the centroids of the two coincide and eliminate the translation difference. The formula is:

[0044] ;

[0045] ;

[0046] where N is the total number of point positions on the target object; is the centroid of the calculated point position ; is the centroid of the actual point position ;

[0047] Calculate the covariance matrix of the centered reference point and the point to be rectified. The formula is:

[0048] ;

[0049] wherein, is the centroid the transpose matrix of the coordinate point matrix calculated after centering; is the actual point matrix after centering; is the covariance matrix; T represents the transpose matrix;

[0050] Perform SVD decomposition on the covariance matrix The formula is:

[0051] ;

[0052] wherein, U and V are orthogonal matrices of SVD decomposition respectively; Σ is a diagonal matrix, that is, the singular value matrix;

[0053] Calculate the rotation matrix R through U and V, and calculate the translation vector t in combination with the centroid of the two points. The formula is:

[0054] ;

[0055] ;

[0056] wherein, R is the rotation matrix; t is the translation vector;

[0057] Apply the rotation matrix R and the translation vector t to each target object point position in the point set to be corrected to obtain the new coordinates after correction and alignment. The formula is:

[0058] ;

[0059] wherein, is the point after correction and alignment.

[0060] Preferably, when inputting the corrected three-dimensional coordinates into the pre-built BIM model, import the corrected three-dimensional coordinates into an Excel file, and use the Dynamo plug-in to import the Excel file into the BIM model to complete the mapping of the digital twin scenario.

[0061] Preferably, it further includes: adjusting the target object coordinates in the BIM model through the grid drawing auxiliary calibration method to achieve the precise mapping of the actual scene and the digital scene; specifically:

[0062] In the BIM model, construct a three-dimensional coordinate system consistent with the physical coordinate system of the actual scene;

[0063] Each point in the scene is calibrated through a square grid, and the nodes of the grid should correspond one by one to the reference points in the actual scene;

[0064] Each grid point represents a reference coordinate in the actual scene in the BIM model;

[0065] According to the accuracy requirements, the grid is subdivided into sizes;

[0066] Combined with the data of the target objects in the actual scene, the target objects within the grid are adjusted.

[0067] The present invention also provides a digital twin scene dynamic mapping generation device based on standard band calibration, including:

[0068] An image acquisition unit, which is used to acquire a dynamic image data set of the construction site and process each image in the image data set to frame and classify each target object in the image;

[0069] A target recognition unit, which is used to set and adjust the standard band of the image to correct the image perspective, and frame all the target objects in the image within the standard band for target recognition; wherein, the standard band is: based on the camera imaging principle, a position area within a specific range from the camera and with a weak influence on the image distortion; and it can move this area by simulating and adjusting the pitch angle between the camera and the horizontal direction to dynamically capture and frame the target objects in the image, ensuring that the coordinates of all target objects at all points can be accurately calculated. Specifically: Let the height of the camera be h, that is, the two-dimensional coordinate is (0, h); the reference plane is the plane where the target object is located, that is, the X-axis; the included angle between the camera and the projection range of the standard band is γ, and the included angle γ is calculated according to the height h of the camera and the distance between the camera and the standard band. The formula is:

[0070] ;

[0071] Among them, 、 are respectively the lower limit value and the upper limit value of the standard band on the X-axis;

[0072] If the optical axis of the camera is used as the basis for moving the standard band, when moving the standard band, according to the fact that the included angle γ of each camera remains unchanged, the actual distance of each target object on the reference plane is calculated, so as to obtain the two-dimensional position of the target object. The expression is:

[0073] ;

[0074] Among them, $x_i$ is the position of the target point $i$ on the $X$-axis, that is, the distance between the camera and the target projected on the reference plane; $i$ is the serial number of the target point; $n$ is the number of times the optical axis point moves; $\theta$ is the pitch angle between the camera and the horizontal direction;

[0075] A three-dimensional conversion unit for converting the two-dimensional pixel coordinates of the recognized target into three-dimensional coordinates in the actual scene through perspective projection calculation;

[0076] A three-dimensional coordinate correction unit for correcting the calculated three-dimensional coordinates based on the adjacent point correction algorithm of Procrustes analysis;

[0077] A real-time mapping unit for inputting the corrected three-dimensional coordinates into a pre-built BIM model to realize real-time mapping update of dynamic objects in the digital twin scene.

[0078] The present invention also provides a digital twin scene dynamic mapping generation device based on standard band correction, including a processor and a memory. The memory stores a computer program, and the computer program can be executed by the processor to implement a digital twin scene dynamic mapping generation method based on standard band correction as described above.

[0079] The present invention also provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, a digital twin scene dynamic mapping generation method based on standard band correction as described above is realized.

[0080] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0081] (1) The present invention enhances the scene adaptability in complex environments. Through sensors (such as cameras) and deep learning technologies, the present invention improves the robustness of target detection and coordinate calculation in complex environments, ensuring the stability and accuracy of the technology under complex conditions such as target occlusion and light changes.

[0082] (2) The present invention solves the problem of insufficient coordinate calculation accuracy caused by camera distortion. By introducing the "standard band" with the least influence of camera distortion obtained from relevant experimental tests, the present invention weakens and effectively avoids the influence of camera distortion on target coordinate calculation, and improves the coordinate calculation accuracy of the target in the actual scene. Especially in the calculation of targets within the "standard band" distance range, the accuracy of coordinate calculation can be ensured, and by combining the change of the calculated angle and the projection principle, the actual position of the target construction can be deduced, and the function of dynamic tracking of the image device can be feedback.

[0083] (3) The present invention realizes the real-time mapping and updating of dynamic objects, meeting the real-time requirements of dynamic scenes. The present invention deeply integrates the BIM model with dynamic data collection through the Dynamo plug-in in the BIM software, and updates the dynamic object information in the digital twin model in real time, ensuring that the digital twin model can accurately reflect the dynamic changes of the construction site. The coordinate position is corrected by the adjacent point correction algorithm based on Procrustes analysis, and then the corrected coordinate data is input into the script instructions written by Dynamo through an Excel file, so as to realize the automatic adjustment of the target object in the BIM model, improve the accuracy of coordinate calculation, and meet the real-time and accuracy requirements of dynamic scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0085] Figure 1 A schematic diagram of a method for generating dynamic mapping of a digital twin scene based on standard band correction provided in Example 1.

[0086] Figure 2 A flowchart of a method for generating dynamic mapping of a digital twin scene based on standard band correction is provided in Example 1.

[0087] Figure 3 A schematic diagram of a method for generating dynamic mapping of a digital twin scene based on standard band correction provided in Example 1.

[0088] Figure 4 This is a schematic diagram of the standard tape provided in Example 1.

[0089] Figure 5 This is a schematic diagram of the standard band sight range provided in Example 1.

[0090] Figure 6 This is a schematic diagram of the field of view (FOV) model provided in Example 1.

[0091] Figure 7 This is a schematic diagram of the diagonal field of view model provided in Example 1.

[0092] Figure 8 A perspective projection principle model diagram provided in Example 1.

[0093] Figure 9 A schematic diagram of a digital twin scene dynamic mapping generation device based on standard band correction provided in Example 2.

[0094] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. Specific Embodiments

[0095] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0096] Embodiment 1

[0097] Embodiment 1 of the present invention provides a method for dynamically mapping and generating a digital twin scenario based on standard band calibration, which can be implemented by a device for dynamically mapping and generating a digital twin scenario based on standard band calibration (hereinafter referred to as the dynamic mapping generation device), and in particular, is executed by one or more processors in the dynamic mapping generation device.

[0098] In this embodiment, the dynamic mapping generation device may be an electronic device equipped with a processor, and the processor has a computer program for the method for dynamically mapping and generating a digital twin scenario based on standard band calibration and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited herein.

[0099] In this embodiment, the core of digital twin technology is to collect data of the physical world in real time through devices such as sensors and cameras, and construct a dynamic mapping model in the virtual space to achieve real-time interaction and synchronization between the physical world and the digital world. In the construction field, BIM technology provides rich static building information for digital twins, including geometric information, attribute information, and management information of buildings. However, in dynamic environments such as construction sites, how to generate dynamic object mappings in real time and accurately still faces many challenges.

[0100] The following are the main defects of the prior art in the generation of dynamic object mappings in digital twin scenarios:

[0101] (1) Poor adaptability to scenes in complex environments. Existing technologies have poor adaptability to complex construction sites and are difficult to handle problems such as target occlusion, illumination changes, and multi-target interference. For example, at a construction site, the target may be obscured by other objects, or the image quality may degrade due to illumination changes, affecting the accuracy of target detection and coordinate calculation. In these situations, the on-site surveillance cameras cannot accurately identify and filter out valid images, so it is necessary to filter the on-site surveillance images saved by the camera to improve the quality of the image data.

[0102] (2) Camera distortion leads to large coordinate errors in the calculated objects. In the prior art, the coordinate calculation based on the perspective projection of the camera does not fully consider the influence of camera distortion. Camera distortion is a geometric deviation caused by the optical characteristics of the lens, including radial distortion and tangential distortion. Radial distortion includes barrel distortion or pincushion distortion. This distortion will cause the image to bend and deform, resulting in the constructed position of the target object in the image being inconsistent with the position in the actual scene. In addition, distortion correction involves many factors, and it cannot fully guarantee that the image can meet the requirements of coordinate calculation after distortion correction.

[0103] (3) Poor real-time performance in dynamic scenes. In dynamic environments such as construction sites, the position and state of the target objects are constantly changing, and the digital twin model needs to be updated in real time. Existing technologies are usually static modeling based on BIM and lack real-time mapping capabilities. For example, at a construction site, the position and state of target objects such as building materials, machinery and equipment, and construction personnel are constantly changing, and existing technologies make it difficult to update the information of these dynamic objects in real time. In addition, the efficiency of dynamic data collection and processing is insufficient, resulting in poor real-time performance of dynamic object mapping. However, using the Dynamo plug-in that comes with the BIM software, combined with perspective projection to calculate the coordinates of dynamic objects, can meet the real-time performance of dynamic scenes to a certain extent.

[0104] like Figures 1-3 As shown, a method for generating dynamic mapping of a digital twin scene based on standard band correction includes steps S1 to S5.

[0105] S1, collects dynamic image datasets of the construction site and processes each image in the image dataset to select and label each target object in the image.

[0106] In this embodiment, according to the actual situation of the construction site, a surveillance camera or a drone is used to capture images of the site to form an image data set. The camera automatically saves the on-site surveillance images every second and automatically uploads them to the cloud disk for storage.

[0107] The target objects in this embodiment are safety factors existing at the construction site, including personnel, piled materials, construction tools, construction objects, ladders and construction machinery with a risk of falling. Therefore, the on-site images captured by the camera must contain one or more of the above factors to form an image data set.

[0108] Then, each image is processed, including: filtering the image data set to remove images that do not contain the target object in the scene or the target object is not fully captured, images in which the target object is blurred due to environmental influences, and repeated images;

[0109] Preprocess the collected images, including denoising, contrast enhancement, brightness adjustment, etc., to improve image quality and facilitate subsequent target recognition and category labeling;

[0110] Label the objects in the filtered image according to the labeling scheme. This embodiment uses deep learning technology to select and label each target object in the image data. The labeling information should be accurate and detailed so that the subsequent target recognition model can accurately identify the target object. For example, the labeling scheme of the PASCAL VOC dataset is used to select and label each target object in the image data, and the label information is the category of the target object.

[0111] In this step, through target recognition and category labeling, the key information in the image can be extracted in a structured form, allowing the computer to "understand" the image content; it is also convenient for subsequent import into the BIM model, so that the entity objects in the physical world can be accurately mapped to the digital world, ensuring the authenticity and accuracy of the digital twin scene.

[0112] In the process of target recognition, an existing deep learning target recognition model, such as the YOLO series target detection model, can be used for target recognition and labeling of the present invention after training. The identified target object needs to be framed, that is, a bounding box is drawn around it to clarify its position and range. At the same time, each target object also needs to be labeled with its category (such as workers, machinery, building materials, etc.) for subsequent classification and analysis.

[0113] S2, set the standard band of the image and adjust it to correct the image viewing angle, and frame all targets in the image within the standard band for target recognition; wherein, the standard band is: based on the camera imaging principle, a position area that is within a specific range from the camera and has a weak effect on image distortion; and the area can be moved by simulating the adjustment of the pitch angle of the camera with respect to the horizontal direction to dynamically capture and frame the targets in the image, ensuring that the coordinates of the targets at all points can be accurately calculated.

[0114] This step aims to improve the accuracy and efficiency of target recognition. By calibrating the perspective with a standard band and framing the range of the target object, unnecessary computational effort can be reduced.

[0115] Due to certain distortions in lens shooting, the coordinate data calculated by the perspective projection algorithm will deviate. Lens distortion is caused by the manufacturing precision of the lens and the deviation of the assembly process, and it can usually be described by a mathematical model. Common distortion models include radial distortion and tangential distortion, and radial distortion includes barrel distortion or pincushion distortion.

[0116] After distortion correction using common distortion models, the image has improved, but it still does not meet the expected assumptions. This may be because the distortion correction formula implicitly changes the position ratio of pixel points, and the relationship between this change and the focal length is not linear. For example, the scaling ratio introduced during the correction process is related to the focal length, resulting in different degrees of coordinate magnification or reduction; the corrected coordinates are calculated based on the ideal projection of the camera's internal parameters and may not be aligned with the actual real physical measurement coordinates. Image distortion correction requires interpolation operations (such as bilinear interpolation or resampling), and numerical errors may be introduced during the interpolation process, leading to complex reasons such as inconsistent coordinates with the actual situation.

[0117] Therefore, using general distortion correction methods cannot meet the actual requirements. The embodiment of the present invention proposes the idea of a "standard band". The purpose of setting the standard band is to correct the image perspective and ensure that the target object in the image can be recognized and analyzed under a relatively unified and accurate perspective. By adjusting the standard band, the image distortion caused by the change of the camera perspective can be corrected, enabling the target object in the image to be more accurately framed and recognized.

[0118] The standard band is as follows: Imitating the camera imaging principle, the position area within a specific range from the camera (this specific range can be obtained through experimental tests) is defined as the standard band. Within this area, the influence of the camera on image distortion is relatively small, and the area can be moved by simulating the adjustment of the pitch angle θ between the camera and the horizontal direction to dynamically capture and frame the target object in the image, ensuring that the target object coordinates at all points can be accurately calculated.

[0119] Without considering distortion correction, to determine the influence range of lens distortion, the camera model used in this embodiment is the Dahua camera DH-IPC-HFW1230M-A-I1 with a focal length of 6mm. The captured image shows that when the target object is within a distance of 300 - 420 cm from the camera, the distortion effect of the camera on the image is relatively small, and the calculated coordinates of the target object are also relatively accurate. Also, due to the limitation of the camera's shooting angle, there will be a situation where the target object at the boundary of the captured image is incomplete. And due to the imaging principle of the camera, the calculated X-axis coordinate of the target object shows an opposite trend to the actual scene coordinate, while the Z-axis is not affected. These problems are all caused by the inherent parameters inside the camera and are difficult to be corrected by a third party such as algorithms or manual adjustment.

[0120] As Figure 4 shown, for this reason, the position area between 300 - 420 cm from the camera is defined as the "standard zone" of this camera. As Figure 5 shown, by adjusting the pitch angle θ between the camera and the horizontal direction, the "standard zone" can be moved up or down in the scene, so that the farther or nearer target objects and the incompletely captured target objects can be completely framed within the "standard zone" for dynamic capture, thereby ensuring that the coordinates of the target objects at all points in the scene can be accurately calculated.

[0121] By dynamically adjusting the standard zone and the pitch angle θ of the analog camera, dynamic capture and framing of the target object in the image can be achieved. When adjusting, the real-time image feedback needs to be combined to gradually optimize the position of the standard zone. This ensures that no matter where the target object is in the image, its coordinates can be accurately calculated, ensuring the accuracy and stability of the recognition result, and providing a reliable data basis for subsequent target tracking, behavior analysis, etc.

[0122] Let the height of camera H be h, that is, the two-dimensional coordinate is (0, h); the reference plane is the plane where the target object is located, that is, the X-axis; the included angle between the camera and the projection range of the standard zone is γ. According to the height h of the camera and the distance to the standard zone, the included angle γ is calculated by the formula:

[0123] ;

[0124] where 、 are respectively the lower limit value and the upper limit value of the standard zone on the X-axis.

[0125] Using the above formula, the included angle γ of this camera can be deduced to be 7.93°.

[0126] As Figure 5As shown, if the optical axis of the camera is used as the basis for the movement of the standard belt, when the optical axis is in the original position, the pitch angle θ between the camera and the horizontal direction is -27.51°, that is, the angle θ between the optical axis and the horizontal direction is -27.51°. The vertical field of view angle of the selected camera for the experiment is 27°, which means that the calculation of the standard belt is accurate and useful. When moving the standard belt (for example, calculated by moving 60 cm each time), when the camera model remains unchanged, the angle γ of each camera remains fixed. If the distance factor of the target object in the actual scene is uncertain, γ is calculated according to the formula, and then the actual distance coordinates of the target object are calculated from γ, and then the actual position of the target object is deduced inversely. Thus, the actual distance of each target object on the reference plane is calculated to obtain the two-dimensional coordinates of the target object. The expression is:

[0127] ;

[0128] wherein, is the position of the target point i on the X-axis, that is, the distance between the projections of the camera and the target object on the reference plane; i is the serial number of the target point; n is the number of times the optical axis point moves.

[0129] For example, Figure 4 the position coordinates of point D in

[0130] .

[0131] S3. Convert the two-dimensional pixel coordinates of the recognized target object into three-dimensional coordinates in the actual scene through perspective projection calculation.

[0132] Specifically, as Figure 6 shown, the angle formed by the two edges of the maximum range between the vertex of the camera H and the measured target object through the lens is called the field of view angle FOV; if the target object exceeds the range of the field of view angle FOV, it will not be captured by the camera. The included angle of the field of view angle is , that is, the diagonal field of view angle.

[0133] The field of view angle includes the horizontal field of view angle (HFOV, ), the vertical field of view angle (VFOV, ), and the diagonal field of view angle (DFOV, ), wherein, , , are the included angles of the corresponding field of view angles respectively.

[0134] Usually, the default FOV without special instructions is generally the horizontal field of view angle. However, for optical devices such as cameras and video cameras, since their photosensitive surfaces are rectangular. Therefore, the method of the present invention calculates the field of view angle based on the imaging object diameter of the diagonal line of the rectangular photosensitive surface, that is, the diagonal field of view angle is used for calculation.

[0135] As shown Figure 7 in the figure, according to the imaging object diameter of the diagonal of the rectangular photosensitive surface, the object distance s is calculated using the diagonal field of view angle . Since the field of view angle of the camera is fixed and known, and the range that can be captured is also related to the focal length of the lens, the formula for the object distance s that the camera can capture is:

[0136] ;

[0137] where is the imaging object diameter of the visible range of the lens, a and b are the two endpoints of the diameter respectively, c is the center point of the target object image, and also the intersection point of the optical axis and the reference plane, is the diagonal field of view angle, and s is the object distance that the camera can capture.

[0138] Based on the fact that the two-dimensional image captured by the camera is the projection of the three-dimensional space on the plane, the position of the target object in the real world is obtained based on the camera pose and the coordinates of the target object in the image, that is, the target box.

[0139] As Figure 8 shown in the figure, let the three-dimensional coordinates of the camera be the point H(0, h, 0), and the pitch angle between H and the horizontal direction is θ; the reference plane is the plane where the target object is located and is set on the same plane as the center of the target object. According to Figure 7 , it can be known that the straight line s is the object distance that the camera can capture, that is, the optical axis. The point c( , 0, 0) is the point of the target object and also the intersection point of the optical axis and the reference plane. The coordinate of point c and the trigonometric function relationship of angle θ are obtained as:

[0140] ;

[0141] ;

[0142] According to the perspective projection principle, the image captured by the camera is perpendicular to the optical axis, and its projection on the reference plane is a trapezoid. Therefore, in order to simplify the calculation, the midpoint of the target object image is set as point c. Then, for any two-dimensional coordinate point Q( , ) in the target object image, its value is obtained through measurement transformation, and the corresponding three-dimensional coordinate of Q in the three-dimensional coordinate system is , and the formula is:

[0143] ;

[0144] Thus, the equation of the straight line l passing through the point and H is obtained:

[0145] ;

[0146] Then the coordinates of the intersection point P of the straight line l and the reference plane are ( , , ), and the three-dimensional coordinate formula of point P is:

[0147] ;

[0148] Thus, the two-dimensional coordinate point Q is mapped to the corresponding world coordinate point P in the digital twin scenario;

[0149] Convert the two-dimensional coordinates of each target in the image into three-dimensional coordinates; then multiply the three-dimensional coordinates by the transformation matrix M to perform the perspective projection transformation of the camera, so as to obtain the world coordinates corresponding to all the target objects mapped in the digital twin scenario.

[0150] The transformation matrix M of the camera is a combination of the internal parameter matrix K and the external parameter matrix [R|t], and the formula is:

[0151] M = K[R|t];

[0152] And the form of the camera internal parameter matrix K is usually shown in the following formula:

[0153] ;

[0154] Among them, and are the focal lengths of the camera in the x and y directions respectively, in pixels; and are the principal point coordinates of the image (i.e., the pixel coordinates of the center of the image).

[0155] The external parameter matrix [R|t] describes the position and orientation of the camera. In the model proposed in the embodiment of the present invention, the selected coordinate system is aligned with the origin of the world coordinate system and there is no rotation. At this time, the transformation matrix M is the internal parameter matrix K. That is:

[0156] ;

[0157] t = ;

[0158] Among them, R is the rotation matrix, is the identity matrix, indicating no rotation; t is the translation vector, indicating that the camera is located at the origin of the world coordinate.

[0159] S4. Based on the adjacent point position deviation correction algorithm of Procrustes analysis, correct the calculated three-dimensional coordinates.

[0160] In this embodiment, Procrustes Analysis is a method for comparing the consistency of two sets of data by analyzing the shape distribution. Mathematically speaking, it is to continuously iterate to find the standard shape and use the least squares method to find the affine transformation method from the shape of each object to this standard shape. This process is also called the least squares orthogonal mapping.

[0161] The adjacent point position rectification algorithm based on Procrustes Analysis performs rectification by finding the optimal rotation matrix through singular value decomposition (SVD). Specifically:

[0162] Taking the three-dimensional coordinates of the calculated target object as the set of points to be rectified, select a point to be rectified from the set of points to be rectified and select the three-dimensional coordinates of the actual point position of a corresponding target object as the reference point. In this embodiment, the center point of the target object within the range of the camera standard belt and not blocked is preferentially selected as the reference point.

[0163] Calculate and of the centroids, and translate the centroid of to the centroid of to make the centroids of the two coincide and eliminate the translation difference. The formula is:

[0164] ;

[0165] ;

[0166] where N is the total number of point positions on the target object; is the centroid of the calculated point position ; is the centroid of the actual point position ;

[0167] Calculate the covariance matrix of the centered reference point and the point to be rectified. The formula is:

[0168] ;

[0169] where is the transposed matrix of the coordinate point position matrix calculated after centering the centroid ; is the actual point position matrix after centering; is the covariance matrix; T represents the transposed matrix;

[0170] Perform SVD decomposition on the covariance matrix . The formula is:

[0171] ;

[0172] Among them, U and V are orthogonal matrices of SVD decomposition; Σ is a diagonal matrix, i.e., the singular value matrix;

[0173] Calculate the rotation matrix R through U and V, and calculate the translation vector t in combination with the centroids of two points. The formula is:

[0174] ;

[0175] ;

[0176] Among them, R is the rotation matrix; t is the translation vector;

[0177] Apply the rotation matrix R and the translation vector t to each target object point position in the point set to be corrected to obtain the new coordinates after correction and alignment. The formula is:

[0178] ;

[0179] Among them, is the point after correction and alignment.

[0180] S5. Input the corrected three-dimensional coordinates into the pre-built BIM model to realize the real-time mapping of dynamic objects in the digital twin scenario.

[0181] When inputting the corrected three-dimensional coordinates into the pre-built BIM model, first import the corrected three-dimensional coordinates into an Excel file, and use the Dynamo plug-in to import the Excel file into the BIM model to complete the mapping of the digital twin scenario.

[0182] Import the calculated three-dimensional coordinates of the target object into the BIM model using the Dynamo plug-in and align them with the target object in the actual scenario. Although the coordinates imported into the BIM are corrected, due to factors such as calculation and lens distortion, there may still be inaccurate matching of the target object positions in the model and the actual scenario (such as a certain distance between the positions in the BIM and the actual scenario, or the target object that should be at point A is matched to point C, etc.). Therefore, it can be further adjusted in the BIM model to achieve accurate mapping between the actual scenario and the digital scenario.

[0183] The following methods can be used for adjustment:

[0184] (1) Auxiliary calibration through grid drawing

[0185] In the BIM model, a three-dimensional coordinate system consistent with the physical coordinate system of the actual scene is constructed, and then each point in the scene is calibrated by drawing a square grid. The nodes of the grid should correspond one by one to the reference points in the actual scene. Each grid point can represent a reference coordinate in the actual scene in the BIM model. Then, according to the accuracy requirements, the grid is subdivided into a size of (such as 6 cm × 6 cm).

[0186] Set the relative position of the target object on the grid and perform preliminary placement according to the coordinate data obtained in steps S3 to S4. Combining the data in the actual scene, adjust the position of the target object in the grid through camera calibration or manual adjustment to make it as aligned as possible with the points in the actual scene.

[0187] (2) Optimization alignment method based on a three-dimensional coordinate system

[0188] First, select some clear and easy-to-calibrate control points (such as corner points, reference planes, etc.) in the actual scene. The three-dimensional coordinates of these points need to be accurately corresponding in the BIM model. Through the coordinate data obtained in steps S3 to S4, these points can be transformed from the camera coordinate system to the world coordinate system, and then according to the differences between the existing points in the BIM model and the actual coordinate points, further adjustments are made using the node instructions in Dynamo.

[0189] (3) Automatic correction and matching

[0190] Based on the point position errors between the existing BIM points and the actual scene, write a script in Dynamo. After importing the three-dimensional coordinates obtained in steps S3 to S4 into the BIM model, automatically calculate and update the point positions, so as to adjust the position of the target object in the BIM to minimize the difference from the actual point positions. And use a dynamic update method in the BIM to gradually correct the errors. Implement a dynamic feedback mechanism in the BIM model. When new actual scene data is input, the system can automatically adjust and output the adjusted coordinates to ensure the accurate alignment of the position of the target object in the scene.

[0191] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0192] (1) The present invention enhances the scene adaptability in complex environments. Aiming at the poor adaptability of the prior art to complex construction sites and the difficulty in dealing with problems such as target object occlusion, light changes, and multi-target interference, through sensors (such as cameras) and deep learning technologies, the robustness of target object detection and coordinate calculation in complex environments is improved, ensuring the stability and accuracy of the technology under complex conditions such as target object occlusion and light changes.

[0193] (2) The present invention solves the problem of insufficient coordinate calculation accuracy caused by camera distortion. The prior art does not fully consider the impact of camera distortion, resulting in large errors in the coordinate calculation of the target object. By introducing a "standard belt" with minimal camera distortion influence obtained from relevant experimental tests, the influence of camera distortion on the target object coordinate calculation is weakened and effectively avoided, thereby improving the coordinate calculation accuracy of the target object. In particular, in the calculation of targets within the distance range of the "standard belt", the accuracy of coordinate calculation can be ensured, and by combining the change in calculation angle with the projection principle, the actual position of the target object can be inferred, and the dynamic tracking function of the image device can be fed back.

[0194] (3) The present invention realizes the real-time mapping and updating of dynamic objects, meeting the real-time requirements of dynamic scenes. The existing technology is usually based on BIM static modeling, lacks the ability to map dynamic objects in real time, has low coordinate calculation accuracy, and is difficult to meet the real-time update requirements of dynamic scenes such as construction sites. Through the Dynamo plug-in built into the BIM software, the BIM model is deeply integrated with dynamic data collection, and the dynamic object information (such as building materials, machinery and equipment, construction personnel, etc.) in the digital twin model is updated in real time to ensure that the digital twin model can accurately reflect the dynamic changes of the construction site. The coordinate position is corrected by the adjacent point correction algorithm based on Procrustes analysis, and then the corrected coordinate data is input into the script instruction written by Dynamo through an Excel file to realize the automatic adjustment of the target object in the BIM model, improve the accuracy of coordinate calculation, and meet the real-time and accuracy requirements of dynamic scenes.

[0195] Embodiment 2

[0196] like Figure 9 As shown, the second embodiment of the present invention further provides a digital twin scene dynamic mapping generation device based on standard band correction, comprising:

[0197] The image acquisition unit is used to collect dynamic image data sets of the construction site and process each image in the image data set to select and classify each target object in the image;

[0198] The target recognition unit is used to set and adjust the standard band of the image to correct the image perspective, frame all the targets in the image within the standard band for target recognition; wherein, the standard band is: based on the camera imaging principle, a position area within a specific range from the camera and with a weak influence on the image distortion; and it can move this area by simulating the adjustment of the pitch angle between the camera and the horizontal direction to dynamically capture and frame the targets in the image, ensuring that the target coordinates at all points can be accurately calculated. Specifically: Let the height of the camera be h, that is, the two-dimensional coordinate is (0, h); the reference plane is the plane where the target is located, that is, the X-axis; the angle between the camera and the projection range of the standard band is γ, and the angle γ is calculated according to the height h of the camera and the distance between the camera and the standard band. The formula is:

[0199] ;

[0200] Among them, 、 are respectively the lower limit value and the upper limit value of the standard band on the X-axis;

[0201] If the camera optical axis is used as the basis for moving the standard band, when moving the standard band, according to the fixed angle γ of each camera, the actual distance of each target on the reference plane is calculated, so as to obtain the two-dimensional position of the target. The expression is:

[0202] ;

[0203] Among them, is the position of the target point i on the X-axis, that is, the distance between the projections of the camera and the target on the reference plane; i is the serial number of the target point; n is the number of times the optical axis point moves; θ is the pitch angle between the camera and the horizontal direction;

[0204] The three-dimensional conversion unit is used to convert the two-dimensional pixel coordinates of the recognized target into three-dimensional coordinates in the actual scene through perspective projection calculation;

[0205] The three-dimensional coordinate deviation correction unit is used to correct the calculated three-dimensional coordinates based on the adjacent point deviation correction algorithm of Procrustes analysis;

[0206] The real-time mapping unit is used to input the corrected three-dimensional coordinates into the pre-built BIM model to realize the real-time mapping update of the dynamic objects in the digital twin scene.

[0207] Embodiment III

[0208] The third embodiment of the present invention also provides a digital twin scenario dynamic mapping generation device based on standard band calibration, which includes a memory and a processor. A computer program is stored in the memory and can be executed by the processor to implement the digital twin scenario dynamic mapping generation method based on standard band calibration as described above.

[0209] Embodiment 4

[0210] The fourth embodiment of the present invention also provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, the digital twin scenario dynamic mapping generation method based on standard band calibration as described above is implemented.

[0211] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0212] In addition, the functional modules in each embodiment of the present invention can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0213] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs. It should be noted that in this article, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including the said element.

[0214] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.

[0215] It should be understood that the term "and / or" used herein is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0216] Depending on the context, the word "if" as used herein can be interpreted as "when", "while", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined", "in response to determining", "when detecting (stated condition or event)", or "in response to detecting (stated condition or event)".

[0217] The "first / second" mentioned in the embodiments is only used to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged with a specific order or sequence under allowable circumstances. It should be understood that the objects distinguished by the "first / second" can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0218] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for generating dynamic mapping of a digital twin scenario based on standard belt calibration, characterized in that Including: S1. Collect a dynamic image dataset of the construction site, and process each image in the image dataset to box and classify each target object in the image; S2. Set and adjust the standard band of the image to correct the image perspective, and frame all target objects in the standard band for target recognition; wherein, the standard band is: based on the camera imaging principle, a position area within a specific range from the camera and with weak distortion influence on the image; and it is possible to move this area by simulating and adjusting the pitch angle between the camera and the horizontal direction to dynamically capture and frame the target objects in the image, ensuring that the target object coordinates at all points can be accurately calculated. Specifically: Let the height of the camera be h, that is, the two-dimensional coordinate is (0, h); the reference plane is the plane where the target object is located, that is, the X-axis; the included angle between the camera and the projection range of the standard band is γ, and the included angle γ is calculated according to the distance between the camera height h and the standard band. The formula is: ; Among them, and are the lower limit value and the upper limit value of the standard band on the X-axis, respectively; If the camera optical axis is used as the basis for moving the standard band, when moving the standard band, according to the fixed included angle γ of each camera, calculate the actual distance of each target object on the reference plane, so as to obtain the two-dimensional position of the target object. The expression is: ; Among them, is the position of the target point i on the X-axis, that is, the distance between the camera and the target projected on the reference plane; i is the serial number of the target point; n is the number of times the optical axis point moves; θ is the pitch angle between the camera and the horizontal direction. S3. Convert the two-dimensional pixel coordinates of the identified target objects into three-dimensional coordinates in the actual scene through perspective projection calculation; S4. Correct the calculated three-dimensional coordinates based on the adjacent point correction algorithm of Procrustes analysis; S5. Input the corrected three-dimensional coordinates into the pre-built BIM model to realize the real-time mapping update of the dynamic objects in the digital twin scene.

2. The dynamic mapping generation method of a digital twin scenario based on standard tape calibration according to claim 1, wherein , The target objects are safety factors existing at the construction site, including personnel and construction items.

3. A method for dynamically mapping and generating a digital twin scenario based on standard tape calibration according to claim 1, characterized in that , The operations for processing each image include: Filter and screen the image dataset to remove images in the scene that do not contain target objects or have incomplete target object shooting, images with blurred target objects due to environmental impacts, and duplicate images; Use deep learning technology to box and classify each target object in the image data.

4. A method for dynamically mapping and generating a digital twin scenario based on standard belt calibration according to claim 1, characterized in that , The specific content of S3 is: The angle formed by taking the camera as the vertex and the two outermost edges of the maximum range of the measured target object passing through the lens is called the field of view angle FOV; if the target object exceeds the range of the field of view angle FOV, it will not be captured by the camera, and the included angle of the field of view angle is , that is, the diagonal field of view angle; According to the imaging object diameter of the diagonal of the rectangular photosensitive surface, the diagonal field of view angle is adopted Calculate the object distance s. Since the field of view angle of the camera is fixed and known, and the range that can be photographed is also related to the focal length of the lens, the formula for the object distance s that the camera can photograph is as follows: ; Among them, is the imaging object diameter within the visible range of the lens, c is the point of the target object image, and it is also the intersection point of the optical axis and the reference plane; According to the two-dimensional image captured by the camera being the projection of the three-dimensional space on the plane, based on the camera pose and the coordinates of the target object in the image, that is, the target box, obtain the position of the target object in the real world; Let the three-dimensional coordinates of the camera be the point H(0, h, 0), and the pitch angle between H and the horizontal direction be θ; the reference plane is the plane where the target object is located and is set on the same plane as the center of the target object. s is the object distance that the camera can capture, that is, the optical axis. Let the point c( , 0, 0) be the point of the target object and also the intersection point of the optical axis and the reference plane. The trigonometric function relationship of the coordinates of point c and the angle θ is as follows: ; ; According to the perspective projection principle, the image captured by the camera is perpendicular to the optical axis, and its projection on the reference plane is trapezoidal. Therefore, the midpoint of the target object image is set as point c. Then, for any two-dimensional coordinate point Q ( , ) in the target object image, its value is obtained through measurement transformation, and the corresponding three-dimensional coordinates of Q in the three-dimensional coordinate system are calculated as , and the formula is: ; Thus, the equation of the straight line l passing through the point and H is obtained as follows: ; Then the coordinates of the intersection point P of the straight line l and the reference plane are ( , , ), and the three-dimensional coordinate formula of point P is: ; Thus, obtain the two-dimensional coordinate point Q mapped to the corresponding world coordinate point P in the digital twin scene; Convert the two-dimensional coordinates of each target in the image into three-dimensional coordinates; then perform perspective projection transformation of the three-dimensional coordinates by the camera, so as to obtain the world coordinates corresponding to all target objects mapped to the digital twin scene.

5. A method for dynamically mapping and generating a digital twin scenario based on standard belt calibration according to claim 1, characterized in that , The adjacent point correction algorithm based on Procrustes analysis corrects by finding the optimal rotation matrix through singular value decomposition SVD. Specifically: Taking the three-dimensional coordinates of the calculated target as the set of points to be corrected, select a point to be corrected from the set of points to be corrected and select the three-dimensional coordinates of the actual position of a corresponding target as the reference point; Calculate and centroid, and translate the centroid of to the centroid of so that the centroids of the two coincide, eliminating the translational difference. The formula is: ; ; Where N is the total number of points on the target object; is the calculated centroid of the point ; is the centroid of the actual point ; Calculate the covariance matrix of the centered reference point and the point to be corrected , and the formula is: ; Among them, is the centroid which is the transpose matrix of the coordinate point matrix calculated after centering; is the actual point matrix after centering; is the covariance matrix; T represents the transpose matrix; For the covariance matrix perform SVD decomposition, and the formula is: ; Among them, U and V are the orthogonal matrices of SVD decomposition; Σ is a diagonal matrix, that is, the singular value matrix; Calculate the rotation matrix R through U and V, and calculate the translation vector t in combination with the centroid of the two points. The formula is: ; ; Among them, R is the rotation matrix; t is the translation vector; Apply the rotation matrix R and the translation vector t to each target object position in the point set to be rectified to obtain the new coordinates after rectification and alignment. The formula is as follows: ; Among them, is the point after deviation correction and alignment.

6. A method for generating a dynamic mapping of a digital twin scenario based on standard tape calibration according to claim 1, wherein , When inputting the rectified three-dimensional coordinates into the pre-built BIM model, import the rectified three-dimensional coordinates into an Excel file, and use the Dynamo plug-in to import the Excel file into the BIM model to complete the mapping of the digital twin scenario.

7. A method for generating a dynamic mapping of a digital twin scenario based on standard belt calibration according to claim 1, characterized in that , It also includes: adjusting the target object coordinates in the BIM model through the grid drawing-assisted calibration method to achieve the precise mapping of the actual scenario and the digital scenario; specifically: In the BIM model, construct a three-dimensional coordinate system consistent with the physical coordinate system of the actual scenario; Calibrate each point in the scenario through a square grid, and the nodes of the grid should correspond one-to-one with the reference points in the actual scenario; Each grid point represents a reference coordinate in the actual scenario in the BIM model; According to the accuracy requirements, the grid is subdivided into in size; Combine the data of the target object in the actual scenario to adjust the target object within the grid.

8. A digital twin scenario dynamic mapping generation device based on standard belt calibration, characterized in that An image acquisition unit, configured to acquire a dynamic image data set of the construction site, and process each image in the image data set to perform bounding box selection and category annotation on each target object in the image; A target recognition unit, configured to set and adjust the standard belt of the image to correct the image perspective, and frame all target objects in the image within the standard belt for target recognition; wherein, the standard belt is: based on the camera imaging principle, a position area at a specific range from the camera and with a weak influence on the image distortion; and it is possible to move this area by simulating and adjusting the pitch angle of the camera with respect to the horizontal direction to dynamically capture and frame the target objects in the image, ensuring that the target object coordinates of all points can be accurately calculated. Specifically: Let the height of the camera be h, that is, the two-dimensional coordinate is (0, h); the reference plane is the plane where the target object is located, that is, the X-axis; the included angle between the camera and the projection range of the standard belt is γ, and the included angle γ is calculated according to the height h of the camera and the distance between the camera and the standard belt. The formula is: ; Among them, and are the lower limit value and the upper limit value of the standard belt on the X-axis, respectively; If the camera optical axis is used as the basis for moving the standard belt, when moving the standard belt, according to the fact that the included angle γ of each camera remains unchanged, calculate the actual distance of each target object on the reference plane, so as to obtain the two-dimensional position of the target object. The expression is: ; Among them, is the position of the target point i on the X-axis, that is, the distance between the camera and the target projected on the reference plane respectively; i is the serial number of the target point; n is the number of times the optical axis point moves; θ is the pitch angle between the camera and the horizontal direction; A three-dimensional conversion unit, configured to convert the two-dimensional pixel coordinates of the recognized target object into three-dimensional coordinates in the actual scenario through perspective projection calculation; A three-dimensional coordinate rectification unit, configured to rectify the calculated three-dimensional coordinates based on the adjacent point rectification algorithm based on Procrustes analysis; A real-time mapping unit, configured to input the rectified three-dimensional coordinates into the pre-built BIM model to achieve real-time mapping update of the dynamic objects in the digital twin scenario.

Citation Information

Patent Citations

  • High-precision alarm positioning method and device based on PTZ (Pan / Tilt / Zoom) camera calibration

    CN119131150A

  • Multi-modal digital twinning scene building method and device based on image semantic fusion

    CN119251687A