Indoor construction scene component positioning method and device, computer equipment and storage medium

By acquiring construction videos and combining them with OpenVSLAM and affine transformation matrices, along with improved YOLOv8-DB and ZoeDepth algorithms, precise positioning of components in indoor construction scenarios was achieved. This solved the problems of inaccurate positioning and low efficiency in existing technologies, and improved the intelligence and efficiency of construction management.

CN120953494APending Publication Date: 2025-11-14TONGJI UNIV +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511043069.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies are inaccurate and inefficient in indoor construction scenarios. Traditional measurement methods are time-consuming, labor-intensive, and prone to errors. Existing automation technologies perform poorly in complex environments.

Method used

By acquiring indoor construction videos, the inspection path is extracted using the OpenVSLAM algorithm and projected onto a two-dimensional planar drawing. The scale information is supplemented by affine transformation matrix. The construction components are detected using an improved YOLOv8-DB model and the ZoeDepth algorithm, and a three-dimensional topological network of path-drawing-target is constructed.

Benefits of technology

It enables precise positioning of components in indoor construction scenes with dimensional information, improving the intelligence level and work efficiency of construction management, and is applicable to indoor construction projects of various sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953494A_ABST
    Figure CN120953494A_ABST
Patent Text Reader

Abstract

The invention discloses an indoor construction scene component positioning method and device, computer equipment and a storage medium. The method comprises the steps of obtaining an indoor construction video; determining an inspection path based on the indoor construction video, and extracting a two-dimensional plane coordinate taking an initial position as an original point from the indoor construction video to obtain a plane drawing; determining an affine transformation matrix based on the plane drawing, mapping the inspection path to the plane drawing by using the affine transformation matrix, and supplementing scale information of the path to obtain an aligned inspection path; detecting construction components and determining pixel coordinates of the construction components on a plane drawing based on the aligned inspection paths in combination with an improved YOLOv8-DB model and a depth estimation algorithm, and forming a path-drawing-target three-dimensional topology network; and outputting the three-dimensional topology network. By implementing the method, the intelligent level and the working efficiency of construction management can be greatly improved, and the method is suitable for indoor construction projects of various scales.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a component positioning method, and more specifically to a component positioning method, device, computer equipment, and storage medium for indoor construction scenarios. Background Technology

[0002] In the construction industry, precise component positioning is crucial for ensuring project quality and schedule. Traditionally, component positioning on construction sites has relied primarily on manual measurement and recording. This method is not only time-consuming and labor-intensive but also prone to human error, especially in complex or large-scale projects. For example, traditional measurement methods require the use of tools such as measuring tapes and total stations for physical measurements and manual recording of coordinates, leading to inefficiency and difficulty in ensuring data consistency and accuracy. Furthermore, while 2D CAD drawings provide detailed information during the design phase, the information on the drawings often cannot be directly applied to the precise positioning of components on-site due to changes in the actual construction environment. In recent years, 3D scanning technologies such as laser scanning have been gradually introduced into construction sites to obtain high-precision spatial data; however, these technologies are expensive and complex to operate, making them unsuitable for large-scale deployment.

[0003] Existing automated positioning technologies also have their limitations. While monocular SLAM (Simultaneous Localization and Mapping) can recover camera motion trajectories to some extent, the generated paths lack scale information and cannot be directly applied to component positioning in real-world construction scenarios. MVG (Multi-View Geometry) methods reconstruct 3D structures by analyzing images from multiple perspectives; however, their computational complexity is high, and they perform poorly in dynamically changing construction scenarios. Methods combining object detection and depth estimation, such as the YOLO series of deep learning-based object detection algorithms, have achieved significant results in many fields, but their performance still needs further optimization in complex indoor construction environments due to factors such as lighting changes and occlusion.

[0004] Therefore, it is necessary to design a new method to achieve accurate and dimensional component positioning in indoor construction scenarios, which will significantly improve the intelligence level and work efficiency of construction management, be applicable to indoor construction projects of various scales, and solve the problems of inaccurate component positioning and low efficiency in existing technologies. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method, device, computer equipment and storage medium for positioning components in indoor construction scenarios.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a method for positioning components in an indoor construction scene, comprising:

[0007] Obtain indoor construction videos;

[0008] The inspection route is determined based on the indoor construction video, and two-dimensional plane coordinates with the initial position as the origin are extracted from the indoor construction video to obtain a plan drawing.

[0009] Based on the planar drawing, an affine transformation matrix is ​​determined, and the inspection path is mapped onto the planar drawing using the affine transformation matrix. The scale information of the path is then added to obtain an aligned inspection path.

[0010] Based on the aligned inspection path, combined with the improved YOLOv8-DB model and depth estimation algorithm, construction components are detected and their pixel coordinates on the planar drawing are determined, forming a three-dimensional topology network of path-drawing-target.

[0011] Output the three-dimensional topology network.

[0012] The further technical solution is as follows: The process of determining the inspection path based on the indoor construction video and extracting two-dimensional plane coordinates from the indoor construction video with the initial position as the origin to obtain a planar drawing includes:

[0013] The indoor construction video data is processed using the OpenVSLAM algorithm to extract key frames and corresponding pose information in order to obtain the inspection path.

[0014] The inspection path is projected onto a two-dimensional plane, and the two-dimensional coordinates of each trajectory point are recorded to obtain a planar drawing.

[0015] The further technical solution is as follows: The step of determining the affine transformation matrix based on the planar drawing, mapping the inspection path onto the planar drawing using the affine transformation matrix, and supplementing the path's scale information to obtain an aligned inspection path includes:

[0016] Select three reference points on the plan view, and determine the position coordinates of the three reference points in the inspection path and the position coordinates on the plan view to obtain the path coordinates and the plan coordinates;

[0017] Determine the affine transformation matrix based on the path coordinates and drawing coordinates;

[0018] The affine transformation matrix is ​​used to map all trajectory points on the inspection path onto the planar drawing, and the scale information of the path is supplemented to obtain an aligned inspection path.

[0019] The further technical solution is as follows: determining the affine transformation matrix based on the path coordinates and drawing coordinates includes:

[0020] Create an affine transformation relationship between the path coordinate system and the drawing coordinate system, and use the path coordinates and drawing coordinates in combination with the affine transformation matrix to calculate the optimal affine transformation parameters, obtain the best estimate of the transformation matrix and offset, and thus determine the affine transformation matrix.

[0021] The further technical solution is as follows: The step of mapping all trajectory points on the inspection path onto the planar drawing using the affine transformation matrix, and supplementing the path's scale information to obtain an aligned inspection path, includes:

[0022] The affine transformation matrix is ​​applied to all trajectory points on the inspection path, and the inspection path is completely mapped to the coordinate system of the planar drawing through the coordinate transformation formula.

[0023] Based on the transformed trajectory point coordinates, the inspection path curve is drawn on the plan view, and the true scale information of the path is restored by the map scale to obtain the path-map mapping relationship with actual physical dimensions.

[0024] Mark the actual coordinates of three pairs of reference points on the plan drawing to further confirm the precise alignment between the path and the drawing, so as to obtain the aligned inspection path.

[0025] Its further technical solution is as follows: Based on the aligned inspection path, combined with the improved YOLOv8-DB model and depth estimation algorithm, the construction components are detected and their pixel coordinates on the planar drawing are determined, forming a three-dimensional topology network of path-drawing-target, including:

[0026] An improved YOLOv8-DB model was used to analyze the indoor construction video to obtain the pixel coordinates and detection boxes of the construction components.

[0027] The absolute depth distance of the center point of the construction component is extracted using the ZoeDepth algorithm. Based on the offset of the pixel coordinates relative to the image center and combined with the perspective characteristics of the panoramic video, the azimuth angle of the construction component relative to the camera's principal optical axis is calculated.

[0028] The aligned inspection path is processed by affine transformation to determine the position of the camera corresponding to the current frame in the coordinate system of the planar drawing, so as to obtain the camera position.

[0029] By applying the polar coordinate to Cartesian coordinate conversion formula and combining it with the camera position, the absolute depth distance and azimuth of the center point of the construction component are converted into pixel coordinates on the planar drawing;

[0030] By associating the pixel coordinates of all construction components on the planar drawing with the timestamps of the inspection path, a three-dimensional topology network of path-drawing-target is constructed.

[0031] The further technical solution is as follows: the improved YOLOv8-DB model includes a backbone network that uses deformable convolutional modules to replace standard convolutions, a neck network structure that introduces a weighted bidirectional feature pyramid structure and fuses multi-level features through learnable weight parameters, and a detection head that adds dynamic convolutional layers; and the improved YOLOv8-DB model uses a method of randomly adjusting the input image size for multi-scale training.

[0032] The further technical solution is as follows: The ZoeDepth algorithm is used to extract the absolute depth distance of the center point of the construction component, and based on the offset of the pixel coordinates relative to the image center, combined with the viewpoint characteristics of the panoramic video, the azimuth angle of the construction component relative to the camera's principal optical axis is calculated, including:

[0033] The ZoeDepth algorithm is used to generate a pixel-by-pixel depth map of the indoor construction video, and the depth value of the corresponding detection box center point is recorded.

[0034] The pixel coordinates of each view in the six-view panoramic video corresponding to the indoor construction video are normalized to the center of their respective image coordinate system. By calculating the horizontal offset and known parameters, the pixel coordinates are converted into azimuth angles relative to the camera's principal optical axis, so as to obtain the azimuth angle of the construction component relative to the camera's principal optical axis.

[0035] The present invention also provides a component positioning device for indoor construction scenes, comprising:

[0036] The video acquisition unit is used to acquire indoor construction videos.

[0037] The plan drawing acquisition unit is used to determine the inspection path based on the indoor construction video and extract two-dimensional planar coordinates from the indoor construction video with the initial position as the origin to obtain a plan drawing.

[0038] The alignment unit is used to determine the affine transformation matrix based on the planar drawing, and to map the inspection path onto the planar drawing using the affine transformation matrix, and to supplement the scale information of the path to obtain the aligned inspection path.

[0039] The topology network generation unit is used to detect construction components and determine the pixel coordinates of the construction components on the planar drawing based on the aligned inspection path combined with the improved YOLOv8-DB model and depth estimation algorithm, and form a three-dimensional topology network of path-drawing-target.

[0040] The output unit is used to output the three-dimensional topology network.

[0041] The present invention also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the above-described method.

[0042] The advantages of this invention compared to existing technologies are as follows: This invention generates a planar drawing by acquiring indoor construction video and extracting two-dimensional planar coordinates with the initial position as the origin. Then, an affine transformation matrix is ​​used to accurately map the inspection path onto this planar drawing, and path scale information is supplemented to obtain an aligned inspection path. Combining an improved YOLOv8-DB model and depth estimation algorithm, construction components are detected based on the aligned inspection path, and their pixel coordinates on the planar drawing are determined, thereby constructing a three-dimensional topological network of path-drawing-target, and finally outputting this network model. This method achieves accurate and scale-informed component positioning in indoor construction scenes, greatly improving the intelligence level and work efficiency of construction management. It is applicable to indoor construction projects of various scales and effectively solves the problems of inaccurate component positioning and low efficiency in existing technologies.

[0043] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Attached Figure Description

[0044] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a schematic diagram illustrating an application scenario of the component positioning method for indoor construction scenes provided in this embodiment of the invention;

[0046] Figure 2 A flowchart illustrating the component positioning method for indoor construction scenarios provided in an embodiment of the present invention;

[0047] Figure 3 A schematic diagram of a panoramic image of an indoor construction video provided in an embodiment of the present invention;

[0048] Figure 4 A schematic diagram illustrating feature extraction and matching of panoramic images provided in an embodiment of the present invention;

[0049] Figure 5 A schematic diagram of the inspection path provided in an embodiment of the present invention;

[0050] Figure 6 This is a schematic diagram of an aligned inspection path provided in an embodiment of the present invention.

[0051] Figure 7 This is a schematic diagram of the inspection of construction components provided in an embodiment of the present invention;

[0052] Figure 8 A schematic diagram illustrating the combination of construction component detection and depth estimation provided in an embodiment of the present invention;

[0053] Figure 9 This is a schematic diagram of the positioning of the construction component provided in an embodiment of the present invention;

[0054] Figure 10 A schematic block diagram of an indoor construction scene component positioning device provided in an embodiment of the present invention;

[0055] Figure 11 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0058] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0059] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0060] Please see Figure 1 and Figure 2 , Figure 1 This is a schematic diagram illustrating the application scenario of the component positioning method for indoor construction scenes provided in this embodiment of the invention. Figure 2This is a schematic flowchart illustrating the component positioning method for indoor construction scenes provided in this embodiment of the invention. This method is applied in a server. The server interacts with the terminal, processes indoor construction video, extracts the inspection path using the OpenVSLAM algorithm, and projects it onto a two-dimensional plane to generate an initial planar drawing. Then, based on selected reference points, an affine transformation matrix is ​​determined to accurately map the inspection path onto the planar drawing and supplement its scale information, ensuring precise alignment between the path and the drawing. Furthermore, combining an improved YOLOv8-DB model and the ZoeDepth depth estimation algorithm, construction components are automatically detected, their pixel coordinates and absolute depth on the planar drawing are determined, and a three-dimensional topological network of "path-drawing-target" is constructed through polar coordinate to Cartesian coordinate conversion, achieving precise component positioning. This method not only significantly improves the intelligence level and work efficiency of construction management but is also applicable to indoor construction projects of various scales, effectively solving the problems of inaccurate component positioning and low efficiency in existing technologies.

[0061] Figure 2 This is a flowchart illustrating the component positioning method for indoor construction scenarios provided in an embodiment of the present invention. Figure 2 As shown, the method includes the following steps S110 to S150.

[0062] S110. Obtain indoor construction videos.

[0063] In this embodiment, indoor construction video refers to continuous image data of indoor construction scenes collected by a monocular camera mounted on inspection equipment (such as a mobile robot, drone, or handheld shooting device). Specifically, to fully utilize multi-view information, such as... Figure 3 As shown, the video was recorded using panoramic video technology, meaning that each video frame can cover a 360-degree horizontal field of view. It is usually composed of multiple cameras or fisheye lenses to ensure comprehensive and blind-spot-free monitoring of the entire indoor space.

[0064] This panoramic video not only records the various components and environmental layout within the construction site, but also captures information about the camera's motion trajectory, providing necessary input for subsequent processing. Specifically, the panoramic video should meet the following conditions to support the method of this embodiment:

[0065] Resolution and sharpness: Even in panoramic mode, the video needs sufficient resolution and sharpness to accurately identify and locate various components in the construction scene. This means the video should be high-definition or even ultra-high-definition to ensure that even smaller or highly detailed components can be clearly observed.

[0066] Multi-view integration: Panoramic videos are typically composed of multiple viewpoints, each with a 90-degree horizontal field of view, collectively forming a complete 360-degree view. These viewpoints need to be precisely synchronized and calibrated to ensure seamless stitching into a panoramic image and to provide consistent spatial awareness.

[0067] Stability and Dynamic Range: Although panoramic video can provide a full-view perspective, it is still necessary to pay attention to image stability and adapt to different lighting conditions during shooting to improve the accuracy of visual SLAM algorithms (such as OpenVSLAM) in recovering camera motion trajectories. Furthermore, good dynamic range helps to cope with changes in indoor and outdoor lighting, ensuring image quality under different brightness conditions.

[0068] Time synchronization: For situations involving multi-sensor data fusion (such as depth information), precise time synchronization is required between video frames and other types of data (such as depth maps) to ensure data consistency and reliability.

[0069] Applying specific algorithms: Due to the unique nature of panoramic video, specially tuned algorithms (such as the improved YOLOv8-DB model combined with the ZoeDepth algorithm) are required for object detection and depth estimation to adapt to the characteristics of panoramic images and improve detection accuracy.

[0070] In conclusion, acquiring high-quality panoramic indoor construction videos is a crucial step in achieving precise component positioning. By processing and analyzing these videos, rich information can be extracted to construct a three-dimensional topology network, thereby enabling digital mapping and management of the construction scene. This method significantly improves the efficiency and accuracy of traditional construction management processes, making project progress tracking and quality control more intelligent and scientific.

[0071] S120. Determine the inspection path based on the indoor construction video, and extract the two-dimensional plane coordinates of the indoor construction video with the initial position as the origin to obtain the plan drawing.

[0072] In this embodiment, the planar drawing refers to a simplified diagram representing the inspection route and positional relationship formed by projecting the inspection path onto a two-dimensional plane, establishing a coordinate system with the initial position as the origin, and recording the two-dimensional coordinates of each trajectory point.

[0073] In one embodiment, step S120 described above may include S121 to S122.

[0074] S121. Process the indoor construction video data using the OpenVSLAM algorithm to extract key frames and corresponding pose information to obtain the inspection path.

[0075] In this embodiment, as Figure 5As shown, the inspection path refers to the trajectory formed by the inspection equipment moving along a predetermined route within the construction site. This path not only includes the equipment's movement trajectory in space but also records the equipment's attitude (i.e., orientation) at various locations. Specifically, at each point in time, the equipment's position can be represented by three-dimensional coordinates, while its attitude can be described as the rotation angle around three principal axes (usually the X, Y, and Z axes). However, since a monocular camera cannot directly provide dimensional information, other methods are needed to reconstruct the actual dimensions.

[0076] In this embodiment, the OpenVSLAM algorithm is used to process the continuous video stream acquired from a monocular camera. OpenVSLAM is an open-source visual SLAM system that can automatically identify feature points from the input video sequence and estimate the camera's motion trajectory using the correspondence between these feature points.

[0077] The specific process is as follows:

[0078] like Figure 4 As shown, OpenVSLAM first performs feature detection on each frame of the image, using commonly used features such as SIFT, SURF, or ORB. Next, the algorithm searches for common feature points between adjacent frames to establish a matching relationship between features.

[0079] Based on the feature matching results above, OpenVSLAM uses geometric constraints (such as the fundamental matrix or essential matrix) and the PnP (Perspective-n-Point) problem-solving method to calculate the camera pose (i.e., position and orientation) corresponding to each frame of the image.

[0080] As more frames are processed, the system gradually builds a map containing a large number of feature points and their associated information. To improve accuracy, graph optimization techniques are usually applied to make global adjustments to the entire trajectory.

[0081] During the process, not all frames are saved as keyframes; only those frames that contain enough new information or are located in important positions are selected as keyframes for subsequent analysis.

[0082] Through the above steps, a series of keyframes representing the inspection path and their precise pose information can be obtained, which forms the basis for subsequent operations.

[0083] S122. Project the inspection path onto a two-dimensional plane and record the two-dimensional coordinates of each trajectory point to obtain a planar drawing.

[0084] In this embodiment, although OpenVSLAM provides path information in three-dimensional space, in many application scenarios, especially when comparing with two-dimensional architectural drawings, we need to simplify the three-dimensional path to a two-dimensional form. Therefore, this step performs the following operations:

[0085] Given that most indoor environments have relatively flat floor structures, we can assume that the variation in the Z-axis direction is small. Therefore, we can choose to ignore the Z-axis component and project the 3D coordinates onto the XY plane. This approach requires ensuring that the area of ​​interest is indeed close to a horizontal plane; otherwise, more complex transformation methods may be necessary.

[0086] A specific point is selected as the origin (usually the starting point of the inspection), and a local two-dimensional coordinate system is established with this point as the center. The positions of all other points will be given their X and Y coordinates relative to this origin.

[0087] For each trajectory point after projection transformation, record its X and Y coordinates in the new coordinate system. This data not only reflects the actual movement of the equipment on the horizontal plane, but also provides the necessary input for subsequent path matching with drawings.

[0088] Ultimately, this process yielded a two-dimensional plan drawing with the initial position as the origin. This drawing clearly shows the specific shape of the inspection path, laying the foundation for further construction scenario analysis and component positioning.

[0089] S130. Determine the affine transformation matrix based on the planar drawing, and use the affine transformation matrix to map the inspection path onto the planar drawing, and supplement the scale information of the path to obtain an aligned inspection path.

[0090] In this embodiment, the aligned inspection path refers to the inspection path presented on the plan drawing with actual physical dimension information after the actual movement trajectory acquired by the inspection equipment is accurately matched and coordinate transformed with the plan drawing of the construction scene through an affine transformation method. This path not only reflects the movement trajectory of the inspection equipment in the indoor space, but also incorporates information such as the scale of the plan drawing, achieving consistency between the real-world scale and the drawing representation.

[0091] In one embodiment, step S130 described above may include steps S131 to S133.

[0092] S131. Select three reference points on the plan drawing and determine the position coordinates of the three reference points in the inspection path and the position coordinates on the plan drawing to obtain the path coordinates and the drawing coordinates.

[0093] In this embodiment, path coordinates refer to the position coordinates of specific points on the camera (or inspection equipment) motion trajectory reconstructed using the OpenVSLAM algorithm based on data collected by a monocular camera during the inspection process. These coordinates are values ​​in a two-dimensional coordinate system established with the inspection starting point as the origin.

[0094] Drawing coordinates refer to the positional coordinates of three selected reference points on a drawing. These reference points typically correspond to fixed landmarks in the real world or pre-set calibration boards, and their positions on the drawing can be determined through manual measurement or drawing annotation.

[0095] S132. Determine the affine transformation matrix based on the path coordinates and drawing coordinates.

[0096] In this embodiment, the affine transformation matrix is ​​a mathematical model used to describe the transformation relationship from one coordinate system to another.

[0097] Specifically, an affine transformation relationship is created between the path coordinate system and the drawing coordinate system. The optimal affine transformation parameters are calculated using the path coordinates and drawing coordinates in combination with the affine transformation matrix to obtain the best estimates of the transformation matrix and offset, thereby determining the affine transformation matrix.

[0098] In this context, it is used to calculate how to accurately map the trajectory points in the inspection path coordinate system to the planar drawing coordinate system. Specifically, by selecting three pairs of path coordinates and drawing coordinates, an affine transformation matrix can be constructed and solved. This matrix contains translation, rotation, scaling, and shearing parameters to achieve accurate transformation between the two coordinate systems.

[0099] S133. Using the affine transformation matrix, map all trajectory points on the inspection path onto the planar drawing, and supplement the path's scale information to obtain an aligned inspection path.

[0100] like Figure 6 As shown, this step involves applying the previously calculated affine transformation matrix to all trajectory points along the inspection path, thereby converting them to the coordinate system of the planar drawing. Simultaneously, to ensure the true scale of the path is accurately reflected, the path length and other scale information need to be adjusted according to the drawing scale. In this way, the final generated inspection path not only accurately displays on the planar drawing but also reflects its relative distances and positional relationships in the real world, facilitating subsequent target location and construction management.

[0101] In one embodiment, step S133 described above may include steps S1331 to S1333.

[0102] S1331. Apply the affine transformation matrix to all trajectory points on the inspection path, and map the inspection path completely to the coordinate system of the planar drawing through the coordinate transformation formula;

[0103] S1332. Based on the transformed trajectory point coordinates, draw the inspection path curve on the planar drawing, and restore the true scale information of the path through the drawing scale to obtain the path-drawing mapping relationship with actual physical dimensions.

[0104] S1333. Mark the actual coordinates of three pairs of reference points on the plan drawing to further confirm the precise alignment between the path and the drawing, so as to obtain the aligned inspection path.

[0105] Specifically, after recovering the inspection trajectory from the inspection video, the path trajectory with the initial position as the origin is extracted based on the keyframe path information, and then projected onto a two-dimensional plane to record the two-dimensional coordinates (x, y, x) of each trajectory point. i y i Select three reference points P on the drawing. j =(X j Y j These reference points correspond to fixed landmarks or pre-set calibration boards in the real world, and their actual coordinates on the drawings are obtained through manual measurement or annotation on the drawings.

[0106] Next, an affine transformation model is established between the path coordinate system and the drawing coordinate system. Based on the corresponding coordinates of three pairs of reference points, the optimal affine transformation parameters are solved using the least squares method, thus obtaining the optimal solutions for the transformation matrix A and the offset B. The solved affine transformation parameters are applied to all trajectory points, and the coordinate transformation formula is used to completely map the inspection path to the drawing coordinate system, achieving precise alignment between the path and the drawing.

[0107] Based on the transformed trajectory point coordinates (X) i Y i The inspection path curve is drawn on a 2D drawing, and the actual scale information of the path is restored using the drawing scale to establish a path-drawing mapping relationship with actual physical dimensions. Specifically, the actual coordinates of these three pairs of reference points are manually marked on the drawing; then, based on the transformed trajectory point coordinates (X... i Y iThe inspection path curve is drawn on a 2D drawing, and the drawing scale is used again to ensure that the actual scale information of the path is accurately reflected, ultimately forming a path-drawing mapping relationship with actual physical dimensions. The purpose of this is to ensure that the inspection path is not only accurately displayed on the 2D drawing but also reflects its relative distance and position in the real world, facilitating subsequent target location and construction management. The entire process emphasizes how to ensure consistency and accuracy between the path and the drawing through precise calculations and conversions.

[0108] Select three pairs of reference points on the 2D drawing, establish an affine transformation relationship between the inspection path coordinate system and the drawing pixel coordinate system, and denote the trajectory coordinates of the three points that need to be repositioned as follows: The pixel coordinates on the drawing are Calculate the affine transformation matrix to map the inspection path to the drawing coordinate system and supplement the path's scale information. The formula is as follows: in, Represents the coordinate matrix of the inspection trajectory. Represents the pixel coordinate matrix of a panoramic image, a ij Let t represent the linear transformation matrix. i It is a translation vector matrix.

[0109] Specifically, supplementing the scale information of the path usually involves converting the inspection path obtained from the video into a representation with actual size units (such as meters and centimeters). This ensures that the position and size of the construction components are accurate and quantifiable in the final generated 3D topology network.

[0110] S140. Based on the aligned inspection path, combined with the improved YOLOv8-DB model and depth estimation algorithm, the construction components are detected and the pixel coordinates of the construction components on the planar drawing are determined, and a three-dimensional topology network of path-drawing-target is formed.

[0111] In this embodiment, the three-dimensional topology network refers to the integration of inspection path data, drawing coordinate system, and location information of target objects (i.e., construction components) identified by target detection algorithms to form a network structure that reflects the spatial and temporal relationships between these elements. This network not only supports the precise positioning of construction components on the 2D drawings but also enables construction progress tracking and quality verification.

[0112] In one embodiment, step S140 described above may include steps S141 to S145.

[0113] S141. The indoor construction video is analyzed using an improved YOLOv8-DB model to obtain the pixel coordinates and detection boxes of the construction components.

[0114] In this embodiment, as Figure 7 As shown, the pixel coordinates of the construction component refer to the position of the target component in the video frame, usually expressed in the form of (x, y), where x and y are the horizontal and vertical distances of the target center point relative to the upper left corner of the image, respectively.

[0115] A detection box is a rectangular area used to mark the location of a target component, containing the target's boundary information. It helps to further accurately locate the target component and provide its dimensional information.

[0116] The improved YOLOv8-DB model includes a backbone network that replaces standard convolutions with deformable convolutional modules, a neck network structure that introduces a weighted bidirectional feature pyramid structure and fuses multi-level features through learnable weight parameters, and a detection head with added dynamic convolutional layers. Furthermore, the improved YOLOv8-DB model employs a method of randomly adjusting the input image size for multi-scale training.

[0117] In this embodiment, the improved YOLOv8-DB model has been innovated and optimized in several aspects to adapt to the needs of target detection in complex indoor construction scenarios. The following is a detailed description of its main improvements:

[0118] Traditional convolution operations assume that each local region in the input image has the same structural pattern. However, in practical applications, especially for construction components with irregular shapes and large size variations, this assumption often fails to hold true.

[0119] Introducing deformable convolutions into the backbone network: Unlike standard convolutions, DCN (Deformable Convolution Network) can dynamically learn the spatial offset of sampling points, allowing the convolution kernel to adaptively adjust its shape and position based on the content of the input feature map. This greatly enhances the model's ability to recognize construction components of different shapes.

[0120] Through this mechanism, the model can capture the key features of the target object more flexibly, and can ensure high detection accuracy even if these objects have large differences in shape or size.

[0121] In object detection tasks, multi-scale information fusion is an important means to improve the performance of small object detection. Traditional methods such as Feature Pyramid Networks (FPNs) are effective, but they are not efficient in cross-scale feature fusion. The neck network adopts an improved weighted bidirectional feature pyramid structure: BiFPN (Bidirectional Feature Pyramid Network), which not only allows information to flow from high-level to low-level layers, but also supports information flow in the opposite direction, achieving more efficient multi-scale feature fusion.

[0122] The calculation formula is: Where γi is a learnable weight parameter, and ∈ is a local minimum to prevent the denominator from being zero.

[0123] By assigning a trainable weight to each connection, the model can automatically learn which connections are more important, further improving the feature fusion effect.

[0124] When faced with complex backgrounds and diverse construction scenarios, fixed convolution kernels may not be able to fully extract all the necessary features.

[0125] This embodiment adds a dynamic convolutional layer to the detection head, dynamically generating convolutional kernel parameters based on input features, instead of using a fixed convolutional kernel. This approach makes the model more flexible in processing different input data and better adaptable to various complex environmental conditions. It enhances robustness, making it particularly suitable for target detection tasks in construction environments affected by factors such as changing lighting and occlusion.

[0126] Construction components range in size from small parts to large machinery, making single-scale training insufficient for all scenarios. Therefore, this embodiment simulates targets of different scales during training by randomly changing the size of the input image, forcing the model to learn how to accurately locate and classify targets in various situations. This method helps the model achieve better generalization ability, ensuring its effective handling of objects of different sizes in practical applications.

[0127] In summary, through targeted improvements to the backbone network, neck network, and detection head, and combined with an effective multi-scale training strategy, the improved YOLOv8-DB model significantly enhances the detection accuracy and robustness of various targets in complex indoor construction scenarios, providing strong technical support for achieving efficient and accurate construction management.

[0128] S142. The absolute depth distance of the center point of the construction component is extracted using the ZoeDepth algorithm, and the azimuth angle of the construction component relative to the camera's principal optical axis is calculated based on the offset of the pixel coordinates relative to the image center and the perspective characteristics of the panoramic video.

[0129] In this embodiment, the azimuth angle refers to the directional angle of the construction component relative to the camera's principal optical axis, used to describe the specific direction of the target within the camera's field of view.

[0130] like Figures 8 to 9As shown, in indoor construction scenarios, to achieve precise positioning and navigation of components, it is necessary to accurately obtain the position and orientation of each construction component relative to the camera's principal optical axis. This involves recovering depth information from monocular video input and calculating the azimuth angle of the target object relative to the camera based on the panoramic video's viewpoint characteristics.

[0131] In one embodiment, step S142 described above may include steps S1421 to S1422.

[0132] S1421. Generate a pixel-by-pixel depth map of the indoor construction video using the ZoeDepth algorithm, and record the depth value of the corresponding detection box center point.

[0133] In this embodiment, a pixel-by-pixel depth map means that each pixel has a corresponding depth value, which represents the actual distance from the camera to the spatial point represented by that pixel.

[0134] Depth value refers to the distance measurement from a specific point to the camera, usually measured in meters.

[0135] ZoeDepth is an advanced monocular depth estimation technique that infers the actual distance from each pixel to the camera from a single-view image. This algorithm does not require stereo vision or multi-view image input, simplifying the operation process.

[0136] After the image is processed by the ZoeDepth algorithm, each pixel is assigned a corresponding depth value, which represents the actual distance between the spatial point represented by that pixel and the camera. For example, if the depth value of a pixel is 3 meters, it means that the spatial point corresponding to that pixel is 3 meters away from the camera.

[0137] For each bounding box, special attention is paid to the center point of the construction component inside it, and the depth value of that center point is recorded. These depth values ​​will be used for position calculations in subsequent steps.

[0138] Suppose there is a construction component within a detection frame, and the coordinates of its center point are (u c v c After processing using the ZoeDepth algorithm, the depth value d of the center point is obtained. raw .

[0139] S1422. Normalize the pixel coordinates of each view in the six-view panoramic video corresponding to the indoor construction video to the center of their respective image coordinate system. By calculating the horizontal offset and known parameters, convert the pixel coordinates into azimuth angles relative to the camera's principal optical axis to obtain the azimuth angle of the construction component relative to the camera's principal optical axis.

[0140] First, the pixel coordinates of each view in the six perspectives of the panoramic video are normalized to the center of their respective image coordinate system. This means that regardless of the resolution of the original image, all pixels will be redefined based on their position relative to the image center.

[0141] Next, based on the normalized coordinates, the horizontal offset Δu = u - u0 of the center point of the construction component detection box relative to the center of the image is calculated, where u is the actual pixel coordinate and u0 is the pixel coordinate of the image center.

[0142] Assuming that some basic parameters such as focal length f and field of view (FOV) are known, these parameters are crucial for converting pixel-level offsets into actual angle measurements.

[0143] Combining the depth and horizontal offset values ​​obtained above, the specific azimuth angle θ of the construction component relative to the camera's principal optical axis is calculated using an appropriate mathematical model (such as trigonometric functions). The formula is as follows: Where W is the image width, and θ represents the deflection angle of the target relative to the camera's principal optical axis (positive for right deflection and negative for left deflection).

[0144] Continuing with the detection box example above, assuming the image width W is 1920 pixels, the image center u0 is 960 pixels, and the detection box center point u is 1000 pixels, then the horizontal offset Δu is 40 pixels. Assuming the FOV is 90 degrees and the focal length f is known, the azimuth angle θ of the construction component relative to the camera's principal optical axis can be calculated: This indicates that the construction component is located approximately 1.875 degrees to the right of the camera's main optical axis.

[0145] Through steps S1421 and S1422, not only can the depth information of the construction component be acquired, but its azimuth angle relative to the camera's principal optical axis can also be accurately calculated. This method fully utilizes the advantages of modern computer vision technology, providing an efficient and high-precision way to locate components in indoor construction scenes. It not only improves work efficiency but also enhances the safety and accuracy of the construction process. Combined with other parts of the entire method, a complete solution from data acquisition and path matching to final target localization is achieved.

[0146] S143. The aligned inspection path is processed by affine transformation to determine the position of the camera corresponding to the current frame in the coordinate system of the planar drawing, so as to obtain the camera position.

[0147] In this embodiment, the camera position refers to the mapping of the two-dimensional coordinates of the camera on the inspection path to the corresponding position in the coordinate system of the planar drawing through affine transformation.

[0148] Affine transformation is a linear transformation between two-dimensional coordinate systems, which can include translation, rotation, scaling, and shearing operations. In this process, the affine transformation matrix is ​​first calculated using the three pairs of reference points selected in the previous steps (i.e., the points on the inspection trajectory and their corresponding points on the drawing).

[0149] Based on the calculated affine transformation matrix, the coordinates of each trajectory point on the inspection path are transformed from its original coordinate system to the drawing coordinate system. For the current frame, based on its timestamp information, the two-dimensional coordinates of the camera position on the inspection path corresponding to that frame are found, and these coordinates are mapped to the drawing coordinate system through an affine transformation, thereby obtaining the precise position of the camera in the drawing coordinate system.

[0150] Assuming the camera position corresponding to a certain frame has coordinates (xt, yt) in the inspection path coordinate system, the affine transformation formula is used: Where A is the affine transformation matrix, B is the offset vector, (x p y p () represents the camera's new coordinates in the drawing coordinate system. This allows us to determine the camera's precise position on the drawing.

[0151] S144. Using the polar coordinate to Cartesian coordinate conversion formula and the camera position, convert the absolute depth distance and azimuth of the center point of the construction component into pixel coordinates on the planar drawing.

[0152] In this embodiment, after determining the camera position, the next step is to calculate the position of the construction component relative to the camera. This involves converting the absolute depth distance (d) of the target component and its azimuth angle (θ) relative to the camera's principal optical axis into pixel coordinates in the drawing coordinate system.

[0153] Use the following formula to convert polar coordinates to Cartesian coordinates: in, Here, θ is the pixel coordinate of the camera (path point) on the plan view, θ is the previously estimated azimuth angle, and d is the depth distance from the construction component to the camera.

[0154] S145. Associate the pixel coordinates of all construction components on the plan drawing with the timestamps of the inspection path to construct a three-dimensional topology network of path-drawing-target.

[0155] The pixel coordinates of each detected construction component are associated with the timestamp of its corresponding inspection path. This establishes the relationship between the component's position and its changes over time, forming a dynamic three-dimensional topological network.

[0156] This network not only contains the static positions of each construction component on the blueprints, but also records how these positions change and are updated over time. It allows users to track changes to components, such as whether they were installed or moved as planned.

[0157] Example: Suppose three construction components are detected at times t1, t2, and t3, with pixel coordinates P1, P2, and P3 respectively. By associating these coordinates with their respective timestamps, a dataset containing the time dimension can be generated. For example, a component might initially be located at P1 and move to P2 after a period of time. This information can help managers understand the actual progress of the construction site and ensure the project is completed on time.

[0158] Through steps S143 to S145, the coordinate transformation of the inspection path, the positioning of construction components, and the dynamic monitoring of the entire construction scene were achieved. This method not only improves the accuracy and efficiency of component positioning in indoor construction scenes but also provides strong support for subsequent quality inspection and progress management. The entire process demonstrates the application potential of modern computer vision technology and data analysis methods in intelligent building construction.

[0159] In this embodiment, an improved YOLOv8-DB model is used to analyze each frame of the inspection video to identify the category of the target component, the coordinates of the detection box, and its corresponding timestamp. Simultaneously, the ZoeDepth depth estimation algorithm is invoked to generate a depth map for each pixel based on the monocular image and record the depth value corresponding to the center point (uc, vc) of the target detection box. To ensure data consistency, the detection box and the depth map are strictly aligned using timestamps.

[0160] Based on the six-view layout of the panoramic video (each view has a horizontal field of view of 90°), we normalize the target pixel coordinates (u, v) to the center of the image coordinate system (u0, v0) and calculate the horizontal offset Δu = u - u0. Then, based on the focal length f and the field of view range FOV = 90°, we proportionally map this offset to the azimuth angle θ = Δu / (W / 2) × FOV / 2, where W represents the image width and θ represents the deflection angle of the target relative to the camera's principal optical axis (right deflection is considered positive, and left deflection is considered negative).

[0161] The depth value draw of the center point (uc, vc) of the target detection box is extracted from the depth map generated by ZoeDepth, and scale calibration is performed in combination with the camera intrinsic parameter matrix: d = draw × s, where s is a depth scaling factor determined in advance by the calibration board, which is used to convert the relative depth into absolute physical distance (unit: meters).

[0162] Based on the inspection path data processed by affine transformation, the pixel coordinates of the current frame camera position in the drawing coordinate system are obtained. Combined with the target azimuth angle and calibrated depth information, the absolute position of the detected target is calculated through polar coordinate-Cartesian coordinate transformation, thereby constructing a topology network of "inspection path-drawing-target object". The positions of all targets are finally converted into actual spatial coordinates according to the scale of the construction drawings, supporting construction progress tracking and quality inspection.

[0163] S150, Output the three-dimensional topology network.

[0164] This step aims to integrate all the previously processed and analyzed data, including inspection paths, target objects, and their coordinates in actual space, into a three-dimensional topology network model and output it. This three-dimensional model can intuitively display the spatial relationship between the inspection path and the target object, supporting the tracking of construction progress and quality inspection.

[0165] The aforementioned method for locating components in indoor construction scenes involves acquiring indoor construction video and extracting two-dimensional planar coordinates with the initial position as the origin to generate a planar drawing. Then, an affine transformation matrix is ​​used to accurately map the inspection path onto this planar drawing, and path scale information is supplemented to obtain an aligned inspection path. Combining an improved YOLOv8-DB model and depth estimation algorithm, construction components are detected based on the aligned inspection path, and their pixel coordinates on the planar drawing are determined, thereby constructing a three-dimensional topological network of path-drawing-target, and finally outputting this network model. This method achieves accurate and scale-informed component location in indoor construction scenes, greatly improving the intelligence level and work efficiency of construction management. It is applicable to indoor construction projects of various sizes and effectively solves the problems of inaccurate component location and low efficiency in existing technologies.

[0166] Figure 10 This is a schematic block diagram of an indoor construction scene component positioning device 300 provided in an embodiment of the present invention. Figure 10 As shown, corresponding to the above-described method for locating components in an indoor construction scene, the present invention also provides an indoor construction scene component positioning device 300. This indoor construction scene component positioning device 300 includes a unit for performing the above-described method for locating components in an indoor construction scene, and the device can be configured in a server. Specifically, please refer to... Figure 10 The indoor construction scene component positioning device 300 includes a video acquisition unit 301, a plan drawing acquisition unit 302, an alignment unit 303, a topology network generation unit 304, and an output unit 305.

[0167] The video acquisition unit 301 is used to acquire indoor construction video; the plan drawing acquisition unit 302 is used to determine the inspection path based on the indoor construction video and extract two-dimensional planar coordinates with the initial position as the origin from the indoor construction video to obtain a plan drawing; the alignment unit 303 is used to determine the affine transformation matrix based on the plan drawing, and use the affine transformation matrix to map the inspection path onto the plan drawing, and supplement the scale information of the path to obtain an aligned inspection path; the topology network generation unit 304 is used to detect construction components and determine the pixel coordinates of the construction components on the plan drawing based on the aligned inspection path combined with an improved YOLOv8-DB model and a depth estimation algorithm, and form a three-dimensional topology network of path-drawing-target; the output unit 305 is used to output the three-dimensional topology network.

[0168] In one embodiment, the plan drawing acquisition unit 302 includes:

[0169] The path determination subunit is used to process the indoor construction video data using the OpenVSLAM algorithm, extract key frames and corresponding pose information to obtain the inspection path; the projection subunit is used to project the inspection path onto a two-dimensional plane and record the two-dimensional coordinates of each trajectory point to obtain a planar drawing.

[0170] In one embodiment, the alignment unit 303 includes:

[0171] The reference point selection subunit is used to select three reference points on the plan view and determine the position coordinates of the three reference points in the inspection path and the position coordinates on the plan view to obtain path coordinates and plan coordinates; the matrix determination subunit is used to determine the affine transformation matrix based on the path coordinates and plan coordinates; the mapping supplement subunit is used to map all trajectory points on the inspection path to the plan view using the affine transformation matrix and supplement the scale information of the path to obtain an aligned inspection path.

[0172] In one embodiment, the matrix determination sub-unit is used to create an affine transformation relationship between the path coordinate system and the drawing coordinate system, and to calculate the optimal affine transformation parameters using the path coordinates and drawing coordinates in combination with the affine transformation matrix, thereby obtaining the best estimate of the transformation matrix and offset to determine the affine transformation matrix.

[0173] In one embodiment, the mapping supplement subunit includes:

[0174] The transformation module applies the affine transformation matrix to all trajectory points on the inspection path, mapping the inspection path completely to the coordinate system of the planar drawing using coordinate transformation formulas. The drawing module draws the inspection path curve on the planar drawing based on the transformed trajectory point coordinates, and restores the true scale information of the path using the drawing scale to obtain a path-drawing mapping relationship with actual physical dimensions. The alignment module marks the actual coordinates of three pairs of reference points on the planar drawing to further confirm the precise alignment between the path and the drawing, thus obtaining an aligned inspection path.

[0175] In one embodiment, the topology network generation unit 304 includes:

[0176] The analysis subunit analyzes the indoor construction video using an improved YOLOv8-DB model to obtain the pixel coordinates and detection boxes of the construction components. The azimuth calculation subunit uses the ZoeDepth algorithm to extract the absolute depth distance of the center point of the construction component and, based on the offset of the pixel coordinates relative to the image center and the perspective characteristics of the panoramic video, calculates the azimuth angle of the construction component relative to the camera's principal optical axis. The camera position determination subunit processes the aligned inspection path through affine transformation to determine the position of the camera corresponding to the current frame in the coordinate system of the planar drawing, thus obtaining the camera position. The transformation subunit applies the polar coordinate to Cartesian coordinate conversion formula and combines the camera position to convert the absolute depth distance and azimuth angle of the center point of the construction component into pixel coordinates on the planar drawing. The association subunit associates the pixel coordinates of all construction components on the planar drawing with the timestamps of the inspection path, constructing a three-dimensional topology network of path-drawing-target.

[0177] In one embodiment, the azimuth calculation subunit includes:

[0178] The recording module is used to generate a pixel-by-pixel depth map of the indoor construction video using the ZoeDepth algorithm, and record the depth value of the corresponding detection box center point; the calculation module is used to normalize the pixel coordinates of each view in the six-view panoramic video corresponding to the indoor construction video to the center of their respective image coordinate system, and convert the pixel coordinates into azimuth angles relative to the camera's principal optical axis by calculating the horizontal offset and known parameters, so as to obtain the azimuth angle of the construction component relative to the camera's principal optical axis.

[0179] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned indoor construction scene component positioning device 300 and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.

[0180] The aforementioned indoor construction scene component positioning device 300 can be implemented as a computer program, which can, for example... Figure 11 It runs on the computer device shown.

[0181] Please see Figure 11 , Figure 11 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.

[0182] See Figure 11 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.

[0183] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform a method for locating components in an indoor construction scene.

[0184] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.

[0185] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a method for locating components in an indoor construction scene.

[0186] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 11 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0187] The processor 502 is used to run a computer program 5032 stored in the memory to perform the following steps:

[0188] Acquire indoor construction video; determine inspection path based on indoor construction video, and extract two-dimensional plane coordinates from indoor construction video with the initial position as the origin to obtain a planar drawing; determine affine transformation matrix based on the planar drawing, and use the affine transformation matrix to map the inspection path onto the planar drawing, and supplement the scale information of the path to obtain an aligned inspection path; based on the aligned inspection path, combine an improved YOLOv8-DB model and depth estimation algorithm to detect construction components and determine the pixel coordinates of the construction components on the planar drawing, and form a three-dimensional topology network of path-drawing-target; output the three-dimensional topology network.

[0189] In one embodiment, when the processor 502 implements the steps of determining the inspection path based on the indoor construction video and extracting two-dimensional plane coordinates from the indoor construction video with the initial position as the origin to obtain a plan drawing, the specific implementation steps are as follows:

[0190] The indoor construction video data is processed using the OpenVSLAM algorithm to extract keyframes and their corresponding pose information to obtain the inspection path; the inspection path is then projected onto a two-dimensional plane, and the two-dimensional coordinates of each trajectory point are recorded to obtain a planar drawing.

[0191] In one embodiment, when the processor 502 implements the steps of determining the affine transformation matrix based on the planar drawing, mapping the inspection path onto the planar drawing using the affine transformation matrix, and supplementing the scale information of the path to obtain an aligned inspection path, the processor specifically implements the following steps:

[0192] Three reference points are selected on the plan view, and the position coordinates of the three reference points in the inspection path and the position coordinates on the plan view are determined to obtain path coordinates and plan coordinates; an affine transformation matrix is ​​determined based on the path coordinates and plan coordinates; the affine transformation matrix is ​​used to map all trajectory points on the inspection path onto the plan view, and the scale information of the path is supplemented to obtain an aligned inspection path.

[0193] In one embodiment, when the processor 502 implements the step of determining the affine transformation matrix based on the path coordinates and the drawing coordinates, it specifically implements the following steps:

[0194] Create an affine transformation relationship between the path coordinate system and the drawing coordinate system, and use the path coordinates and drawing coordinates in combination with the affine transformation matrix to calculate the optimal affine transformation parameters, obtain the best estimate of the transformation matrix and offset, and thus determine the affine transformation matrix.

[0195] In one embodiment, when the processor 502 implements the step of mapping all trajectory points on the inspection path onto the planar drawing using the affine transformation matrix and supplementing the path's scale information to obtain an aligned inspection path, the specific implementation steps are as follows:

[0196] The affine transformation matrix is ​​applied to all trajectory points on the inspection path, and the inspection path is completely mapped to the coordinate system of the planar drawing using the coordinate transformation formula. Based on the transformed trajectory point coordinates, the inspection path curve is drawn on the planar drawing, and the true scale information of the path is restored using the drawing scale to obtain a path-drawing mapping relationship with actual physical dimensions. The actual coordinates of three pairs of reference points are marked on the planar drawing to further confirm the precise alignment between the path and the drawing, thus obtaining an aligned inspection path.

[0197] In one embodiment, when the processor 502 implements the steps of detecting construction components and determining the pixel coordinates of the construction components on the planar drawing based on the aligned inspection path combined with the improved YOLOv8-DB model and depth estimation algorithm, and forming a three-dimensional topology network of path-drawing-target, the specific implementation is as follows:

[0198] An improved YOLOv8-DB model was used to analyze the indoor construction video to obtain the pixel coordinates and detection boxes of the construction components. The ZoeDepth algorithm was used to extract the absolute depth distance of the center point of the construction component, and based on the offset of the pixel coordinates relative to the image center, combined with the viewpoint characteristics of the panoramic video, the azimuth angle of the construction component relative to the principal optical axis of the camera was calculated. The aligned inspection path was processed by affine transformation to determine the position of the camera corresponding to the current frame in the coordinate system of the planar drawing, so as to obtain the camera position. The polar coordinate to Cartesian coordinate conversion formula was applied to convert the absolute depth distance and azimuth angle of the center point of the construction component into pixel coordinates on the planar drawing in combination with the camera position. The pixel coordinates of all construction components on the planar drawing were associated with the timestamps of the inspection path to construct a three-dimensional topology network of path-drawing-target.

[0199] The improved YOLOv8-DB model includes a backbone network that replaces standard convolution with deformable convolutional modules, a neck network structure that introduces a weighted bidirectional feature pyramid structure and fuses multi-level features through learnable weight parameters, and a detection head with added dynamic convolutional layers. Furthermore, the improved YOLOv8-DB model employs a method of randomly adjusting the input image size for multi-scale training.

[0200] In one embodiment, when the processor 502 implements the step of extracting the absolute depth distance of the center point of the construction component using the ZoeDepth algorithm, and calculating the azimuth angle of the construction component relative to the principal optical axis of the camera based on the offset of the pixel coordinates relative to the image center and the viewpoint characteristics of the panoramic video, the specific implementation steps are as follows:

[0201] The ZoeDepth algorithm is used to generate a pixel-by-pixel depth map of the indoor construction video, and the depth value of the corresponding detection box center point is recorded. The pixel coordinates of each view in the six-view panoramic video corresponding to the indoor construction video are normalized to the center of their respective image coordinate system. By calculating the horizontal offset and known parameters, the pixel coordinates are converted into azimuth angles relative to the camera's principal optical axis, so as to obtain the azimuth angle of the construction component relative to the camera's principal optical axis.

[0202] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0203] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0204] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform the following steps:

[0205] Acquire indoor construction video; determine inspection path based on indoor construction video, and extract two-dimensional plane coordinates from indoor construction video with the initial position as the origin to obtain a planar drawing; determine affine transformation matrix based on the planar drawing, and use the affine transformation matrix to map the inspection path onto the planar drawing, and supplement the scale information of the path to obtain an aligned inspection path; based on the aligned inspection path, combine an improved YOLOv8-DB model and depth estimation algorithm to detect construction components and determine the pixel coordinates of the construction components on the planar drawing, and form a three-dimensional topology network of path-drawing-target; output the three-dimensional topology network.

[0206] In one embodiment, when the processor executes the computer program to determine the inspection path based on the indoor construction video and extract two-dimensional plane coordinates from the indoor construction video with the initial position as the origin to obtain a plan drawing, the processor specifically implements the following steps:

[0207] The indoor construction video data is processed using the OpenVSLAM algorithm to extract keyframes and their corresponding pose information to obtain the inspection path; the inspection path is then projected onto a two-dimensional plane, and the two-dimensional coordinates of each trajectory point are recorded to obtain a planar drawing.

[0208] In one embodiment, when the processor executes the computer program to implement the steps of determining the affine transformation matrix based on the planar drawing, mapping the inspection path onto the planar drawing using the affine transformation matrix, and supplementing the path's scale information to obtain an aligned inspection path, the processor specifically implements the following steps:

[0209] Three reference points are selected on the plan view, and the position coordinates of the three reference points in the inspection path and the position coordinates on the plan view are determined to obtain path coordinates and plan coordinates; an affine transformation matrix is ​​determined based on the path coordinates and plan coordinates; the affine transformation matrix is ​​used to map all trajectory points on the inspection path onto the plan view, and the scale information of the path is supplemented to obtain an aligned inspection path.

[0210] In one embodiment, when the processor executes the computer program to implement the step of determining the affine transformation matrix based on the path coordinates and the drawing coordinates, it specifically implements the following steps:

[0211] Create an affine transformation relationship between the path coordinate system and the drawing coordinate system, and use the path coordinates and drawing coordinates in combination with the affine transformation matrix to calculate the optimal affine transformation parameters, obtain the best estimate of the transformation matrix and offset, and thus determine the affine transformation matrix.

[0212] In one embodiment, when the processor executes the computer program to implement the step of mapping all trajectory points on the inspection path onto the planar drawing using the affine transformation matrix and supplementing the path's scale information to obtain an aligned inspection path, the specific implementation steps are as follows:

[0213] The affine transformation matrix is ​​applied to all trajectory points on the inspection path, and the inspection path is completely mapped to the coordinate system of the planar drawing using the coordinate transformation formula. Based on the transformed trajectory point coordinates, the inspection path curve is drawn on the planar drawing, and the true scale information of the path is restored using the drawing scale to obtain a path-drawing mapping relationship with actual physical dimensions. The actual coordinates of three pairs of reference points are marked on the planar drawing to further confirm the precise alignment between the path and the drawing, thus obtaining an aligned inspection path.

[0214] In one embodiment, when the processor executes the computer program to implement the steps of detecting construction components and determining the pixel coordinates of the construction components on the planar drawing based on the aligned inspection path combined with the improved YOLOv8-DB model and depth estimation algorithm, and forming a three-dimensional topology network of path-drawing-target, the specific implementation is as follows:

[0215] An improved YOLOv8-DB model was used to analyze the indoor construction video to obtain the pixel coordinates and detection boxes of the construction components. The ZoeDepth algorithm was used to extract the absolute depth distance of the center point of the construction component, and based on the offset of the pixel coordinates relative to the image center, combined with the viewpoint characteristics of the panoramic video, the azimuth angle of the construction component relative to the principal optical axis of the camera was calculated. The aligned inspection path was processed by affine transformation to determine the position of the camera corresponding to the current frame in the coordinate system of the planar drawing, so as to obtain the camera position. The polar coordinate to Cartesian coordinate conversion formula was applied to convert the absolute depth distance and azimuth angle of the center point of the construction component into pixel coordinates on the planar drawing in combination with the camera position. The pixel coordinates of all construction components on the planar drawing were associated with the timestamps of the inspection path to construct a three-dimensional topology network of path-drawing-target.

[0216] The improved YOLOv8-DB model includes a backbone network that replaces standard convolution with deformable convolutional modules, a neck network structure that introduces a weighted bidirectional feature pyramid structure and fuses multi-level features through learnable weight parameters, and a detection head with added dynamic convolutional layers. Furthermore, the improved YOLOv8-DB model employs a method of randomly adjusting the input image size for multi-scale training.

[0217] In one embodiment, when the processor executes the computer program to extract the absolute depth distance of the center point of the construction component using the ZoeDepth algorithm, and calculates the azimuth angle of the construction component relative to the principal optical axis of the camera based on the offset of the pixel coordinates relative to the image center and the viewpoint characteristics of the panoramic video, the specific steps are as follows:

[0218] The ZoeDepth algorithm is used to generate a pixel-by-pixel depth map of the indoor construction video, and the depth value of the corresponding detection box center point is recorded. The pixel coordinates of each view in the six-view panoramic video corresponding to the indoor construction video are normalized to the center of their respective image coordinate system. By calculating the horizontal offset and known parameters, the pixel coordinates are converted into azimuth angles relative to the camera's principal optical axis, so as to obtain the azimuth angle of the construction component relative to the camera's principal optical axis.

[0219] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.

[0220] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0221] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0222] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0223] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0224] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for positioning components in an indoor construction scene, characterized in that, include: Obtain indoor construction videos; The inspection route is determined based on the indoor construction video, and two-dimensional plane coordinates with the initial position as the origin are extracted from the indoor construction video to obtain a plan drawing. Based on the planar drawing, an affine transformation matrix is ​​determined, and the inspection path is mapped onto the planar drawing using the affine transformation matrix. The scale information of the path is then added to obtain an aligned inspection path. Based on the aligned inspection path, combined with the improved YOLOv8-DB model and depth estimation algorithm, construction components are detected and their pixel coordinates on the planar drawing are determined, forming a three-dimensional topology network of path-drawing-target. Output the three-dimensional topology network.

2. The method for positioning components in an indoor construction scene according to claim 1, characterized in that, The process of determining the inspection path based on the indoor construction video and extracting two-dimensional plane coordinates from the indoor construction video with the initial position as the origin to obtain a plan drawing includes: The indoor construction video data is processed using the OpenVSLAM algorithm to extract key frames and corresponding pose information in order to obtain the inspection path. The inspection path is projected onto a two-dimensional plane, and the two-dimensional coordinates of each trajectory point are recorded to obtain a planar drawing.

3. The method for positioning components in an indoor construction scene according to claim 1, characterized in that, The process of determining the affine transformation matrix based on the planar drawing, mapping the inspection path onto the planar drawing using the affine transformation matrix, and supplementing the path's scale information to obtain an aligned inspection path includes: Select three reference points on the plan view, and determine the position coordinates of the three reference points in the inspection path and the position coordinates on the plan view to obtain the path coordinates and the plan coordinates; Determine the affine transformation matrix based on the path coordinates and drawing coordinates; The affine transformation matrix is ​​used to map all trajectory points on the inspection path onto the planar drawing, and the scale information of the path is supplemented to obtain an aligned inspection path.

4. The method for positioning components in an indoor construction scene according to claim 3, characterized in that, The step of determining the affine transformation matrix based on the path coordinates and drawing coordinates includes: Create an affine transformation relationship between the path coordinate system and the drawing coordinate system, and use the path coordinates and drawing coordinates in combination with the affine transformation matrix to calculate the optimal affine transformation parameters, obtain the best estimate of the transformation matrix and offset, and thus determine the affine transformation matrix.

5. The method for positioning components in an indoor construction scene according to claim 3, characterized in that, The process of mapping all trajectory points on the inspection path onto the planar drawing using the affine transformation matrix and supplementing the path's scale information to obtain an aligned inspection path includes: The affine transformation matrix is ​​applied to all trajectory points on the inspection path, and the inspection path is completely mapped to the coordinate system of the planar drawing through the coordinate transformation formula. Based on the transformed trajectory point coordinates, the inspection path curve is drawn on the plan view, and the true scale information of the path is restored by the map scale to obtain the path-map mapping relationship with actual physical dimensions. Mark the actual coordinates of three pairs of reference points on the plan drawing to further confirm the precise alignment between the path and the drawing, so as to obtain the aligned inspection path.

6. The method for positioning components in an indoor construction scene according to claim 1, characterized in that, The process, based on the aligned inspection path and combined with an improved YOLOv8-DB model and depth estimation algorithm, detects construction components and determines their pixel coordinates on the planar drawing, forming a 3D topology network of path-drawing-target, including: An improved YOLOv8-DB model was used to analyze the indoor construction video to obtain the pixel coordinates and detection boxes of the construction components. The absolute depth distance of the center point of the construction component is extracted using the ZoeDepth algorithm. Based on the offset of the pixel coordinates relative to the image center and combined with the perspective characteristics of the panoramic video, the azimuth angle of the construction component relative to the camera's principal optical axis is calculated. The aligned inspection path is processed by affine transformation to determine the position of the camera corresponding to the current frame in the coordinate system of the planar drawing, so as to obtain the camera position. By applying the polar coordinate to Cartesian coordinate conversion formula and combining it with the camera position, the absolute depth distance and azimuth of the center point of the construction component are converted into pixel coordinates on the planar drawing; By associating the pixel coordinates of all construction components on the planar drawing with the timestamps of the inspection path, a three-dimensional topology network of path-drawing-target is constructed.

7. The method for positioning components in an indoor construction scene according to claim 6, characterized in that, The improved YOLOv8-DB model includes a backbone network that replaces standard convolutions with deformable convolutional modules, a neck network structure that introduces a weighted bidirectional feature pyramid structure and fuses multi-level features through learnable weight parameters, and a detection head with added dynamic convolutional layers. Furthermore, the improved YOLOv8-DB model employs a method of randomly adjusting the input image size for multi-scale training.

8. The method for positioning components in an indoor construction scene according to claim 6, characterized in that, The method involves using the ZoeDepth algorithm to extract the absolute depth distance of the center point of the construction component, and based on the offset of the pixel coordinates relative to the image center, combined with the viewpoint characteristics of panoramic video, to calculate the azimuth angle of the construction component relative to the principal optical axis of the camera, including: The ZoeDepth algorithm is used to generate a pixel-by-pixel depth map of the indoor construction video, and the depth value of the corresponding detection box center point is recorded. The pixel coordinates of each view in the six-view panoramic video corresponding to the indoor construction video are normalized to the center of their respective image coordinate system. By calculating the horizontal offset and known parameters, the pixel coordinates are converted into azimuth angles relative to the camera's principal optical axis, so as to obtain the azimuth angle of the construction component relative to the camera's principal optical axis.

9. A component positioning device for indoor construction scenes, characterized in that, include: The video acquisition unit is used to acquire indoor construction videos. The plan drawing acquisition unit is used to determine the inspection path based on the indoor construction video and extract two-dimensional planar coordinates from the indoor construction video with the initial position as the origin to obtain a plan drawing. The alignment unit is used to determine the affine transformation matrix based on the planar drawing, and to map the inspection path onto the planar drawing using the affine transformation matrix, and to supplement the scale information of the path to obtain the aligned inspection path. The topology network generation unit is used to detect construction components and determine the pixel coordinates of the construction components on the planar drawing based on the aligned inspection path combined with the improved YOLOv8-DB model and depth estimation algorithm, and form a three-dimensional topology network of path-drawing-target. The output unit is used to output the three-dimensional topology network.

10. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 8.