Structural object butt joint monitoring method fusing computer vision and laser radar

By integrating computer vision and lidar, neural networks are used to perform image segmentation and attitude estimation, the large monitoring error and inconvenience in operation during docking of large structures are solved, and high-precision real-time motion monitoring is achieved, which is suitable for wind power and offshore oil and gas installation.

CN120374671APending Publication Date: 2025-07-25TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510225212.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art has problems such as large measurement error, inconvenient operation and high cost during docking of large structures, making it difficult to achieve accurate relative motion monitoring.

Method used

Using a method of fusion computer vision and lidar, the image data of the target structure is obtained through cameras and lidar, and the neural network is used to perform image segmentation and attitude estimation, so as to monitor the movement of the structure in real time.

Benefits of technology

Real-time motion monitoring of multiple structures is realized, monitoring accuracy and operating efficiency are improved, operation complexity and cost are reduced, and it is suitable for wind power installation and offshore oil and gas installation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374671A_ABST
    Figure CN120374671A_ABST
Patent Text Reader

Abstract

The invention discloses a structure docking monitoring method fusing computer vision and a laser radar. The method comprises the following steps: setting a camera and a laser radar for capturing a picture of a target structure, and obtaining an internal reference of the camera and a three-dimensional model of the target structure; a camera head of the camera and the laser radar are started, image data of the target structure are obtained, and the image data comprise a depth map and a color map; carrying out image segmentation based on the image data, and obtaining an initial state mask image corresponding to each target structure; estimating the initial position and attitude of each target structure by using a neural network based on the color map, the depth map and the initial state mask map of the first frame, the three-dimensional model of the target structure and the internal reference of the camera; and comparing the current input image with the target structure at the previous moment by using the neural network to obtain a motion tracking result. According to the invention, a plurality of structures can be tracked at the same time, and real-time motion monitoring of the whole engineering operation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of mechatronic control, and more specifically, to a method for monitoring the docking of structures by integrating computer vision and lidar. Background Art

[0002] During the docking process of large structures, due to the complex aerial environment and the limitations of lifting equipment, the structures often swing during the docking process, and it is necessary to accurately monitor the relative movement between the docking components for on-site decision-making.

[0003] In the existing methods for monitoring the docking of structures, the common solutions include using total stations or the integration of inertial navigation and satellite navigation. After measuring the poses of two objects respectively, the relative position is calculated. This indirect measurement method has technical defects such as complex calibration between multiple devices, inconvenient operation, and large measurement errors, significantly increasing the operation cost and risk.

[0004] In summary, it is necessary to further improve the method for monitoring the docking of structures to improve the monitoring accuracy and operation efficiency and enhance the convenience of operation. Summary of the Invention

[0005] The object of the present invention is to overcome the defects of the above-mentioned prior art and provide a method for monitoring the docking of structures by integrating computer vision and lidar. The method includes the following steps:

[0006] Set up a camera and a lidar to capture the images of the target structure, and obtain the internal parameters of the camera and the three-dimensional model of the target structure;

[0007] During the docking operation, start the camera and the lidar to obtain the image data of the target structure, which includes a depth map and a color map;

[0008] Based on the image data, perform image segmentation to obtain the initial state mask map corresponding to each target structure;

[0009] Based on the color map, the depth map, and the initial state mask map of the first frame, as well as the three-dimensional model of the target structure and the internal parameters of the camera, use a neural network to estimate the initial position and pose of each target structure;

[0010] Use the neural network to compare the current input image with the target structure at the previous moment and update its position and pose to obtain the motion tracking result.

[0011] Compared with the prior art, the advantages of the present invention are as follows: A non-contact monitoring method for the relative position during the docking of structures integrating an optical lens and a lidar is proposed based on computer vision theory. By taking the color images, depth maps, mask maps obtained by a camera and a lidar, as well as the three-dimensional model of the target object and the camera internal parameters as inputs, a neural network is used to estimate the pose of the initial state of the target object. The neural network will calculate the position parameters of the current state based on the estimation result of the previous state and the input images of the current state. With the present invention, multiple structures can be tracked simultaneously, thereby realizing real-time motion monitoring of the entire engineering operation, and improving the monitoring accuracy and operation efficiency of the docking operation of large structures. It can serve scenarios such as wind power installation and offshore oil and gas installation.

[0012] Other features and advantages of the present invention will become clear from the following detailed description of exemplary embodiments of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings incorporated in and constituting a part of this specification illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.

[0014] Figure 1 is a flowchart of a method for monitoring the docking of structures integrating computer vision and lidar according to an embodiment of the present invention;

[0015] Figure 2 is a schematic diagram of the process of a method for monitoring the docking of structures integrating computer vision and lidar according to an embodiment of the present invention;

[0016] Figure 3 is a schematic diagram of the scoring result of the initial pose estimation according to an embodiment of the present invention;

[0017] Figure 4 is a schematic diagram showing the simulation result of the motion monitoring of an experimental structure model according to an embodiment of the present invention;

[0018] Figure 5 is a schematic diagram of the digital display of the motion monitoring result of an experimental structure model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0019] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present invention.

[0020] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present invention or its application or use.

[0021] Techniques, methods, and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered as part of the specification.

[0022] In all examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.

[0023] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, further discussion thereof in subsequent figures is not required.

[0024] Generally speaking, the structure docking monitoring method based on computer vision and lidar provided by the present invention includes: before the structure docking, installing a camera and a lidar at appropriate positions to ensure that both can effectively capture the target structure, and calibrating the images of both according to the internal and external parameters of the camera and the lidar; when the structure is lifted and before the docking work starts, after ensuring that the camera can capture all target structures at the same time, starting the video stream to obtain the color and depth image data of the structure in the current state; using the object image segmentation method to respectively cut out the mask map of each structure according to the color image data; inputting the color map, depth map, mask map, 3D model, and camera internal parameter, and using a neural network to respectively realize the estimation of the current position and attitude of each structure; according to the attitude estimation at the previous moment and the current image input, using a neural network to calculate the current attitude, so as to realize the real-time motion monitoring of the structure during the docking process.

[0025] See Figure 1 As shown, the provided structure docking monitoring method integrating computer vision and lidar includes the following steps:

[0026] Step S1, fixing the camera and lidar for capturing the structure image and calibrating the image.

[0027] Before the structure is installed, install the camera and lidar at appropriate positions, and ensure that during the tracking process, the images of the camera and lidar can effectively capture the structure. It is not necessary to require that all structures must appear completely in the image, but it is necessary to ensure that most of the structures of each structure are included in the image. The camera and lidar should be installed as horizontally as possible with the target object or facing away from the sun to avoid image interference caused by direct sunlight. Before the docking work starts, obtain the parameters of the camera and the 3D model of the target structure. If there is no ready-made model, a neural network can be used to reconstruct the 3D model of the target structure. At the same time, ensure that the camera and lidar have completed image calibration according to their internal and external parameters.

[0028] Step S2, start the camera and lidar to capture the image data of the target structure, including depth maps and color maps.

[0029] When the docking work starts, start the camera and lidar of the camera to capture the image data of the structure, including depth maps, color maps, etc.

[0030] Step S3, perform image segmentation on the structure and input it into the neural network to estimate the initial positions and postures of each structure respectively.

[0031] For example, use the object image segmentation method, take the first-frame color image as a reference, and segment the target structure from the background. For the docking monitoring of multiple structures, it is necessary to perform image segmentation on each structure separately to obtain the initial state mask map of each structure. This mask map will be used for the estimation of the initial posture and ensure that the output mask map is a single-channel normalized grayscale map, where the pixel value of the target structure is 1 and the pixel value of the background is 0:

[0032]

[0033] where I is the pixel value of the original grayscale image (range [0, 255]). I 归一化 is the normalized pixel value, with a range of [0, 1].

[0034] Step S4, take the color map, depth map, mask map, as well as the 3D model of the target structure and the camera internal parameters as inputs, and use the neural network to estimate the initial positions and postures of each structure respectively.

[0035] In one embodiment, according to the color image, depth map, mask map of the first frame, as well as the 3D model of the target structure and the camera internal parameters as inputs, use the neural network to estimate the initial positions and postures of each structure respectively, see Figure 2 shown.

[0036] The neural network can adopt various types, such as convolutional neural network, Tranformer neural network, etc. The neural network renders multiple hypothesized postures based on the existing 3D model and compares them with the actual results captured by the camera. The comparison results of multiple hypothesized postures are evaluated through a scoring mechanism to select the optimal posture, see Figure 3 shown.

[0037] Step S5, the neural network calculates the current position according to the position information of the structure at the previous moment to realize the motion monitoring of the structure, and then fuses the position information of each structure to realize the real-time motion monitoring of the entire engineering operation.

[0038] Use a neural network to compare the input image with the structure at the previous moment, and update its position and attitude. When tracking multiple objects, it is necessary to separately perform individual motion tracking on each structure first, and then fuse the tracking results of each structure to achieve real-time motion monitoring during the installation and docking process of multiple structures. This method is used to verify the experimental model, Figure 4 and shows the results of real-time monitoring of the experimental model.

[0039] The attitude tracking result of the structure digitizes the motion monitoring result of the structure. This is extremely important for other related projects such as lifting control, engineering operation modeling, and digital twin ecosystem. The tracking result is a homogeneous coordinate transformation matrix, which represents the object's in the camera coordinate system:

[0040]

[0041] where R is a 3×3 rotation matrix, representing the rotation of the object around the origin. Each element r of the matrix ij represents the cosine relationship of the rotated coordinate axis directions. T is a 3×1 column vector, representing the translation of the object. t x , t y , t z are the translation amounts of the object along the x, y, and z axes in three-dimensional space. Figure 5 shows the calculated motion state of the object using this matrix.

[0042] In summary, compared with the prior art, the present invention has the following advantages:

[0043] 1) The present invention can implement a non-contact measurement method for large equipment in engineering operations. This method combines computer vision and lidar technologies and is used to monitor the docking work of structures during engineering construction or equipment installation. In addition, this method can also effectively solve the problem of motion monitoring of large equipment in engineering operations by only performing motion tracking on partial feature areas of large equipment (such as regular structures like pile legs, bases, etc.) and calculating the specific position and attitude of the entire equipment using inverse dynamics.

[0044] 2) The present invention expands the two-dimensional data of the image into three-dimensional position data containing depth information through lidar, which can significantly improve the measurement accuracy compared with traditional visual measurement. In addition, digital processing of the motion data information of the structure is realized, and these data can be applied to other projects such as lifting control, engineering operation modeling, and engineering digital twin ecosystem.

[0045] 3) When the present invention monitors a new object or an object without a three-dimensional model, by inputting color maps, depth maps, and mask maps of different angles of the object, it can use a neural network to reconstruct the three-dimensional model of the new object, thereby monitoring the motion of the new object.

[0046] 4) The motion monitoring method provided by the present invention employs a neural network trained with a large number of data sets. Therefore, this method has a certain degree of robustness against local exposure, occlusion of objects, and missing images caused by harsh environments. In engineering, in the face of large and complex operations, it can improve the stability and safety of motion monitoring to a certain extent.

[0047] In summary, the monitoring method provided by the present invention obtains color images, depth maps, mask maps generated by object segmentation, 3D models, and camera internal parameters through cameras and lidar, and uses a neural network to estimate the pose of the initial state of the target structure. The neural network will combine the pose estimation result of the previous moment and the current image input to monitor the motion state of the target structure in real time. This method can monitor multiple structures simultaneously, so it can be widely used in the real-time motion monitoring of the docking work of structures in engineering operations. Compared with the traditional visual inspection method, the present invention significantly improves the monitoring accuracy. At the same time, compared with measurement technologies such as GPS + IMU, the present invention has the advantage of non-contact measurement, avoiding a large amount of installation, calibration, and disassembly costs. In addition, the digitized monitoring results can provide important data support for applications such as hoisting control, engineering operation modeling, and digital twin ecosystems, which has important practical significance.

[0048] The present invention may be a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0049] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0050] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0051] The computer program instructions for carrying out operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, Python, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present invention.

[0052] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0053] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0054] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0055] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction may include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are equivalent.

[0056] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.

Claims

1. A method for monitoring the docking of structures by integrating computer vision and lidar, comprising the following steps: Set up a camera and a lidar to capture images of the target structure, and obtain the internal parameters of the camera and the three-dimensional model of the target structure; During the docking process, start the camera and the lidar to obtain image data of the target structure, which includes a depth map and a color map; Based on the image data, perform image segmentation to obtain an initial state mask map corresponding to each target structure; Based on the color map, the depth map, and the initial state mask map of the first frame, as well as the three-dimensional model of the target structure and the internal parameters of the camera, use a neural network to estimate the initial position and pose of each target structure; Use the neural network to compare the current input image with the target structure at the previous moment, and update its position and pose to obtain a motion tracking result.

2. The method according to claim 1, wherein The initial state mask map is a single-channel normalized grayscale map, which contains the pixel values of the target structure and the background.

3. The method according to claim 1, characterized in that Performing image segmentation based on the image data includes: using an object image segmentation method, with the first-frame color map as a reference, to segment the target structure from the background.

4. The method according to claim 1, wherein The motion tracking result of the target structure is a homogeneous coordinate transformation matrix, expressed as: Among them, R is a 3×3 rotation matrix representing the rotation of an object around the origin, and the element r ij represents the cosine relationship of the coordinate axis directions after rotation. T is a 3×1 column vector representing the translation of the object, and t x , t y , t z are the translation amounts of the object along the x, y, and z axes in three-dimensional space.

5. The method according to claim 1, characterized in that, The three-dimensional model of the target structure is an existing three-dimensional model or a three-dimensional model reconstructed using a neural network.

6. The method according to claim 1, wherein The neural network is a Transformer neural network.

7. The method according to claim 2, wherein The single-channel normalized grayscale map of the mask map is expressed as: where I is the pixel value of the original grayscale image, with a range of [0, 255], and I 归一化 is the pixel value after normalization, with a range of [0, 1]. The pixel value of the target structure is 1, and the pixel value of the background is 0.

8. The method according to claim 1, wherein The number of the target structures is one or more.

9. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

10. A computer device, comprising a memory and a processor, wherein a computer program capable of running on the processor is stored on the memory, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.