Remote three-dimensional dimensional measurement and field marking system and method
Patent Information
- Application Number
- CN202611350691.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-09-02
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]为解决现有技术存在的高压电缆附件施工质量检验高度依赖专家到场、二维影像无法直接提供可靠空间尺寸、远程专家选定位置难以在现场空间中准确呈现,以及自动图像匹配结果缺少有效性验证等缺陷,本申请一种提供基于增强现实设备多相机图像的远程三维尺寸测量及现场标示系统和方法
本申请对增强现实、计算机视觉、双目或多目视觉测量、远程协作及工业施工质量管控等领域方案进行融合,将远程交互选点、多相机几何测量、结果可信度校验以及增强现实空间回显相结合,通过本方案,基于头戴式增强现实设备的异构双相机同步图像和远程专家交互选点实现现场长度测量,远程专家能够在查看现场实时画面的同时选取待测组件的起点和终点,利用头戴式增强现实设备的多个已标定相机同步采集现场图像,由远程专家在图像中选取待测端点,自动恢复两个端点的三维坐标、计算三维直线距离尺寸,系统自动获得的具有实际尺度的三维测量结果可以避免仅凭二维影响估算尺寸,并将测量端点及测量结果在增强现实设备中进行空间标示,叠加显示在现场人员视野中。持续标示使现场人员能够准确理解测量位置并即时完成复核或整改,实现了对施工关键尺寸的即时远程测量和现场复核。通过同步多相机采集、已标定相机几何关系和双目或多目三角测量,作为辅助手段,配合强标要求的旁站检查、现场测量等规定,由现场人员配合完成关键尺寸远程核验。施工专家可同时远程支持多个施工现场,减少差旅和等待时间,降低问题回溯及返工成本,提高高压电缆附件施工质量管控的及时性和人效。本方案测量值与真实值的偏差不超过5%,特别适合于高压电缆附件施工过程中的关键尺寸预检、远程复核与多现场并发支持,作为监理平行检验与厂家技术服务人员现场指导的辅助手段。
Smart Images

Figure CN122835246A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of AR-assisted measurement technology, specifically a remote three-dimensional dimension measurement and on-site marking system and method based on multi-camera images from augmented reality devices. Background Technology
[0002] During the construction of high-voltage cable accessories, key parameters such as the reserved length of the cable core, the length of insulation stripping, the length of the semi-conductive layer treatment, the crimping position, and the dimensions after crimping directly affect the construction quality and operational reliability. Current quality inspections typically rely on construction experts arriving on-site to verify the measurements using measuring tools, or on-site personnel taking photos and videos which are then submitted to experts for retrospective analysis. The former method consumes significant expert resources and incurs travel costs, while the latter method is significantly delayed, and ordinary two-dimensional images lack reliable scale information and are easily affected by shooting distance, viewing angle, and lens distortion, making it difficult to directly obtain the true spatial dimensions. For the latter method, some remote video collaboration solutions exist: One related prior art is remote video collaboration or augmented reality remote assistance technology. For example, patent document CN113936121B discloses an AR annotation setting method and a remote collaboration system. This type of technology can transmit first-person view video from on-site personnel to remote experts, allowing experts to guide on-site personnel through voice, arrows, wireframes, or visual annotations. Such annotations are typically based on two-dimensional video image coordinates, capable of expressing operational intentions, but cannot directly provide the absolute depth and actual distance of the annotation points in real space. When the on-site camera moves, the video image is scaled or cropped, or the object to be measured has significant spatial undulations, the two-dimensional annotations also struggle to stably correspond to the actual measurement position. In other words, while this expert retrospective analysis can help experts view the construction site in real time and make voice or visual annotations, it typically only provides prompts on the two-dimensional video image and cannot stably and accurately obtain the three-dimensional endpoint coordinates and actual length of the object to be measured solely from the remote image, nor can it accurately reflect the measurement position selected by the expert back to the augmented reality view of the on-site personnel.
[0003] Another related prior art is dimension measurement based on monocular images. For example, this involves setting known-sized rulers, markers, or reference objects in the image, establishing homography relationships based on the known camera pose and the plane where the target is located, or using monocular depth estimation algorithms to calculate the target distance and size. This type of technology typically relies on known-scale references, targets approximately located on the same plane, the camera and target maintaining a specific pose, or statistical depth provided by training data. However, for cylindrical surfaces, curved surfaces, obstructed areas, and reflective materials at high-voltage cable accessory construction sites, these assumptions are often difficult to consistently meet; monocular depth results may also lack stable absolute scale, making them difficult to directly use for accurate verification of critical construction dimensions.
[0004] Another category of related prior art is binocular or multi-view vision measurement technology, as well as active depth measurement technologies such as structured light, time-of-flight cameras, and lidar. For example, the design of a fixed dual-camera multi-view reconstruction measurement method disclosed in the Journal of the National University of Defense Technology (Vol. 40, No. 4); a geometric length measurement method and system based on AR glasses disclosed in patent application publication number CN116188562A; and a BAR / VR glasses and a system, method, electronic device, and storage medium for measuring height disclosed in patent authorization announcement number CN114022543. This binocular or multi-view vision can perform triangulation based on the baseline and corresponding points between calibrated cameras, thereby recovering three-dimensional coordinates with actual scale. Active depth sensors can directly output depth information, such as a remote positioning method for nuclear reactor internal components based on binocular vision disclosed in patent authorization announcement number CN118887287A. However, conventional binocular measurement systems often use specially arranged, synchronized cameras with similar imaging characteristics, and typically perform local measurements on pre-detected feature points, regular targets, or dense parallax. High-definition color cameras on head-mounted augmented reality devices may differ significantly from wide-angle or fisheye positioning cameras in terms of field of view, resolution, distortion model, and exposure characteristics. Furthermore, the measurement endpoints selected by remote experts in high-definition images may be located in low-texture, edge, highly reflective, or non-feature point positions. Therefore, conventional feature matching or dense binocular measurement methods cannot be directly applied to this scenario. Active depth sensors may also be limited by device size, power consumption, measurement range, reflective surfaces, and on-site deployment conditions. Summary of the Invention
[0005] To address the shortcomings of existing technologies, such as the high dependence of high-voltage cable accessory construction quality inspection on on-site experts, the inability of two-dimensional images to directly provide reliable spatial dimensions, the difficulty in accurately representing remotely selected locations in the on-site space, and the lack of validity verification of automatic image matching results, this application provides a remote three-dimensional dimension measurement and on-site marking system and method based on multi-camera images from augmented reality devices.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a remote three-dimensional dimension measurement and on-site marking system, comprising: The on-site head-mounted augmented reality device is equipped with a first optical camera, a second optical camera, and a spatial positioning module. The first optical camera is a high-definition color optical camera used to provide clear images of selected points for remote users, supporting precise point selection on targets. The second optical camera has a wider field of view than the first optical camera, providing a wider epipolar search range to assist in corresponding point search and depth measurement. The first and second optical cameras are rigidly mounted and share a common field of view. The spatial positioning module is used to acquire the dynamic pose of the first optical camera in the augmented reality world coordinate system. The remote expert module at the remote end is used to receive image data directly or indirectly transmitted from the head-mounted augmented reality device at the field end, respond to the annotation operation of the remote user, and generate at least two annotation points on the image captured by the first optical camera. The processing unit is used to search for corresponding points in the epipolar search area of the image acquired by the second optical camera based on the image coordinates of the marked points, optimize the searched corresponding points to obtain a corresponding point positioning accuracy of less than one pixel, calculate the three-dimensional world coordinates of the marked points in real space through binocular parallax triangulation based on the optimized corresponding point coordinates, and verify the validity of the calculation results. After the verification is passed, the real spatial distance between any two marked points is calculated based on the three-dimensional world coordinates. The processing unit is located in any one of the head-mounted augmented reality device at the field end, the remote expert module at the remote end, or an independent server that is communicatively connected to both the field end and the remote end. The head-mounted augmented reality device also includes a projection display screen. After receiving the real spatial distance calculated by the processing unit directly or indirectly, the head-mounted augmented reality device overlays and renders the marked points and the corresponding real spatial measurement results on its projection display screen in an augmented reality manner.
[0007] Based on the same inventive concept, this application also provides a remote three-dimensional dimension measurement and on-site marking method, applied to the processing unit of the aforementioned remote three-dimensional dimension measurement and on-site marking system, comprising the following steps: Image pose acquisition and world coordinate system establishment steps: acquire image data acquired by the first and second optical cameras of the head-mounted augmented reality device at the field end, the world coordinate system established at the field end, and the dynamic pose sequence of the device in the world coordinate system acquired by the spatial positioning module of the head-mounted augmented reality device as the dynamic pose sequence of the first optical camera in the world coordinate system; the first optical camera is a high-definition color optical camera, the second optical camera has a field of view larger than that of the first optical camera, and the first optical camera and the second optical camera are rigidly mounted and have a common field of view.
[0008] Measurement frame synchronous acquisition steps: Receive measurement instructions from the remote expert module at the remote end to synchronously acquire images; determine measurement frame pairs through a preset time synchronization threshold; and perform asynchronous exposure time motion compensation on the position vector and attitude quaternion corresponding to the exposure time of the measurement frame pair through linear interpolation to obtain the equivalent relative pose aligned to the asynchronous exposure time. Point selection and image coordinate acquisition steps: Receive the image coordinates of at least two annotation points sent by the remote expert module at the remote end. The annotation points are generated by the remote user in response to the annotation operation when the image captured by the first optical camera is transmitted to the display window of the remote expert module. The steps for establishing polar geometric constraints are as follows: distortion correction is performed on the measurement frame pairs, and polar geometric constraints are established based on the relative pose between the first optical camera and the second optical camera to obtain the epipolar line or epipolar line search region. Corresponding point search steps: Based on the image coordinates of the marked points, candidate points are extracted from the epipolar search area of the image acquired by the second optical camera in the measurement frame pair, and local feature matching is performed to find the corresponding points; Triangulation and 3D coordinate solution verification steps: Sub-pixel optimization is performed on the searched corresponding points to obtain a corresponding point positioning accuracy of less than one pixel; based on the sub-pixel optimized corresponding point coordinates, the 3D world coordinates of the at least two marked points in real space are solved by binocular parallax triangulation; the validity of the solved 3D world coordinates of the marked points is verified; if the verification fails, the process jumps to the measurement frame synchronization acquisition step and re-executes the steps that started in the measurement frame synchronization acquisition step; if the verification passes, the next step of the triangulation and 3D coordinate solution verification steps is executed. Steps for calculating and outputting distance between annotation points: Calculate the actual spatial distance between annotation points based on their three-dimensional world coordinates, and then send the calculation results and the valid status of the annotation points back to the remote expert module for display.
[0009] Augmented Reality Reflection Steps: The marked points and corresponding measurement results are overlaid and rendered in an augmented reality manner on the projection display screen of the head-mounted augmented reality device.
[0010] As a preferred technical solution of the present invention, when determining the measurement frame pair, the pixel displacement is estimated by the exposure time difference and compared with the preset pixel displacement threshold to exclude unsuitable measurement frame pairs. When an unsuitable measurement frame pair is found, the head-mounted augmented reality device at the site is prompted to re-acquire the data under the operation of the on-site personnel.
[0011] As a preferred technical solution of the present invention, the polar geometric constraint establishment step uses whole image correction, applies a distortion model to all pixels of the measurement frame to remove distortion, calculates the respective correction rotation matrix based on the fixed or time-compensated relative pose between the first optical camera and the second optical camera, constructs a remapping table for each coordinate, traverses all pixels, and resamples the coordinates in the original image according to the remapping table to generate their respective corrected images, so that the vertical coordinate axes of the two corrected images are parallel to each other.
[0012] As a preferred embodiment of the present invention, the polar geometry constraint establishment step uses the annotation points selected by remote experts and the candidate pixels in the second optical camera to perform distortion removal using a distortion model and transform them to normalized observation rays in their respective coordinate systems. The pixels of the two cameras are mapped to an ideal pinhole plane. Based on the relative rotation and translation between the first and second optical cameras, an essential matrix and a fundamental matrix are constructed. A straight epipolar line is determined on the ideal pinhole plane to satisfy the relationship between the essential matrix and the normalized observation ray. The epipolar line vector uses... It means that in the formula F The fundamental matrix, v Ai It is a marker point on the ideal pinhole plane of the first optical camera. i , l Bi It is a marker point i The corresponding polar line in the ideal pinhole plane of the second optical camera.
[0013] As a preferred technical solution of the present invention, the polar geometric constraint establishment step uses the labeled pixel selected by the remote expert to perform distortion removal using the distortion model of the first optical camera and transforms it into a normalized observation ray in the coordinate system of the first optical camera. Adaptive sampling is performed along the aforementioned observation ray and the three-dimensional points formed by different assumed depths, so that the Euclidean distance between image points at adjacent sampling depths is no greater than a preset pixel step size. The three-dimensional points are then projected onto the original second image through the projection model of the second optical camera and connected to form a curved finite polar line search trajectory.
[0014] As a preferred technical solution of the present invention, the method for determining the epipolar search area in the corresponding point search step is to perform a one-dimensional search within the parallax range along the horizontal epipolar line with the same row number as each marked point in the second corrected image, and to expand the preset pixel width in the vertical direction to form an epipolar search band.
[0015] As a preferred embodiment of the present invention, the polar search region is determined along the polar equation in the corresponding point search step. A one-dimensional search is performed within the parallax range along the determined straight line direction, and a preset pixel width is extended along the epipolar normal direction to form an epipolar search band, where x and y are the horizontal and vertical coordinate components of the homogeneous pixel coordinates of the candidate point. These are the three components of the homogeneous coordinate expansion of the polar lines.
[0016] As a preferred technical solution of the present invention, the method for determining the epipolar search area in the corresponding point search step is to set a search band of preset width near the generated curved finite epipolar search trajectory.
[0017] As a preferred technical solution of the present invention, the local feature matching uses a coarse matching followed by fine matching. The coarse matching uses normalized cross-correlation matching cost for similarity measurement. After the local feature matching, similarity screening is also performed to exclude the case where the candidate point and the labeled point cannot match. The similarity screening method is not limited to the discrimination screening of the best and second-best matching or the left-right consistency verification. If the candidate point and the labeled point cannot match, the process jumps to the measurement frame synchronization acquisition step and re-executes the steps that started the measurement frame synchronization acquisition step.
[0018] The beneficial effects of this invention are: This application integrates solutions from augmented reality, computer vision, binocular or multi-view vision measurement, remote collaboration, and industrial construction quality control. It combines remote interactive point selection, multi-camera geometric measurement, result reliability verification, and augmented reality spatial feedback. This solution achieves on-site length measurement based on synchronous images from heterogeneous dual cameras of a head-mounted augmented reality device and interactive point selection by remote experts. Remote experts can select the start and end points of the component to be measured while viewing real-time on-site images. Multiple calibrated cameras on the head-mounted augmented reality device simultaneously acquire on-site images. The remote expert selects the endpoints to be measured from the images, automatically reconstructs the three-dimensional coordinates of the two endpoints, and calculates the three-dimensional straight-line distance. The system automatically obtains three-dimensional measurement results with actual scale, avoiding the influence of relying solely on two-dimensional estimations. The measurement endpoints and results are spatially marked on the augmented reality device and overlaid on the view of on-site personnel. Continuous marking enables on-site personnel to accurately understand the measurement location and immediately complete verification or rectification, achieving real-time remote measurement and on-site verification of critical construction dimensions. By employing simultaneous multi-camera data acquisition, calibrated camera geometry, and binocular or multi-camera triangulation as supplementary methods, and in conjunction with mandatory on-site inspections and measurements, key dimensions can be remotely verified by on-site personnel. Construction experts can simultaneously provide remote support to multiple construction sites, reducing travel and waiting times, lowering costs associated with problem tracing and rework, and improving the timeliness and efficiency of quality control in high-voltage cable accessory construction. The deviation between measured and actual values in this solution does not exceed 5%, making it particularly suitable for pre-inspection of key dimensions, remote verification, and concurrent support across multiple sites during high-voltage cable accessory construction. It serves as an auxiliary means for parallel inspections by supervisors and on-site guidance from manufacturer technical service personnel. Attached Figure Description
[0019] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the architecture of an embodiment of a remote three-dimensional dimension measurement and on-site marking system according to this application, including a separate service calculation terminal; Figure 2 This is a flowchart illustrating an embodiment of a remote three-dimensional dimension measurement and on-site marking method according to this application; Figure 3 A schematic diagram of an embodiment for searching corresponding points in this application; Figure 4 This application aims to anchor the world coordinates of the measurement endpoints and display the augmented reality back. Detailed Implementation
[0020] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0021] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0022] like Figure 1 As shown, this application also provides an embodiment of a remote 3D dimension measurement and on-site marking system based on multi-camera images from an augmented reality device, including a head-mounted augmented reality device at the on-site end, a remote expert module and a processing unit at the remote end.
[0023] The head-mounted augmented reality device includes at least a rigidly connected and pre-calibrated camera A (serving as the first optical camera), a wide-angle camera B (serving as the second optical camera), and a spatial positioning module. In this embodiment, camera A is a high-definition color camera, which refers to a color camera with a resolution of at least 1920x1080. This is consistent with the industry standard definition of high definition in standards such as "GY / T 295—2015 Technical Requirements and Measurement Methods for Broadcast-Grade High-Definition Cameras" or "GA / T 1211—2014 Technical Requirements for Security and Prevention High-Definition Video Surveillance Systems." Such cameras have a horizontal resolution greater than or equal to 800 TVL, meeting the requirement of providing clear point selection images for remote users to support accurate point selection on small targets. Camera A uses a common pinhole lens, while camera B is a wide-angle or fisheye camera with a wider field of view than camera A. It is used to form a binocular visual measurement combination with camera A, providing a wider epipolar search range than camera A to assist in corresponding point search and depth measurement. Cameras A and B are rigidly mounted and share a common field of view covering the object under test. In this embodiment, the head-mounted augmented reality device (DVR) pre-obtains the intrinsic parameters, distortion parameters, and relative pose relationship between the two cameras at a real-scale scale. The spatial positioning module can use visual-inertial V-SLAM combined with an inertial measurement unit (IMU) to output a 6-DOF pose. If anchor points can be pre-placed, ultra-wideband (UWB) combined with an IMU can also be used to output a 6-DOF pose. Fixed camera calibration parameters are used for binocular triangulation, while dynamic world pose is used to anchor the measurement results and continuously display them in the field space. In this embodiment, the DVR is used to transmit real-time first-view images to a remote end and synchronously acquire auxiliary images from cameras A and B according to measurement image acquisition instructions sent by a remote expert module. The measurement image from camera A is then sent to the remote expert module. The DVR may include a projection display screen for overlaying and rendering the marked points and corresponding real-space measurement results in an augmented reality manner. In this case, the DVR and the projection display screen together form an entity similar to augmented reality glasses.
[0024] The remote expert module is deployed on a remote computer, such as a remote expert computer. The remote expert module selects the start and end points of the object to be measured only in the measurement image of camera A. The remote expert module converts the corresponding image coordinates into the original measurement image coordinates and then sends them back to the head-mounted augmented reality device.
[0025] The remote expert module is used to respond to measurement operations initiated by expert users, thereby sending measurement image acquisition commands to the head-mounted augmented reality device.
[0026] The head-mounted augmented reality device performs distortion correction or epipolar correction on two synchronized images and the coordinates of the start and end points based on the calibration parameters of camera A and camera B, and determines the finite epipolar search regions of the start and end points in the image of camera B, respectively. Within the corresponding search regions, the corresponding points of the start and end points in the image of camera B are determined by one or more of the following conditions: local image feature similarity, epipolar constraint, best-match and second-best-match discrimination, and bidirectional matching consistency.
[0027] After obtaining two sets of corresponding points, the head-mounted augmented reality device performs binocular triangulation directly or indirectly based on the actual scale relative pose relationship between camera A and camera B, calculates the three-dimensional coordinates of the starting point and the ending point respectively, generates the length measurement result based on the Euclidean distance between the two three-dimensional coordinates, and sends the measurement result back to the remote expert module.
[0028] Simultaneously, the head-mounted augmented reality device directly or indirectly transforms the three-dimensional coordinates of the start and end points, which have passed validity verification, from the camera coordinate system at the time of image acquisition to the spatial positioning coordinate system of the augmented reality device. Based on the current left and right eye view matrices, projection matrices, and optical distortion compensation parameters, the measurement start and end points are superimposed as spatial markers in the binocular field of vision of the on-site operator, thus forming a complete measurement process of "remote experts selecting points in the specified measurement image - device-side cross-camera corresponding point search - three-dimensional coordinate recovery with actual scale - length result feedback - on-site spatial marker verification".
[0029] The collaboration between the remote expert module and the head-mounted augmented reality device described in this technical solution specifically refers to the combined process used to trigger the above-mentioned synchronous measurement image acquisition, transmit measurement endpoints, perform cross-camera 3D measurement, and display the measurement endpoints back to the field space. It does not involve any limitation on general remote audio and video communication, general augmented reality annotation, document sharing, or other remote collaboration functions themselves.
[0030] The processing unit is used to search for corresponding points in the epipolar search area of the image acquired by the second optical camera based on the image coordinates of the marked points, optimize the searched corresponding points to obtain a corresponding point positioning accuracy of less than one pixel, calculate the three-dimensional world coordinates of the marked points in real space through binocular parallax triangulation based on the optimized corresponding point coordinates, and verify the validity of the calculation results. After the verification is passed, the real spatial distance between any two marked points is calculated based on the three-dimensional world coordinates. The processing unit is located in any one of the head-mounted augmented reality device at the field end, the remote expert module of the remote end device, or an independent server that is communicatively connected to the field end and the remote end. When the computing power of the head-mounted augmented reality (HAR) device is sufficient, the processing unit's calculations can be performed on the HAR device at the field end. When the computing power of the HAR device is insufficient but the computing power at the remote end is sufficient, the processing unit's calculations can be performed on the computer or workstation housing the remote expert module. When the computing power of both the field-end and remote-end processors is insufficient, or when multi-person collaboration, data retention, or flexible scaling of computing power is required, the processing unit, as the main calculation entity, can be... Figure 1 As shown, the main calculation is placed on a server that is simultaneously connected to the augmented reality glasses and the remote expert computer. The server can be a standalone cloud server or an edge server.
[0031] like Figure 2 As shown, an embodiment of a remote three-dimensional dimension measurement and on-site marking method is provided. The main body executing this method is the processing unit of the remote three-dimensional dimension measurement and on-site marking system. The method is described from a single-side perspective of the processing unit, but for the sake of convenience, this embodiment describes the following steps from multiple perspectives: S1 System Composition and Camera Configuration: Configure and assemble the aforementioned remote three-dimensional dimension measurement and on-site marking system.
[0032] S2 Pre-calibration and coordinate system definition: that is, pre-calibrating the intrinsic parameters, lens distortion parameters, and relative poses between cameras A and B.
[0033] The camera matrix and distortion parameters of camera A are denoted as K. A and D A The camera matrix and distortion parameters of camera B are denoted as K. B and D B For any spatial point X in the camera A coordinate system A Its coordinates X in the camera B coordinate system B satisfy: ,in R BA This is the relative rotation matrix for transforming from camera A coordinate system to camera B coordinate system. t BA Let be the translation vector that transforms the camera coordinate system from camera A to camera B, where t BA The actual length units used are millimeters or meters. The camera's relative pose described above is a fixed calibration parameter of the device. The head-mounted augmented reality device also obtains the dynamic pose of camera A relative to the augmented reality world coordinate system at the moment of shooting through a spatial positioning module. ,in This represents the transformation from the camera A coordinate system to the world coordinate system. In this application, the world coordinate system corresponds to the scene coordinate system definition in GB / T 38247—2019 Information Technology Augmented Reality Terminology.
[0034] Synchronous acquisition of S3 measurement frames. When a remote expert initiates a measurement operation, the remote expert module at the remote end of the remote 3D dimension measurement and field marking system sends a measurement frame acquisition command to the head-mounted augmented reality device via HTTP, WebSocket, or other network communication mechanisms. Upon receiving the command, the head-mounted augmented reality device at the field end establishes a unique measurement task identifier for this measurement and selects a set of images belonging to the same measurement task and with the closest time from the image buffer queues of camera A and camera B, respectively denoted as measurement image group A and auxiliary image group B. It then records the timestamp, frame number, image size, and pose information. The data stored or associated with the measurement images includes at least: measurement task identifier, image frame number, sensor exposure timestamp, exposure duration, original image width and height, image rotation and mirror status, effective cropped area, calibration versions of camera intrinsic parameters and distortion parameters, fixed relative pose of camera A and camera B, equivalent relative pose adopted after time alignment, and the world pose of camera A at the time of measurement image acquisition. Camera calibration parameters can be pre-stored in the augmented reality device and referenced via version number, eliminating the need for repeated transmission with each measurement. The camera timestamp preferably uses the monotonic clock value corresponding to the sensor's start exposure time, exposure midpoint, or frame start time, avoiding the direct use of image encoding completion time or network arrival time. The monotonic clock value refers to a timing source that increases unidirectionally without jumps. Two cases are considered: In the first implementation, where hardware synchronization is available, cameras A and B are exposed by the same trigger signal, or the main camera outputs a synchronization signal to trigger the acquisition by the secondary cameras. The system records the actual exposure time t of both cameras. A and t B And calculate the time difference. , ,in This is the preset time synchronization threshold. When... Greater than the preset time synchronization threshold When two frames are combined into a measurement frame pair, if the time difference exceeds a threshold, the earlier frame is discarded and the search continues for an image with a closer time. The synchronization threshold can be preset based on device movement speed, camera field of view, target distance, allowable pixel error, and actual measurement accuracy requirements. For example... It is set to 10ms, but usually does not exceed one frame period, for example, under 60fps acquisition conditions, it is no more than 16.7ms.
[0035] In the second approach, where simultaneous hardware exposure of the two cameras is not possible, both cameras record timestamps using the same monotonic clock or a time base corrected for deviation. Subsequent post-processing employs time alignment and motion compensation. The spatial positioning module of the head-mounted augmented reality device at the field end outputs the device pose sequence at a frequency higher than or equal to the camera frame rate. For adjacent positioning times... t k and t k+1 Camera exposure time between t It can perform linear interpolation on the position vector and spherical linear interpolation on the pose quaternion to obtain the camera A world pose corresponding to the exposure time. The attitude quaternions in this application conform to the definitions of terms in 3.1.4 and data structures in 5.5.8 of GB / T 44220—2024 Virtual Reality Device Interface Positioning Devices. Compared with Euler angles, quaternions do not encounter gimbal lock and are more suitable for interpolation operations.
[0036] The specific implementation is as follows: Define the interpolation factor α:
[0037] In the formula t For the camera exposure time, t k and t k+1 These are adjacent positioning times.
[0038] Interpolation calculations are performed on position and orientation:
[0039] In the formula p(t) Indicates the device at time t The specific location at that time P k and P k+1 Indicates at discrete time t k and t k+1 Location, q(t) Indicates the device at time t The specific posture at that time, i.e., the orientation of the head-mounted augmented reality device, q k and q k+1 Indicates at discrete time t k and t k+1 The posture.
[0040] Let the fixed calibration transformation T BAThis represents a homogeneous transformation from camera coordinate system A to camera coordinate system B. The exposure time of camera A obtained through interpolation. t A and camera B exposure time t B World pose of camera A and We can find the position of camera B. t B pose relative to the world coordinate system at any moment And further obtained from camera A at t A Transform the coordinate system at time t to that of camera B B Equivalent relative transformation of the time coordinate system .
[0041] Will Decompose into equivalent rotation matrices R eff and equivalent translation vector t eff Then, it can be used in subsequent polar geometric constraint establishment and triangulation. R eff and t eff Alternative R BA and t BA This is to compensate for pose changes caused by the movement of the head-mounted augmented reality device during the exposure of the two cameras.
[0042]
[0043] The aforementioned pose-based time compensation assumes that the object under test remains stationary or nearly stationary between two exposures. The system can determine whether the scene meets this condition based on the translation, rotation angle, global optical flow of the image, or on-site target tracking results between the two exposure times. When the estimated pixel displacement caused by device movement exceeds the preset pixel displacement threshold, or when the object under test itself undergoes movement, occlusion changes, or deformation, the system determines that the set of images is not suitable for measurement and prompts on-site personnel to stabilize the device and the object under test before re-acquiring the images.
[0044] The extended description of the method for excluding unsuitable measurement frame pairs, based on the determination that the estimated pixel displacement caused by device movement exceeds a preset pixel displacement threshold, indirectly constrains the estimated pixel displacement. During image acquisition, the estimated pixel displacement is typically estimated coarsely based on the exposure time difference between the two cameras. For example, assuming a user typically rotates their head at an angular velocity of 30° / s when wearing a head-mounted device, the camera's horizontal field of view is 90°, the camera image width is 1920 pixels, and the exposure time difference between the two cameras is 0.01 seconds, then the estimated pixel displacement is 6.4 pixels. In one embodiment, the threshold can be 8 pixels, but typically does not exceed 10 pixels.
[0045] S4 Remote Point Selection and Coordinate Restoration: This involves a remote expert selecting a start and end point in measurement image A and restoring the interface coordinates. Inverse transformation to original pixel coordinates The head-mounted augmented reality device sends the measurement image A, along with its associated frame number, original size, cropped area, and orientation information, directly or indirectly to the remote expert module via a separate server. The remote expert module displays the pre-processed measurement image A (including cropping, scaling, rotation, mirroring, and margin adjustments) on the image display interface. The expert then selects the start and end points of the length to be measured, obtaining the interface coordinates with the upper left corner of the display interface as the origin. and The remote expert module binds the selected point coordinates to the corresponding measurement task identifier and image frame number to prevent the selected point results from being incorrectly applied to other image frames.
[0046] The following sections explain the methods for coordinate reconstruction in several scenarios: First, a coordinate restoration method is adopted that scales the image proportionally while preserving its completeness. Let the actual image size displayed on the screen after cropping and orientation adjustment be... The size of the area used to display the image in the display interface is There are two display modes: "Full Display" and "Fill and Crop". When using "Full Display" mode, the image scaling factor 's' and the horizontal and vertical margins are adjusted. , It can be represented as:
[0047] The margin is a positive value, which may result in black borders. For the interface point (U, V) located within the actual display area of the image, the margin is first subtracted and then divided by the scaling factor to obtain the pixel coordinates in the image after orientation adjustment. , If the point is located outside the reserved area, the point selection is invalid.
[0048] If the display is set to "fill and crop" mode, the scaling factor mentioned above can be changed to:
[0049] The margin allowance can be negative, used to represent the image edge that is cropped by the display interface.
[0050] Then, following the reverse order of the image display process, reverse mirroring, reverse rotation, and cropped area backfilling are performed sequentially. For example, if the displayed image is horizontally mirrored after orientation adjustment, then first... If the displayed image is rotated 90 degrees clockwise relative to the original cropped image, then it can be displayed according to... Restore the cropped image coordinates to their original positions before rotation, where H c The height of the cropped image before rotation; finally, based on the coordinates of the top left corner of the original cropped area (c x ,c y Obtain sensor pixel coordinates For 180° and 270° rotations or vertical mirroring, the corresponding inverse transformation can be used.
[0051] As a more general implementation of homography matrices, cropping translation, rotation, mirroring, scaling, and margin translation can be uniformly represented from the original sensor pixel coordinates. To interface coordinates u D 3×3 homogeneous transformation matrix The matrix is saved during image display, and the original sensor pixels are recovered in one go through its inverse matrix after remote point selection. coordinate:
[0052] in, In this application, the tilde symbol on the vector indicates homogeneous coordinate normalization. After restoring the coordinates, u and v Restricted to and Within this range, it ensures that the reconstructed pixel coordinates truly fall within the effective imaging area of the physical sensor, preventing the calculation of negative or out-of-resolution coordinates. Floating-point form can be retained to support subsequent sub-pixel calculations, with discretization and rounding only performed in the final output or physical mapping. The pixel coordinates of the start and end points in the original sensor image of camera A are obtained using the above method. and In this step, both the interface coordinates and the original pixel coordinates correspond to the coordinates in the image coordinate system of "GB / T 38247—2019 Information Technology Augmented Reality Terminology".
[0053] S5 distortion correction and polar geometry constraint establishment involves distorting the coordinates of images A, B, and endpoints to obtain epipolar lines or epipolar search regions.
[0054] Camera A and camera B can have different resolutions, fields of view, and distortion models. For any camera... j The imaging process can be represented as a projection function from the three-dimensional observation ray to the pixel coordinates. The process of back-projecting pixels into unit observation rays is represented as follows: For ordinary pinhole lenses, a radial-tangential distortion model can be used; for fisheye lenses, equidistant projection, the Kannala-Brandt model, or other fisheye models consistent with the actual lens calibration can be used. The start and end pixels are first converted into normalized observation rays in the camera's A-coordinate system using a back-projection function.
[0055] In the formula d Ai Indicates the first i Normalized observation rays of each point in the camera A coordinate system normalize ( · The function is a normalization function to transform non-unit length vectors into unit length vectors. The normalized vector represents the direction of a unit observation ray triggered from the camera's optical center and passing through the imaging plane. i For pixel index, in this embodiment it represents the 0th pixel and the 1st pixel.
[0056] Typically, the entire image is calibrated before being used in subsequent steps. In implementations using whole-image binocular epipolar correction, the calibration rotation matrices for camera A and camera B are calculated based on the relative poses of the two cameras, either fixed or after time compensation. R A r , R B r The correction rotation matrix is used to rotate the observation directions of the two cameras to a common correction coordinate system, making one axis of the correction coordinate system parallel to the line connecting the optical centers of the two cameras, and making the vertical coordinate axes of the two corrected images parallel to each other. Then, common correction camera intrinsic parameters are selected. K r Or select separately K A r , K B r The corrected image size and effective area are determined based on the common field of view of the two cameras.
[0057] For pixels in the original image u j Its position in the corrected image u j r This can be achieved through a combination of "original pixel back projection - correction rotation - ideal pinhole projection". After calibration, the device can pre-calculate the original image sampling position corresponding to each corrected image pixel, forming a remapping table M between camera A and camera B. A and M BDuring measurement, bilinear interpolation or other image resampling methods are used to generate corrected images A1 and B1, thereby reducing the amount of real-time computation.
[0058]
[0059] In the formula j It is the camera's index. u j r Indicates camera j Position in the corrected image.
[0060] By combining distortion correction and epipolar correction into a single remapping, image blurring caused by multiple resampling is avoided, while also improving efficiency.
[0061] After completing the binocular epipolar calibration, the vertical coordinates of the same spatial point in A1 and B1 should be the same or only have coordinates less than the threshold ε. epi The residual difference. The formula is expressed as... In the formula and represent Indicates the first i The vertical coordinates of a point in calibrated images A1 and B1. Ideally, the same spatial point should lie on the same horizontal scan line in calibrated images A1 and B1, i.e., the residual difference is 0. However, in practical systems, it is affected by camera calibration errors, fisheye distortion correction residuals, synchronization errors, image resampling errors, and matching errors, so a small residual vertical difference is usually allowed. ε epi The unit is pixels; in one embodiment, ε epi The limit is set to 3 pixels, typically no more than 5 pixels. Therefore, for an endpoint in camera A, a search band with a finite width can be established near the same horizontal scan line of the image calibrated by camera B. Since the common field of view of the heterogeneous cameras may only cover a portion of the calibrated image, the system also generates an effective pixel mask; when an endpoint falls outside the common effective area after calibration, it is directly determined that the auxiliary camera cannot be used for that endpoint measurement.
[0062] In implementations that do not generate a complete corrected image, distortion correction and ray transformation can be performed only on the endpoints selected by the remote expert and candidate pixels in camera B. After mapping the two camera pixels to the ideal pinhole plane, the essential matrix can be constructed based on relative rotation and translation. E and fundamental matrix F :
[0063] in, It is an antisymmetric matrix composed of relative translation vectors. R BA Let be the relative rotation matrix of camera B with respect to camera A. The endpoints of image A in the ideal pinhole plane. vAi The epipolar line in image B can be represented as In the formula v Ai Endpoint i Homogeneous pixel coordinates in image A l Bi Endpoint i The homogeneous coordinates of the corresponding epipolar line in image B.
[0064] The two methods described above are applicable when camera B uses a regular wide-angle lens. When camera B uses an ultra-wide-angle or fisheye lens, the following implementation method that preserves the fisheye imaging plane is used.
[0065] For implementations that preserve the fisheye imaging plane, three-dimensional points formed by the observation ray at the endpoint and different assumed depths can be sampled, and these points can be projected onto the original fisheye image through the fisheye projection function of camera B. The sampled points are connected to form a curved epipolar trajectory, without requiring the trajectory to be a straight line in the original fisheye image.
[0066] Specifically, the minimum allowable measurement distance Z of the device is as follows: min and maximum measurement distance Z max Let Z min <Z max The infinite observation ray corresponding to endpoint A of camera is restricted to Z. min To Z max A finite three-dimensional line segment. For high-voltage cable accessory construction scenarios, on-site operators wear head-mounted displays to observe the workpiece at hand; the location to be measured is typically within the reach of their arm. In one embodiment, Z... min Z is 0.3m. max It is 2.0m, usually Z min The value is taken as 0.25m to 0.5m, Z max The value ranges from 1.5m to 3.0m.
[0067] For a fisheye image B without binocular epipolar correction, the system does not directly limit the search area for the corresponding point to the straight epipolar line in image B. Instead, it generates a finite trajectory search area based on the camera ray model and the fisheye projection model. Specifically, the depth Z along the optical axis of camera A is used as the ray sampling parameter, and the value range of Z is the preset effective measurement distance range [Z]. min Z max For any sampling depth Z, candidate 3D points in the camera A coordinate system can be obtained according to the following formula: in X A ( Z Let Z be the three-dimensional vector of a point in the camera's coordinate system at a distance Z from the optical center. z To take values in [Z min Zmax The depth variable of the interval. d Ai,z Observation of rays per unit d Ai The component along the optical axis of camera A. Then, using the fixed relative pose from camera A to camera B, the candidate 3D points are transformed to the camera B coordinate system: In the formula X B ( Z Let Z be the three-dimensional vector of a point in the camera's B coordinate system at a distance Z from the optical center. R BA and t BA All are the same as the definitions in step S2 above.
[0068] Then, based on the fisheye projection model and internal parameters of camera B... K B and distortion parameters D B ,Will X B ( Z Projecting these points onto image B yields candidate image points: when Z In [Z] min Z max When the range changes, u B ( Z This forms a finite trajectory in image B.
[0069] Once the above function expression is determined, uniform sampling can be chosen to calculate at each depth. However, this method involves a large amount of computation and is inefficient. As an implementation method, the system can adjust the depth parameters. Z Adaptive sampling is performed. This adaptive sampling does not directly sample angles from the fisheye image, but rather controls the sampling density based on the pixel intervals after projecting candidate 3D points onto image B. When two adjacent sampling depths... Z k and Z k+1 The corresponding image points satisfy the condition that the Euclidean distance (calculated as the L2 norm of a vector) between two image points is no greater than a preset pixel step size. At that time, the sampling density of the interval is considered to meet the requirements. s pix The preset pixel step size; when the pixel interval is greater than the preset pixel step size s pix At that time, the depth range is further subdivided for sampling. s pixSet to 1 to 3 pixels. Using the above method, a finite curve search region can be formed in the fisheye image B by connecting or interpolating discrete sampling points, thus avoiding unconstrained searching across the entire fisheye image.
[0070] S6 Corresponding point search. In this embodiment, corresponding point searches are performed for the starting point and the ending point respectively.
[0071] For binocular epipolar calibration that has been completed and the calibration baseline length is... b Corrected focal length is f r The image can be used to calculate the parallax range based on the depth range, with the corrected focal length unit being pixels. According to the parallax-depth relationship, closer targets correspond to larger parallax, while distant targets correspond to smaller parallax. The search range can be expressed as: Therefore, according to the parallax definition formula, the following can be satisfied in the camera B-corrected image: The search is performed within the specified interval, and the vertical expansion is extended by ±w pixels to form an epipolar search band to accommodate residual calibration errors such as ε. epi The parallax direction is determined by the camera coordinate system and the definition of the correction projection matrix. In one embodiment, w is 3 pixels, but typically it is 1 to 5 pixels.
[0072] For the fundamental matrix F The corresponding uncorrected image, let the homogeneous pixel coordinates of the endpoints in image A be... The homogeneous pixel coordinates of the candidate points in image B are: . byu A The polar lines identified in image B are: ,in Polar equation The coefficients, i.e., the epipolar equation, are the horizontal and vertical coordinate components of the homogeneous pixel coordinates of the candidate point, where x and y are the equation. These are the three components of the homogeneous coordinate expansion of the epipolar line. Candidate point u B The normalized point-to-line distance to the polar line can be expressed as: For a fisheye image B that has not undergone binocular epipolar correction, the system sets a search band of preset width near the generated finite search trajectory and extracts candidate points within the search band for local feature matching.
[0073] Since cameras A and B may have different exposures, white balances, resolutions, and spectral responses, local images can be converted to grayscale or gradient maps before matching, and local mean elimination, variance normalization, histogram adjustment, or scale normalization can be performed. An n×n or multi-scale local support region can be selected centered on the endpoint of camera A. At each candidate location in the search band of camera B qSelect support regions of the same physical scale or those that have undergone scale compensation. A feasible similarity measure is Normalized Cross-Correlation (NCC): In the formula For the region and Similarity; I A and I B These are the pixel grayscale values within local windows of image A and image B, respectively. and They are regions and Average gray level, ε Minimum protection value, such as 10 -10 ; For the region and The normalized cross-correlation matching cost.
[0074] Besides normalized cross-correlation, matching costs can also be calculated using methods such as absolute difference, squared difference, Census transform and Hamming distance, gradient direction consistency, ORB or SIFT-type local descriptors, and learned local features. Normalized cross-correlation, absolute difference, and squared difference are all local image patch matching based on grayscale statistics; gradient direction consistency is based on gradient feature matching; Census transform and Hamming distance are based on Census binary matching; and ORB or SIFT-type local descriptors and learned local features are based on floating-point descriptor matching. As an optional further optimization strategy, to improve overall efficiency, the system can first use the normalized cross-correlation matching cost to perform coarse matching to determine a precise search window, and then perform other more refined local feature matching operations within this search window, thereby reducing computational overhead while ensuring accuracy. Traditional stereo matching algorithms require that the matching points must be stable feature points, i.e., feature points that can be detected by the feature point detector, such as spots. However, for cases where arbitrary pixels selected by remote experts are not stable feature points, such as smooth regions, this scheme prefers to use local image blocks or gradient support regions centered on the selected point, rather than requiring that the pixel must be pre-detected by the feature point detector. When the endpoint is located at the edge of an object, an asymmetric support region can be used according to the edge direction, or the matching cost can be calculated separately for both sides of the edge, to reduce the interference of the background region on the matching result. The asymmetric support region refers to constructing a support region only on the foreground side of the cable and not taking the background side at all. The calculation of the matching cost separately for both sides of the edge refers to dividing the support region into foreground and background sides, calculating the matching cost with the candidate region of image B separately, and then obtaining the final cost through weighted fusion. It should be noted that when the endpoint is located in a highly reflective or low-texture gradient region, the above edge optimization strategy can only alleviate background interference. The system can further combine steps such as left-right consistency verification and matching confidence evaluation to comprehensively judge the matching validity. When the confidence is lower than the threshold, it can prompt to reselect the endpoint.
[0075] To balance search range and sub-pixel accuracy, an image pyramid can be constructed. A coarse search is performed on the entire finite-pole line segment at a low-resolution layer, and then the best candidate positions are mapped layer by layer to a higher-resolution image for a finer search within a smaller neighborhood. Let q1 be the candidate point with the minimum matching cost, and q2 be the second smallest candidate point, with costs C1 and C2 respectively. The system can simultaneously require C1 to be below an absolute threshold and C1 / C2 to be below a ratio threshold, or require the difference between the best and second-best correlation coefficients to be above a preset threshold, to exclude multiple solutions caused by duplicate textures.
[0076] Regarding geometric consistency, the distance from candidate points to the theoretical epipolar line can be calculated, and a reverse search can be performed on camera A from candidate points of camera B. The search is considered successful when the distance between the obtained position and the original endpoint is no greater than a threshold. At that time, it is considered to have passed the left-right consistency check. In one embodiment, The value is 3 pixels, but the typical range is 1 to 5 pixels.
[0077] After obtaining the optimal position for integer pixels, parabolic fitting can be performed using the matching cost of the optimal position and its neighboring positions to obtain the sub-pixel disparity correction. Taking a one-dimensional horizontal search as an example, if C0 and C +1 Given the costs to the left, right, and left of the optimal position, respectively, the sub-pixel offset can be estimated using the following formula, which is then constrained to a range of one pixel: The final high-precision disparity value is obtained by superimposing the optimal integer disparity with the subpixel disparity correction.
[0078] As an optional further optimization strategy, the system can also combine texture intensity, best matching cost, best and second-best candidate discrimination, epipolar distance, bidirectional consistency error, and subpixel fitting stability into a matching confidence score. The weights of each matching confidence score index can be dynamically adjusted according to actual working conditions. (The remaining text appears to be incomplete and requires further context.) A0 and u A1 The corresponding point u in image B B0 and u B1 When the texture confidence is below the threshold, the candidate point is not unique, the endpoint is located in a low-texture or high-reflectivity area, the endpoint is occluded, or the endpoint is not within the common field of view of the two cameras, the system does not directly output the measurement result. Instead, it prompts the user to reselect the endpoint, reacquire the image, switch to another auxiliary camera, or send multiple candidate points to a remote expert for manual confirmation. Switching to another auxiliary camera refers to switching the viewpoint or imaging parameters of the two cameras.
[0079] Step S7 involves triangulation, 3D coordinate solving, and validity verification. This involves constructing a projection matrix to obtain the 3D coordinates of the start and end points, followed by validity verification. The coordinate system of camera A at the measurement reference time is used as the 3D reconstruction coordinate system; that is, the current pose of camera A is taken as the origin of the scene coordinate system. Corrected pixel coordinates are used, and a projection matrix containing the corresponding corrected camera intrinsic parameters is constructed. , in K A r , K B r These are the calibrated camera intrinsic parameters for camera A and camera B; I is the unit rotation matrix with a rotation angle of 0, and 0 is the translation zero vector with a translation distance of 0. R r BA , t r BAThese are the relative rotation matrix and relative translation matrix of camera B relative to camera A, respectively. , , R A r , R B r These are the correction rotation matrices for camera A and camera B, respectively.
[0080] One linear triangulation method constructs a homogeneous linear equation based on the cross product constraint between each image point and the projection matrix. ,in For three-dimensional endpoints i homogeneous coordinates A i The coefficient matrix is the coefficient matrix. A i It can be represented as: in, P A (k) and P B (k) Representing the projection matrix respectively P A and P B The k OK, u Ai , v Ai As endpoints i The horizontal and vertical coordinates of the corrected pixels in image A u Bi , v Bi As endpoints i The horizontal and vertical coordinates of the corrected pixels in image B.
[0081] Obtained through singular value decomposition A i The initial 3D coordinates of the start or end point can be obtained by dividing the homogeneous coordinates by the right singular vector corresponding to the minimum singular value and its fourth component. X i Besides the aforementioned algebraic method based on singular value decomposition and linear triangulation to solve homogeneous equations for initial 3D coordinates, a geometric method using the minimum distance intersection of two observation rays can also be employed to obtain initial values. To reduce the impact of pixel noise on the results, the initial coordinates can be used as a starting point to minimize the reprojection error between the two cameras using the Gauss-Newton method, the Levenberg-Marquardt method, or other nonlinear optimization methods to obtain accurate final output 3D coordinates. The optimization function with reprojection error as the objective is constructed as follows:
[0082] In the formula, The three-dimensional point with the smallest reprojection error , u Ai and u Bi For three-dimensional points The actual measured and observed coordinates in camera A and camera B, π A ( X i )and π B ( X i (a) is a three-dimensional point The theoretical pixel coordinates are calculated by projecting them onto the image planes of cameras A and B. This optimization function is solved iteratively: first, linearization is performed, then an incremental equation is constructed. The increment is then solved using the Gauss-Newton method, the Levenberg-Marquardt method, or other nonlinear optimization methods. The coordinates are updated iteratively. Iteration stops if the error between two consecutive iterations is less than a preset threshold, and the final coordinates are output. .
[0083] After triangulation is completed, validity verification is performed on each 3D point. First, [the following steps are taken]. X i The coordinates are transformed to those of camera A and camera B, respectively, requiring that the depth components are both greater than zero and within a preset effective distance range. Next, the angle θ between the two normalized observation rays in the same coordinate system is calculated. θ must be greater than a preset minimum angle to avoid significantly amplifying the depth error when the two rays are nearly parallel. The preset minimum angle is generally 1–5°, and in one embodiment it is set to 2.2°. Then, the... X i Reproject onto images A and B, and calculate the root mean square reprojection error. e i Compared to a preset pixel threshold, which is typically set to 1-5 pixels, in one embodiment, the preset pixel threshold is set to 3 pixels. Root mean square reprojection error. e i The calculation formula is used as follows: In the formula u Ai and u Bi For three-dimensional points X i The actual measured and observed coordinates in camera A and camera B, π A ( X i )and π B ( Xi (a) is a three-dimensional point X i The theoretical pixel coordinates calculated on the image planes projected onto cameras A and B. Only when... e i The 3D coordinates of an endpoint are deemed valid when the pixel value is less than a preset pixel threshold, the matching confidence level meets the requirements, the positive depth condition is met, and the angle between the observed rays satisfies the requirements. Alternatively, the 3D coordinate covariance can be estimated based on the pixel matching uncertainty and the triangulation Jacobian matrix; if the estimated standard deviation along the depth or length direction exceeds the allowable value, the point is marked as having low confidence, even if the reprojection error is small. Length calculation is only performed when both the starting point X0 and the ending point X1 pass verification.
[0084] Step S8: 3D length calculation and result output, i.e., through the formula... The distance between the two points is calculated, and the measured length and valid status are then transmitted back to the remote expert module.
[0085] When the starting point's three-dimensional coordinates 3D coordinates of the endpoint At that time, the three-dimensional straight-line distance between the starting point and the ending point L Calculated using Euclidean distance: Due to the relative translation vector between cameras t BA Calculated using actual length units L It has corresponding actual length units such as millimeters or meters. The system can round the results according to a preset resolution and transmit the measured length, the three-dimensional coordinates of the start and end points, the matching confidence level, the reprojection error, the angle of the observed ray, and the valid measurement status back to the remote expert module for display and archiving. If the covariance of the three-dimensional coordinates of the endpoints is estimated, the length uncertainty or confidence interval can also be calculated through error propagation, and the corresponding confidence level can be displayed next to the measurement results.
[0086] S9 measurement endpoint world coordinate anchoring and augmented reality echo, that is, using T WA ( t 0) Convert X0 and X1 to world coordinates X W0 X W1 The endpoints, connections, length values, and status are superimposed and displayed in the binocular field of view based on the left and right eye view matrices, projection matrices, and optical distortion compensation parameters.
[0087] Using the world pose of camera A at the time of image acquisition T WA ( t0), convert the camera A coordinate system, i.e., the starting point X0 and ending point X1 in the camera coordinate system definition in "GB / T38247—2019 Information Technology Augmented Reality Terminology", into the world coordinates X in the augmented reality world coordinate system. W0 and X W1 For homogeneous coordinates , can be represented as: The system establishes spatial anchor points for this measurement and sets X... W0 X W1 The connection between the two endpoints, the measured values, and the measurement task identifier are stored in the spatial anchor point coordinate system. Spatial anchor points can be directly constructed using the synchronous localization and mapping of augmented reality devices to create SLAM world coordinates, or they can be local anchor points established from the site plane, device structural features, visual markers, or local spatial maps. When the SLAM system relocalizes and outputs a world coordinate system correction transformation, the system applies the same correction transformation to the saved measurement endpoints to maintain their relative relationship with the real objects.
[0088] In each display frame In this process, the head-mounted augmented reality device acquires the latest head pose and, combined with fixed extrinsic parameters from the head coordinate system to the optical centers of the left and right eyes, obtains the view matrices from the world coordinate system to the left and right eye observation coordinate systems, respectively. For any one glance The world coordinate points are sequentially passed through the view matrix. and the eye projection matrix P E Transform to clipping space: when p E When the homogeneous component w is greater than zero and the normalized depth is within the visible range, perspective division is performed to obtain the normalized device coordinates, and the width W of the texture is rendered based on the eye. E and height H E Convert to two-dimensional pixel positions. With the top-left corner of the image as the origin, it can be represented as: , In the formula u E , v E For the coordinate point at E The horizontal and vertical coordinates in the texture rendered by the eye p x , p y To normalize the horizontal and vertical coordinate components of the device, p w is the scale factor for homogeneous coordinates.
[0089] The system can be used in X W0 and XW1 The system renders dots, crosshairs, spheres, or primitives with a fixed world size or fixed viewpoint size, always facing the observer, and renders 3D line segments between two points; length values can be displayed near the midpoint between the two endpoints, always facing the current observer. When the head of the augmented reality device moves, causing the previously selected point to be occluded, for devices with on-site spatial grids or depth information, such as AR headsets equipped with ToF cameras or depth sensors, the measurement marks can participate in depth testing to hide the corresponding parts when occluded by real structures; for optical perspective devices without reliable occlusion depth, dashed lines, transparency, or edge indicators can be used to indicate that the marks are outside the field of view or are occluded.
[0090] Before being sent to the display optical engine, the rendered textures for both the left and right eyes undergo pre-distortion based on the distortion calibration parameters of each eye's optical system. This process can be achieved through a pre-generated distortion compensation mesh: the mesh vertices store the mapping relationship between the ideal distortion-free texture coordinates and the display sampling coordinates. When rendering textures containing endpoints, lines, and text, resampling is performed according to this mesh, ensuring that the light rays after passing through the display lenses align with the ideal projection position in the direction of human eye observation. For optical engines requiring chromatic aberration compensation, different texture coordinate mappings can be used for the red, green, and blue channels respectively.
[0091] Through the aforementioned coordinate anchoring and eye-by-eye rendering process, on-site operators can simultaneously observe the starting point, ending point, 3D lines, and measurement values selected by remote experts within their binocular field of view. They can also confirm the measurement location with experts via voice or other remote collaboration channels. This feedback process uses only the 3D endpoints generated and validated in this measurement and does not limit general augmented reality annotation or other remote collaboration functions.
[0092] It should be noted that the remote measurement scheme provided in this application is not limited to experts selecting two annotation points, but can also select more than two annotation points. The scheme of this application will measure and annotate the distance of multiple annotation points in the order of annotation points, but this method will increase the possibility of the annotation points and candidate points not matching in the measurement frame or the three-dimensional coordinate point validity verification failing.
[0093] It should be noted that the remote measurement scheme provided in this application is not intended to replace the on-site inspection and handover test required by the mandatory standards for the quality of line construction. At the final inspection, acceptance, and supervision stages, remote measurement is only used as an auxiliary verification method to conduct pre-inspection, preliminary judgment and support for multiple concurrent sites for key dimensions, in order to discover problems in advance and reduce rework. It does not mean that it can be used as the sole basis for inspection.
[0094] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A remote three-dimensional dimension measurement and on-site marking system, characterized in that, include: The on-site head-mounted augmented reality device is equipped with a first optical camera, a second optical camera, and a spatial positioning module. The first optical camera is a high-definition color optical camera used to provide clear images of selected points for remote users, supporting precise point selection on targets. The second optical camera has a wider field of view than the first optical camera, providing a wider epipolar search range to assist in corresponding point search and depth measurement. The first and second optical cameras are rigidly mounted and share a common field of view. The spatial positioning module is used to acquire the dynamic pose of the first optical camera in the augmented reality world coordinate system. The remote expert module at the remote end is used to receive image data directly or indirectly transmitted from the head-mounted augmented reality device at the field end, respond to the annotation operation of the remote user, and generate at least two annotation points on the image captured by the first optical camera. The processing unit is used to search for corresponding points in the epipolar search area of the image acquired by the second optical camera based on the image coordinates of the marked points, optimize the searched corresponding points to obtain a corresponding point positioning accuracy of less than one pixel, calculate the three-dimensional world coordinates of the marked points in real space through binocular parallax triangulation based on the optimized corresponding point coordinates, and verify the validity of the calculation results. After the verification is passed, the real spatial distance between any two marked points is calculated based on the three-dimensional world coordinates. The processing unit is located in any one of the head-mounted augmented reality device at the field end, the remote expert module at the remote end, or an independent server that is communicatively connected to both the field end and the remote end. The head-mounted augmented reality device also includes a projection display screen. After receiving the real spatial distance calculated by the processing unit directly or indirectly, the head-mounted augmented reality device overlays and renders the marked points and the corresponding real spatial measurement results on its projection display screen in an augmented reality manner.
2. A remote three-dimensional dimension measurement and on-site marking method, applied to the processing unit of the remote three-dimensional dimension measurement and on-site marking system according to claim 1, characterized in that, Includes the following steps: Image pose acquisition and world coordinate system establishment steps: acquire image data acquired by the first and second optical cameras of the head-mounted augmented reality device at the field end, the world coordinate system established at the field end, and the dynamic pose sequence of the device in the world coordinate system acquired by the spatial positioning module of the head-mounted augmented reality device as the dynamic pose sequence of the first optical camera in the world coordinate system; the first optical camera is a high-definition color optical camera, the second optical camera has a field of view larger than that of the first optical camera, and the first and second optical cameras are rigidly mounted and have a common field of view; Measurement frame synchronous acquisition steps: Receive measurement instructions from the remote expert module at the remote end to synchronously acquire images; determine measurement frame pairs through a preset time synchronization threshold; and perform asynchronous exposure time motion compensation on the position vector and attitude quaternion corresponding to the exposure time of the measurement frame pair through linear interpolation to obtain the equivalent relative pose aligned to the asynchronous exposure time. Point selection and image coordinate acquisition steps: Receive the image coordinates of at least two annotation points sent by the remote expert module at the remote end. The annotation points are generated by the remote user in response to the annotation operation when the image captured by the first optical camera is transmitted to the display window of the remote expert module. The steps for establishing polar geometric constraints are as follows: distortion correction is performed on the measurement frame pairs, and polar geometric constraints are established based on the relative pose between the first optical camera and the second optical camera to obtain the epipolar line or epipolar line search region. Corresponding point search steps: Based on the image coordinates of the marked points, candidate points are extracted from the epipolar search area of the image acquired by the second optical camera in the measurement frame pair, and local feature matching is performed to find the corresponding points; Triangulation and 3D coordinate solution verification steps: Perform sub-pixel optimization on the searched corresponding points to obtain corresponding point positioning accuracy of less than one pixel; Based on the sub-pixel optimized corresponding point coordinates, the three-dimensional world coordinates of the at least two marked points in real space are calculated by binocular parallax triangulation; the validity of the calculated three-dimensional world coordinates of the marked points is verified. If the verification fails, the process will jump to the measurement frame synchronization acquisition step and re-execute the steps that started with the measurement frame synchronization acquisition step. If the verification passes, the process will proceed to the next step of the triangulation and 3D coordinate solution verification step. Steps for calculating and outputting distance between annotation points: Calculate the actual spatial distance between annotation points based on their 3D world coordinates and send the calculation results and the valid status of the annotation points back to the remote expert module for display; Augmented Reality Reflection Steps: The marked points and corresponding measurement results are overlaid and rendered in an augmented reality manner on the projection display screen of the head-mounted augmented reality device.
3. The remote three-dimensional dimension measurement and on-site marking method according to claim 2, characterized in that: When determining the measurement frame pair, the pixel displacement is estimated by the exposure time difference and compared with the preset pixel displacement threshold to exclude unsuitable measurement frame pairs. When an unsuitable measurement frame pair is found, the head-mounted augmented reality device at the site is prompted to re-acquire the data under the operation of the on-site personnel.
4. The remote three-dimensional dimension measurement and on-site marking method according to claim 2, characterized in that: The polar geometric constraint establishment step uses whole-image correction. The distortion model is applied to all pixels of the measurement frame image to remove distortion. The correction rotation matrix of each is calculated based on the fixed or time-compensated relative pose between the first and second optical cameras. Then, a remapping table of each coordinate is constructed. All pixels are traversed. The coordinates in the original image are retrieved according to the remapping table and resampled to generate the corrected images of each, so that the vertical coordinate axes of the two corrected images are parallel to each other.
5. The remote three-dimensional dimension measurement and on-site marking method according to claim 2, characterized in that: The polar geometry constraint establishment step uses marker points selected by remote experts and candidate pixels from the second optical camera to perform distortion correction using a distortion model and transform them to normalized observation rays in their respective coordinate systems. The pixels of both cameras are mapped to an ideal pinhole plane. Based on the relative rotation and translation between the first and second optical cameras, an essential matrix and a fundamental matrix are constructed. A straight epipolar line is determined on the ideal pinhole plane to satisfy the relationship between the essential matrix and the normalized observation ray. The epipolar line vector uses... It means that in the formula F The fundamental matrix, v Ai It is a marker point on the ideal pinhole plane of the first optical camera. i , l Bi It is a marker point i The corresponding polar line in the ideal pinhole plane of the second optical camera.
6. The remote three-dimensional dimension measurement and on-site marking method according to claim 2, characterized in that: The polar geometric constraint establishment step uses the labeled pixel selected by the remote expert to perform distortion removal using the distortion model of the first optical camera and transforms it into a normalized observation ray in the coordinate system of the first optical camera. Adaptive sampling is performed along the aforementioned observation ray and the three-dimensional points formed by different assumed depths, so that the Euclidean distance between image points at adjacent sampling depths is no greater than a preset pixel step size. The three-dimensional points are then projected onto the original second image through the projection model of the second optical camera and connected to form a curved finite polar line search trajectory.
7. The remote three-dimensional dimension measurement and on-site marking method according to claim 4, characterized in that: The method for determining the epipolar search region in the corresponding point search step is to perform a one-dimensional search within the parallax range along the horizontal epipolar line with the same row number as each labeled point in the second corrected image, and to expand the preset pixel width in the vertical direction to form an epipolar search band.
8. The remote three-dimensional dimension measurement and on-site marking method according to claim 5, characterized in that: The polar search region in the corresponding point search step is determined along the polar equation. A one-dimensional search is performed within the parallax range along the determined straight line direction, and a preset pixel width is extended along the epipolar normal direction to form an epipolar search band, where x and y are the horizontal and vertical coordinate components of the homogeneous pixel coordinates of the candidate point. These are the three components of the homogeneous coordinate expansion of the polar lines.
9. The remote three-dimensional dimension measurement and on-site marking method according to claim 6, characterized in that: The method for determining the epipolar search area in the corresponding point search step is to set a search band of preset width near the generated curved finite epipolar search trajectory.
10. The remote three-dimensional dimension measurement and on-site marking method according to claim 2, characterized in that: The local feature matching uses a coarse matching followed by fine matching. The coarse matching uses normalized cross-correlation matching cost for similarity measurement. After local feature matching, similarity filtering is also performed to exclude cases where candidate points and labeled points cannot match. The similarity filtering method is not limited to the discrimination filtering of best and second-best matches or left-right consistency verification. If a candidate point and labeled point cannot match, the process jumps to the measurement frame synchronization acquisition step and re-executes the steps that started the measurement frame synchronization acquisition step.
Citation Information
Patent Citations
An AR annotation setting method and a remote collaboration system
CN113936121B
Geometric length measurement method and system based on AR glasses
CN116188562A
Remote positioning method for nuclear reactor internals based on binocular vision
CN118887287A