An abnormality classification and supplementary acquisition guidance method for visual point cloud mapping
Patent Information
- Application Number
- CN202610906944.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-04
AI Technical Summary
[0006]本申请实施例提供了一种面向视觉点云建图的异常分类与补采引导方法、系统、计算机设备和计算机可读存储介质,以至少解决相关技术中缺乏空间化实时质量诊断能力的问题
[0018] Compared to related technologies, the passive binocular reconstruction anomaly classification and supplementary acquisition guidance method for visual point cloud mapping provided in this application addresses the technical shortcomings of traditional methods that cannot perform real-time spatial quality diagnosis and motion guidance. It avoids misjudging the geometric coverage output by active sensors as a visual mapping ready state. By transforming visual matching failure modes into concrete spatial patch-level anomaly classifications (such as distinguishing between low texture and reflective transparency), it outputs precise motion-level correction commands (such as lateral movement, angle adjustment, and supplementary lighting) to the user. This solution can improve the first-pass yield of front-end data acquisition, providing the back-end visual SLAM and point cloud fusion system with a data source possessing rich perspectives, stable features, and highly reliable pose constraints, thus achieving efficient construction of highly available visual maps.
Smart Images

Figure CN122695342A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of 3D reconstruction, and in particular to an anomaly classification and re-sampling guidance method, system, computer device and computer-readable storage medium for visual point cloud mapping. Background Technology
[0002] In applications such as indoor positioning, robot navigation, augmented reality (AR) scene reconstruction, smart wearable device positioning, and visual map construction, it is usually necessary to perform mapping and data acquisition on the target scene in advance. The existing conventional implementation involves an operator holding a smartphone, tablet, or other mobile terminal, walking along a designated route within the target area and continuously taking pictures. The system simultaneously acquires information such as images, video frames, timestamps, inertial measurement unit (IMU) data, and camera parameters. Subsequently, algorithms such as structure for motion reconstruction (SfM), visual simultaneous localization and mapping (SLAM), visual inertial odometry (VIO), feature matching, and point cloud fusion are used to generate a visual point cloud map or a visual positioning map.
[0003] In the aforementioned data flow process, the quality of front-end acquisition directly determines the accuracy of subsequent mapping output. If the target area suffers from issues such as a single shooting perspective, insufficient translational parallax, sparse image features, motion blur, severe surface reflection, transparent glass, incomplete trajectory loops, or unstable sensor time synchronization, it often leads to sparse point clouds, missing local areas, increased reprojection errors, trajectory drift, loop closure failures, or reduced localization and re-identification success rates. Currently, the specific technical deficiencies in related technologies manifest in the following aspects: First, there is a lack of spatial real-time quality diagnostic capabilities. In the existing acquisition process, the user interface only outputs camera previews or the overall acquisition progress percentage, making it difficult to determine in real time whether the visual data for specific spatial components (such as walls, corridors, desktops, corners, equipment surfaces, etc.) has sufficient conditions for mapping. When the backend algorithm throws a mapping failure exception, the operator is usually far from the acquisition site, leading to costly rework.
[0004] Second, there is a mismatch between the criteria for judging geometric coverage and visual mapping requirements. Some 3D scanning systems (such as active depth sensor devices based on Time-of-Flight (ToF), structured light, or LiDAR) can display the coverage of 3D point clouds or meshes during the acquisition process to indicate the completion of geometric scanning. However, for low-texture or non-diffuse reflection scenes such as white walls, repetitive textures, glass, and reflective surfaces, even if the geometric depth mesh is fully enclosed, subsequent visual feature matching and visual localization still have a high failure rate. Simply relying on the coverage of the active depth mesh can easily lead to the system misjudgment that "the geometric scan meets the requirements, but the visual mapping data is still missing."
[0005] Third, the evaluation metrics are abstract and lack action guidance. Some existing visual SLAM systems provide abstract quality scores based on the number or sharpness of feature points at the image level. This not only makes it difficult to accurately map low-scoring sources to 3D spatial patches, but also fails to provide operators with targeted corrective action suggestions. Due to the heterogeneity of the causes of insufficient data acquisition (such as insufficient parallax, single viewpoint, low texture, reflective transparency, limited lighting, etc.), existing systems cannot perform accurate anomaly classification, and therefore cannot guide users to implement correct supplementary data acquisition interventions. Summary of the Invention
[0006] This application provides an anomaly classification and re-acquisition guidance method, system, computer device, and computer-readable storage medium for visual point cloud mapping, so as to at least solve the problem of lack of spatial real-time quality diagnosis capability in related technologies.
[0007] In a first aspect, embodiments of this application provide an anomaly classification and supplementary acquisition guidance method for visual point cloud mapping, characterized in that it is applied to a visual mapping acquisition scenario on a mobile terminal, and the method includes: Simultaneously acquire the left and right grayscale images input from the passive binocular grayscale imaging component, as well as the RGB image and IMU data input from the mobile terminal; Local depth information is obtained by performing binocular reconstruction on the left grayscale image and the right grayscale image, and the target region is divided into a set of spatial patches in a unified coordinate system based on the local depth information. For each spatial patch, the binocular reconstruction anomaly index and the visual mapping index are calculated respectively, and the visual mapping acquisition quality of the spatial patch is evaluated based on the binocular reconstruction anomaly index and the visual mapping index. If the acquisition quality assessment result indicates that it is unqualified, the acquisition anomaly type corresponding to the spatial patch is obtained through classification diagnosis. Based on the anomaly type and the spatial location of the spatial patch, an action-level supplementary sampling prompt is generated in the user interface to guide the user to perform targeted supplementary sampling actions.
[0008] In some embodiments, binocular reconstruction is performed on the left grayscale image and the right grayscale image to obtain local depth information, and the target region is divided into multiple spatial patches, including: Epipolar correction and disparity matching are performed on the left grayscale image and the right grayscale image to generate a local depth map and a local mesh; The local depth map and the local mesh are divided into spatial patches, and the spatial patches corresponding to different image frames are merged into a unified coordinate system.
[0009] In some embodiments, for each spatial patch, the calculation of binocular reconstruction anomaly indices and visual mapping indices includes: The binocular matching confidence, left-right consistency failure rate, effective disparity ratio, depth hole ratio, cross-frame depth variance, and mesh fusion residual of the spatial patch are calculated as the binocular reconstruction anomaly indicators. The number of feature points, feature point distribution, feature point tracking length, reprojection error, image sharpness, exposure status, pose change corresponding to the IMU data, and VIO pose uncertainty in the RGB image are calculated as the visual mapping indicators.
[0010] In some embodiments, the classification diagnosis yields the following acquisition anomaly types corresponding to the spatial patch: Based on the binocular reconstruction anomaly index and the visual mapping index, the acquisition status is classified into parallax deficiency type, viewing angle deficiency type, low texture type, reflective transparency type, motion blur type, insufficient illumination type, and loop closure deficiency type.
[0011] In some embodiments, based on the anomaly type and the spatial location of the spatial patch, an action-level re-sampling prompt is generated in the user interface, including: When the number of feature points meets the preset conditions, the effective parallax ratio is lower than the first threshold, and the translation amount or triangulation angle of the mobile terminal is lower than the second threshold, the abnormality type is determined to be the insufficient parallax type, and a supplementary sampling prompt is generated to indicate lateral movement or increase the shooting baseline. When the number of feature points is lower than the third threshold, the confidence of binocular matching is lower than the fourth threshold, and the proportion of depth holes is higher than the first preset proportion, the abnormality type is determined to be the low texture type, and a supplementary sampling prompt is generated to indicate the acquisition of surrounding auxiliary reference objects or adjacent texture objects. When the left-right consistency failure rate is higher than the preset failure rate, the cross-frame depth variance or the mesh fusion residual is higher than the fifth threshold, and there are bright reflections or transparent areas in the RGB image, the anomaly type is determined to be the reflective transparency type, and a supplementary sampling prompt is generated to indicate changing the shooting angle or avoiding the reflective direction. When the quality of the local spatial patch is acceptable, and the pose uncertainty continues to increase while the acquisition trajectory is not closed, the anomaly type is determined to be the insufficient loop type, and a prompt is generated to indicate the return to the acquired area to re-acquire the marker.
[0012] In some embodiments, the classification diagnosis of the acquisition anomaly type corresponding to the spatial patch also includes: When the mobile terminal observes the same spatial patch at multiple observation times, it extracts the line connecting the camera optical center position and the patch center at each observation time as the observation vector, and records the local depth estimate and binocular matching confidence at each observation time to construct a multi-view observation set of the spatial patch. The maximum angle between the observation vectors of the spatial patch is calculated as the viewpoint divergence. When the viewpoint divergence exceeds the preset disparity baseline threshold, the multi-view depth weighted variance of the spatial patch is calculated based on the local depth estimate and the binocular matching confidence. Multidimensional feature decoupling is performed based on the multi-view depth-weighted variance and the binocular matching confidence to determine the anomaly type of the spatial patch.
[0013] In some embodiments, multidimensional feature decoupling is performed based on the multi-view depth-weighted variance and the binocular matching confidence to determine the final anomaly type of the spatial patch, including: If, in the multi-view observation set, the binocular matching confidence at the preset observation time is lower than the matching threshold, and the extreme value of the image gradient in the spatial patch area approaches zero, the anomaly type is determined to be low texture type, and an action-level re-sampling prompt for acquiring adjacent texture objects is output. If the confidence level of the binocular matching under a specific observation vector is higher than the preset matching threshold, and the local depth estimation value changes abruptly under different observation angles, causing the multi-view depth weighted variance to be greater than the fusion residual tolerance of the grid, the anomaly type is determined to be reflective transparency type, and an action-level supplementary sampling prompt to change the angle to avoid reflection is output.
[0014] In some embodiments, after guiding the user to perform the supplementary sampling action, the method further includes: The abnormal indicators and acquisition quality scores of the corresponding spatial patch after the supplementary acquisition are recalculated, and the acquisition status of the spatial patch is updated.
[0015] Secondly, embodiments of this application provide an anomaly classification and supplementary acquisition guidance system for visual point cloud mapping, applied to visual mapping acquisition scenarios on mobile terminals. The system includes: The acquisition module is used to simultaneously acquire the left grayscale image and right grayscale image input from the passive binocular grayscale imaging component, as well as the RGB image and IMU data input from the mobile terminal; The evaluation module is used to perform binocular reconstruction on the left grayscale image and the right grayscale image to obtain local depth information, and divide the target region into a set of spatial patches in a unified coordinate system based on the local depth information; and to calculate binocular reconstruction anomaly index and visual mapping index for each spatial patch, and to evaluate the visual mapping acquisition quality of the spatial patch based on the binocular reconstruction anomaly index and the visual mapping index. The anomaly diagnosis module is used to classify and diagnose the acquisition anomaly type corresponding to the spatial patch when the acquisition quality assessment result indicates that it is unqualified. The guidance prompt module is used to generate action-level supplementary sampling prompts in the user interface based on the anomaly type and the spatial location of the spatial patch, so as to guide the user to perform targeted supplementary sampling actions.
[0016] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect above.
[0017] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect above.
[0018] Compared to related technologies, the passive binocular reconstruction anomaly classification and supplementary acquisition guidance method for visual point cloud mapping provided in this application addresses the technical shortcomings of traditional methods that cannot perform real-time spatial quality diagnosis and motion guidance. It avoids misjudging the geometric coverage output by active sensors as a visual mapping ready state. By transforming visual matching failure modes into concrete spatial patch-level anomaly classifications (such as distinguishing between low texture and reflective transparency), it outputs precise motion-level correction commands (such as lateral movement, angle adjustment, and supplementary lighting) to the user. This solution can improve the first-pass yield of front-end data acquisition, providing the back-end visual SLAM and point cloud fusion system with a data source possessing rich perspectives, stable features, and highly reliable pose constraints, thus achieving efficient construction of highly available visual maps. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of a passive binocular reconstruction anomaly classification and supplementary sampling guidance method for visual point cloud mapping according to an embodiment of this application; Figure 2This is a structural block diagram of an anomaly classification and supplementary sampling guidance system for visual point cloud mapping according to an embodiment of this application; Figure 3 This is an architecture diagram of another visual point cloud mapping anomaly classification and supplementary sampling guidance system according to an embodiment of this application; Figure 4 This is a schematic diagram of the internal structure of a computer device according to an embodiment of this application. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.
[0021] Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without any inventive effort. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, any changes to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.
[0022] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0023] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following objects are in an "or" relationship. The terms "first," "second," and "third" used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.
[0024] The technologies described in this application can be used in 3D scene reconstruction systems, indoor space digitization platforms, augmented reality (AR) spatial map generation systems, robot navigation map building platforms, and visual positioning and surveying terminals, etc. The software system provided in this embodiment can be integrated into mobile terminals such as smartphones, tablets, handheld data collectors, and AR glasses, or deployed in edge computing nodes that communicate with the aforementioned mobile terminals. The mobile terminal is equipped with a passive binocular grayscale imaging component (including a left grayscale camera and a right grayscale camera). This component has various optional hardware device forms, specifically including a clamp-on binocular device external to the mobile terminal, an independent binocular camera connected via USB or Type-C interface, a wireless binocular device transmitting data via a wireless network protocol, a built-in binocular grayscale camera directly embedded in the mobile terminal's casing, or a binocular peripheral component integrated into a phone case or dedicated mounting bracket.
[0025] This embodiment provides a computer device. The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor may consist of one or more processors, including a central processing unit (CPU), an ASIC, or one or more integrated circuits configured to implement embodiments of this application. The memory may include a mass storage device for data or instructions, used to store or cache image frames, IMU sequences, depth grid data, and program instructions executed by the processor.
[0026] In this embodiment, the processor is configured to execute a passive binocular reconstruction anomaly classification and re-sampling guidance method for visual point cloud mapping. Figure 1 This is a flowchart of a passive binocular reconstruction anomaly classification and re-sampling guidance method for visual point cloud mapping according to an embodiment of this application, such as... Figure 1 As shown, the process includes the following steps: S101 synchronously acquires the left and right grayscale images input from the passive binocular grayscale imaging component, as well as the RGB image and IMU data input from the mobile terminal.
[0027] In this embodiment, when the acquisition process starts, the timestamps of the multi-source sensors are strictly aligned through the time synchronization module, and the input data stream is continuously read. Specifically, this includes the left grayscale image IL(t), the right grayscale image IR(t), the RGB image IC(t), the inertial measurement data M(t), and the corresponding timestamp T(t) at time t. Furthermore, the synchronization mechanism is set to use the timestamp of the RGB image frame as a reference, searching for matching binocular image pairs and IMU sequences within a preset time tolerance window, and performing interpolation compensation for fixed time deviations between different hardware sensors to ensure data consistency in spatial pose solving.
[0028] This step introduces a high-precision timestamp alignment and a multi-source heterogeneous sensor synchronization flow control mechanism at the data acquisition source, so that the image frames and inertial motion trajectories in the subsequent 3D spatial reconstruction have a deterministic correspondence in the spatiotemporal dimension, which can avoid trajectory drift and geometric distortion caused by data misalignment.
[0029] S102, perform binocular reconstruction on the left and right grayscale images to obtain local depth information, and divide the target region into a set of spatial patches in a unified coordinate system based on the local depth information.
[0030] Specifically, after acquiring the aligned image stream, the pre-calibrated camera intrinsic matrix K and binocular extrinsic matrix E are called to perform distortion correction and epipolar correction operations on the left and right grayscale images. Subsequently, a disparity matching algorithm (such as semi-global matching, block matching, or neural network disparity estimation) is used to generate a local depth map and the corresponding local mesh or local point cloud.
[0031] Furthermore, to achieve spatial state tracking, the local mesh is spatially discretized based on the generated local depth information, dividing it into multiple independent geometric units, defined as a set of spatial patches: Patch={p1,p2,...,pn}. Combined with the mobile terminal pose trajectory P(t) output by the VIO or SLAM system, the spatial patches extracted from different image frames are uniformly transformed and fused into the global coordinate system to eliminate redundant patches and maintain a unified three-dimensional spatial topology.
[0032] Step S102 transforms the local depth map and 3D mesh obtained by passive binocular reconstruction into discretized spatial patches in a unified coordinate system, discretizing the continuous 3D physical world into units that can be independently tracked and evaluated for quality, thus providing a geometric data basis for achieving refined real-time quality diagnosis.
[0033] S103. For each spatial patch, calculate the binocular reconstruction anomaly index and visual mapping index respectively, and evaluate the visual mapping acquisition quality of the spatial patch based on the binocular reconstruction anomaly index and visual mapping index.
[0034] In this embodiment, features are extracted bidirectionally from the image domain and the 3D geometric domain to construct a comprehensive quality assessment model. For each spatial patch p in the set, its binocular reconstruction anomaly index is calculated. Specifically, these include: binocular matching confidence, left-right consistency failure rate, effective disparity ratio, depth hole ratio, cross-frame depth variance, and mesh fusion residual.
[0035] Meanwhile, visual mapping metrics are extracted based on the corresponding RGB images, including: the number of feature points in the RGB image, the uniformity of feature point distribution, the cross-frame tracking length of feature points, the reprojection error, the image sharpness score, the exposure status, and the pose change and VIO pose uncertainty covariance calculated by combining IMU data.
[0036] Based on the extracted indicators, the visual mapping acquisition quality score Q(p) is calculated for the spatial patch p. The evaluation model is executed using the following weighted equation: Q(p)=w1·C(p)+w2·Rdisp(p)+w3·Vang(p)+w4·F(p)+w5·T(p)+w6·Pstab(p)-w7·H(p)-w8·Emesh(p)-w9·B(p) Where C(p) represents the binocular matching confidence, Rdisp(p) represents the effective disparity ratio, Vang(p) represents the coverage of the observation angle, F(p) represents the number and distribution of visual features, T(p) represents the feature tracking stability, and Pstab(p) represents the pose estimation stability; in the negative penalty term, H(p) is the depth hole ratio, Emesh(p) is the mesh fusion residual, and B(p) represents negative factors such as image blurring or exposure abnormalities; w1 to w9 are set as system calibration weight coefficients. When Q(p) is lower than the set quality pass line, the system determines that the acquisition quality evaluation result of the spatial patch is unqualified.
[0037] Step S103 breaks through the limitations of single-index evaluation by fusing passive binocular micro-geometric reconstruction metrics with macroscopic visual mapping features of RGB images.
[0038] S104. If the acquisition quality assessment result indicates that it is not qualified, the acquisition anomaly type corresponding to the spatial surface patch is obtained through classification diagnosis.
[0039] The process involves determining whether the data acquisition quality is acceptable based on a quality score and a preset threshold. For spatial patches marked as unacceptable, a multi-branch logic is used to classify them into specific anomaly categories based on multi-dimensional anomaly features. This classification logic avoids relying solely on the geometric coverage output of the active depth sensor; its core mechanism lies in identifying reconstruction failure modes caused by physical materials and motion states in passive binocular imaging.
[0040] The anomaly classification rules include the following decision branches: First, the parallax deficiency type: When the number of feature points meets the preset extraction conditions, but the effective parallax ratio within the patch is lower than the first threshold, and the translation amount or triangulation angle of the mobile terminal in this area is lower than the second threshold, it is determined to be the parallax deficiency type. This mode is usually caused by in-situ rotation scanning.
[0041] Secondly, the insufficient viewing angle type: when the patch is observed in multiple frames, but the normal angle distribution of all observation cameras is concentrated in a very small angle range, resulting in insufficient coverage of the three-dimensional surface normal, it is judged as the insufficient viewing angle type.
[0042] Third, low-texture type: When the number of RGB feature points is lower than the third threshold, the confidence of passive stereo matching is lower than the fourth threshold, and the proportion of holes in the generated depth map is higher than the first preset proportion, it is judged as low-texture type. This type of anomaly usually corresponds to large areas of white walls or textureless floor tiles.
[0043] Fourth, motion-blurred type: When the image clarity of multiple consecutive frames is lower than the threshold, and the IMU data indicates that the peak value of the mobile terminal's angular velocity or acceleration exceeds the stable boundary, it is judged as motion-blurred type.
[0044] Fifth, insufficient illumination type: When the overall brightness of the image is low, the noise ratio is high, the exposure is abnormal, and it is accompanied by the expansion of the binocular depth hole, it is classified as insufficient illumination type.
[0045] Sixth, insufficient loop closure type: The local spatial patch itself is of acceptable quality, but at the global scale, the VIO pose uncertainty continues to expand over time, and when the trajectory has never crossed or overlapped to close the loop, the output loop closure is abnormally insufficient.
[0046] In some embodiments, considering that the single-frame judgment method cannot accurately distinguish between "low-texture surfaces such as white walls" and "reflective transparent surfaces such as glass and other high-gloss metals," this embodiment further introduces a refined anomaly classification mechanism based on multi-view depth variance analysis. The specific execution steps are as follows: The first step is to construct a multi-view observation set. This is done when the mobile terminal observes the same spatial patch at multiple observation times t1, t2, ..., tk. At that time, the position of the camera optical center at each observation moment is extracted. The observation vector is constructed by connecting the center of the patch to the line. Simultaneously record the local depth estimates corresponding to each observation time. and binocular matching confidence .
[0047] The second step is to extract the viewpoint divergence and calculate the depth conditional variance. The maximum angle between the observation vectors of each spatial patch is calculated as the viewpoint divergence. .when When the parallax baseline threshold is exceeded, the multi-view depth-weighted variance is calculated based on the effective confidence level.
[0048]
[0049] in, This is a weighted average depth estimate. It is a multi-angle depth-weighted variance, used to reflect the waveform of the depth value calculated at the calculated angle. By configuring weights, if the confidence level at a certain angle is higher, the depth value at that angle will have a greater component in the formula, which will lead to an increase in variance and reflect that the depth value is changing drastically.
[0050] The third step is to perform multi-dimensional feature decoupling diagnosis. A non-linear rule is used for judgment: if the binocular matching confidence at preset observation times in the multi-view observation set is generally lower than the matching threshold, and the extreme values of the RGB image gradients within the patch area approach zero, it is confirmed that the patch has no physical features and is judged as low-texture type.
[0051] Conversely, if capturing spurious textures reflected in the environment under a specific observation vector leads to a confidence level higher than the matching threshold, but this confidence level varies with different observation angles... Changes in its local depth estimate The drastic jump caused the multi-view depth-weighted variance to exceed the tolerance of the mesh fusion residual. This phenomenon indicates that the system matched and locked the target as a reflective reflection or an object behind glass, thus confirming that its anomaly type is reflective transparency.
[0052] Step S104 performs multi-branch cascaded anomaly classification diagnosis on non-conforming spatial units, transforming the error into a concrete physical cause classification, providing diagnostic conclusions for subsequent output of targeted action-level correction instructions, and improving the precision of fault diagnosis.
[0053] S105 generates action-level supplementary sampling prompts in the user interface based on the anomaly type and the spatial location of the spatial patch, to guide the user to perform targeted supplementary sampling actions.
[0054] After obtaining the coordinates and specific anomaly types of the spatial facets, a three-dimensional semi-transparent spatial layer or spatial heatmap is rendered and overlaid in the screen display component to highlight the anomaly facets with a prominent color.
[0055] Furthermore, based on the anomaly type dictionary mapping, action-level guidance prompts are output: For parallax deficiency, a lateral movement indicator arrow is generated, or a text prompt is output to increase the shooting baseline and avoid in-situ rotation; for insufficient viewing angle, a prompt is generated requiring intervention from the side or oblique angle and changing the pitch angle for observation; for low texture, an instruction is output to guide the acquisition of auxiliary reference objects such as surrounding door frames, corners, and sockets, or adjacent objects with strong texture; for reflective transparency, a high-priority alarm is triggered, and an action prompt is generated to change the shooting angle or avoid the current reflective incident direction; for motion blur, feedback instructions are issued to decelerate, pause, or stabilize the holding posture; for insufficient lighting, a prompt is made to turn on the supplementary lighting equipment or move closer to the target area; for insufficient loop closure, a return trajectory map is generated, prompting the user to return to the already acquired area to shoot historical landmarks to form a closed loop constraint.
[0056] S106, recalculate the abnormal indicators and acquisition quality score of the corresponding spatial patch after the supplementary acquisition, and update the acquisition status of the spatial patch.
[0057] After the user performs supplementary data collection as prompted, incremental data is captured, and depth maps and feature point sequences are re-injected into the target spatial patch. The aforementioned steps are repeated to recalculate anomaly indicators and comprehensive quality scores.
[0058] It should be noted that this embodiment maintains a dynamic spatial state machine diagram, refreshing the state of each patch through transitions between "not acquired," "observed but of insufficient quality," "acquiring," "needs supplementary acquisition," and "acquisition is sufficient." Once the reassessment meets the standards, the warning marker is removed, and the region status is updated to acquisition complete. At the end of the data acquisition process, high-quality image sequences and IMU pose data packets, filtered using this method and including a spatial patch quality scoring matrix, anomaly classification logs, and supplementary acquisition enhancement correlation records, are output to the backend SFM, visual SLAM, or point cloud fusion server. The backend algorithm then performs frame filtering, adaptive weight adjustment, and mapping optimization calculations, significantly reducing the mapping crash rate caused by low frontend data quality.
[0059] Through steps S101 to S106, the technical shortcomings of traditional methods in real-time spatial quality diagnosis and motion guidance are addressed. This avoids misinterpreting the geometric coverage output by active sensors as a visual mapping readiness state. By transforming abstract visual matching failure patterns into concrete spatial patch-level anomaly classifications (such as distinguishing between low-texture and reflective transparency), precise motion-level correction commands (such as lateral movement, angle adjustment, and supplemental lighting) are output to the user. This mechanism significantly improves the first-pass yield of front-end data acquisition, providing the back-end visual SLAM and point cloud fusion systems with a data source possessing rich perspectives, stable features, and highly reliable pose constraints, thus enabling the efficient construction of highly available visual maps.
[0060] On the other hand, this embodiment also provides an anomaly classification and re-sampling guidance system for visual point cloud mapping. Figure 2 This is a structural block diagram of an anomaly classification and re-acquisition guidance system for visual point cloud mapping according to an embodiment of this application, such as... Figure 2 As shown, the system includes: a data acquisition module 20, an evaluation module 21, an anomaly diagnosis module 22, and a guidance and prompting module 23, wherein: The acquisition module 20 is used to simultaneously acquire the left grayscale image and right grayscale image input from the passive binocular grayscale imaging component, as well as the RGB image and IMU data input from the mobile terminal; Evaluation module 21 is used to perform binocular reconstruction on the left grayscale image and the right grayscale image to obtain local depth information, and divide the target region into a set of spatial patches in a unified coordinate system based on the local depth information; and to calculate binocular reconstruction anomaly index and visual mapping index for each spatial patch, and to evaluate the visual mapping acquisition quality of the spatial patch based on the binocular reconstruction anomaly index and the visual mapping index. Anomaly diagnosis module 22 is used to classify and diagnose the acquisition anomaly type corresponding to the spatial patch when the acquisition quality assessment result indicates that it is unqualified. The guidance prompt module 23 is used to generate an action-level supplementary sampling prompt in the user interface based on the anomaly type and the spatial position of the spatial patch, so as to guide the user to perform targeted supplementary sampling actions; also, Figure 3 This is an architecture diagram of another visual point cloud mapping anomaly classification and supplementary acquisition guidance system according to an embodiment of this application, such as... Figure 3 As shown, the system includes: The data acquisition synchronization module is used to synchronously acquire the left grayscale image and right grayscale image input from the passive binocular grayscale imaging component, as well as the RGB image and IMU data input from the mobile terminal, and to perform spatiotemporal alignment of the multi-source heterogeneous sensor data stream through a time synchronization mechanism. The spatial patch construction module is used to perform binocular reconstruction on the left and right grayscale images to obtain local depth information, and based on the local depth information, divide the target area into a set of spatial patches, a three-dimensional voxel mesh, or a local sub-map region under a unified coordinate system. The quality assessment module is used to combine micro-geometric indicators such as binocular matching confidence with macro-visual indicators such as the number of feature points to assess the visual mapping acquisition quality of each spatial patch, and does not use active depth coverage as the sole readiness criterion. The anomaly classification module is used to diagnose precise acquisition anomaly types such as insufficient parallax, insufficient viewing angle, low texture, reflective transparency, motion blur, or insufficient loop closure when the evaluation is unqualified through multi-branch cascaded logic. The interactive guidance module is used to overlay a highlight layer in the three-dimensional space layer of the user interface based on the type of acquisition anomaly and the spatial position of the spatial patch, and generate a corresponding action-level supplementary acquisition prompt. The state update module is used to dynamically transfer the spatial unit-level acquisition state machine after acquiring incremental data input by the user, and output the complete supplementary acquisition record, including anomaly history and quality score, to the backend mapping system for mapping optimization calculation.
[0061] Compared to the traditional blind data acquisition mode that relies solely on preview images or active depth sensor grid coverage, this system achieves multi-dimensional quality assessment of heterogeneous building components such as walls, corridors, and desktops by cross-domain fusion of passive binocular microscopic reconstruction indicators with macroscopic visual mapping feature points and odometry status. Furthermore, it decouples abstract algorithm failures into concrete physical anomaly categories such as insufficient parallax, reflective transparency, and low texture, transforming them into intuitive action-level supplementary acquisition prompts like lateral movement and angle changes. Combined with a dynamically updated spatial patch state machine and multi-dimensional data output for backend optimization systems, this significantly improves the first-pass yield of visual data acquisition for frontline operators in complex indoor, industrial, or civilian scenarios, eliminating rework costs caused by data loss. This provides crucial pre-construction data assurance for the robust construction and fusion convergence of subsequent high-precision, large-scale visual point cloud maps.
[0062] In one embodiment, Figure 4 This is a schematic diagram of the internal structure of a computer device according to an embodiment of this application. For example... Figure 4 As shown, a computer device is provided, which can be an in-vehicle computing unit or a cloud server. The computer device includes a processor, a network interface, internal memory, and non-volatile memory connected via an internal bus. The non-volatile memory stores the operating system, computer programs, and parameters related to a large language model. The processor provides computing power to execute deep neural network inference and control algorithm operations. When the processor executes the stored computer program, it implements an anomaly classification and supplementary sampling guidance method for visual point cloud mapping according to any of the above embodiments.
[0063] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0064] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by hardware related to computer program instructions. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc. The above embodiments only illustrate several implementation methods of this application. The descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent.
[0065] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for anomaly classification and re-sampling guidance in visual point cloud mapping, characterized in that, The method, applied to visual mapping and acquisition scenarios on mobile terminals, includes: Simultaneously acquire the left and right grayscale images input from the passive binocular grayscale imaging component, as well as the RGB image and IMU data input from the mobile terminal; Local depth information is obtained by performing binocular reconstruction on the left grayscale image and the right grayscale image, and the target region is divided into a set of spatial patches in a unified coordinate system based on the local depth information. For each spatial patch, the binocular reconstruction anomaly index and the visual mapping index are calculated respectively, and the visual mapping acquisition quality of the spatial patch is evaluated based on the binocular reconstruction anomaly index and the visual mapping index. If the acquisition quality assessment result indicates that it is unqualified, the acquisition anomaly type corresponding to the spatial patch is obtained through classification diagnosis. Based on the anomaly type and the spatial location of the spatial patch, an action-level supplementary sampling prompt is generated in the user interface to guide the user to perform targeted supplementary sampling actions.
2. The method according to claim 1, characterized in that, Binocular reconstruction is performed on the left and right grayscale images to obtain local depth information, and the target region is divided into multiple spatial patches, including: Epipolar correction and disparity matching are performed on the left grayscale image and the right grayscale image to generate a local depth map and a local mesh; The local depth map and the local mesh are divided into spatial patches, and the spatial patches corresponding to different image frames are merged into a unified coordinate system.
3. The method according to claim 1 or 2, characterized in that, For each spatial patch, the binocular reconstruction anomaly index and visual mapping index are calculated separately, including: The binocular matching confidence, left-right consistency failure rate, effective disparity ratio, depth hole ratio, cross-frame depth variance, and mesh fusion residual of the spatial patch are calculated as the binocular reconstruction anomaly indicators. The number of feature points, feature point distribution, feature point tracking length, reprojection error, image sharpness, exposure status, pose change corresponding to the IMU data, and VIO pose uncertainty in the RGB image are calculated as the visual mapping indicators.
4. The method according to claim 3, characterized in that, The classification and diagnosis yielded the following types of acquisition anomalies corresponding to the spatial patches: Based on the binocular reconstruction anomaly index and the visual mapping index, the acquisition status is classified into parallax deficiency type, viewing angle deficiency type, low texture type, reflective transparency type, motion blur type, insufficient illumination type, and loop closure deficiency type.
5. The method according to claim 4, characterized in that, Based on the anomaly type and the spatial location of the spatial patch, an action-level re-sampling prompt is generated in the user interface, including: When the number of feature points meets the preset conditions, the effective parallax ratio is lower than the first threshold, and the translation amount or triangulation angle of the mobile terminal is lower than the second threshold, the abnormality type is determined to be the insufficient parallax type, and a supplementary sampling prompt is generated to indicate lateral movement or increase the shooting baseline. When the number of feature points is lower than the third threshold, the confidence of binocular matching is lower than the fourth threshold, and the proportion of depth holes is higher than the first preset proportion, the abnormality type is determined to be the low texture type, and a supplementary sampling prompt is generated to indicate the acquisition of surrounding auxiliary reference objects or adjacent texture objects. When the left-right consistency failure rate is higher than the preset failure rate, the cross-frame depth variance or the mesh fusion residual is higher than the fifth threshold, and there are bright reflections or transparent areas in the RGB image, the anomaly type is determined to be the reflective transparency type, and a supplementary sampling prompt is generated to indicate changing the shooting angle or avoiding the reflective direction. When the quality of the local spatial patch is acceptable, and the pose uncertainty continues to increase while the acquisition trajectory is not closed, the anomaly type is determined to be the insufficient loop type, and a prompt is generated to indicate the return to the acquired area to re-acquire the marker.
6. The method according to claim 1, characterized in that, The classification and diagnosis also yielded the following types of acquisition anomalies corresponding to the spatial patches: When the mobile terminal observes the same spatial patch at multiple observation times, it extracts the line connecting the camera optical center position and the patch center at each observation time as the observation vector, and records the local depth estimate and binocular matching confidence at each observation time to construct a multi-view observation set of the spatial patch. The maximum angle between the observation vectors of the spatial patch is calculated as the viewpoint divergence. When the viewpoint divergence exceeds the preset disparity baseline threshold, the multi-view depth weighted variance of the spatial patch is calculated based on the local depth estimate and the binocular matching confidence. Multidimensional feature decoupling is performed based on the multi-view depth-weighted variance and the binocular matching confidence to determine the anomaly type of the spatial patch.
7. The method according to claim 6, characterized in that, Based on the multi-view depth-weighted variance and the binocular matching confidence, multi-dimensional feature decoupling is performed to determine the final anomaly type of the spatial patch, including: If, in the multi-view observation set, the binocular matching confidence at the preset observation time is lower than the matching threshold, and the extreme value of the image gradient in the spatial patch area approaches zero, the anomaly type is determined to be low texture type, and an action-level re-sampling prompt for acquiring adjacent texture objects is output. If the confidence level of the binocular matching under a specific observation vector is higher than the preset matching threshold, and the local depth estimation value changes abruptly under different observation angles, causing the multi-view depth weighted variance to be greater than the fusion residual tolerance of the grid, the anomaly type is determined to be reflective transparency type, and an action-level supplementary sampling prompt to change the angle to avoid reflection is output.
8. The method according to claim 1, characterized in that, After guiding the user to perform the supplementary sampling action, the method further includes: The abnormal indicators and acquisition quality scores of the corresponding spatial patch after the supplementary acquisition are recalculated, and the acquisition status of the spatial patch is updated.
9. An anomaly classification and re-acquisition guidance system for visual point cloud mapping, characterized in that, The system, applied to visual mapping and acquisition scenarios on mobile terminals, includes: The acquisition module is used to simultaneously acquire the left grayscale image and right grayscale image input from the passive binocular grayscale imaging component, as well as the RGB image and IMU data input from the mobile terminal; The evaluation module is used to perform binocular reconstruction on the left grayscale image and the right grayscale image to obtain local depth information, and divide the target region into a set of spatial patches in a unified coordinate system based on the local depth information; and to calculate binocular reconstruction anomaly index and visual mapping index for each spatial patch, and to evaluate the visual mapping acquisition quality of the spatial patch based on the binocular reconstruction anomaly index and the visual mapping index. The anomaly diagnosis module is used to classify and diagnose the acquisition anomaly type corresponding to the spatial patch when the acquisition quality assessment result indicates that it is unqualified. The guidance prompt module is used to generate action-level supplementary sampling prompts in the user interface based on the anomaly type and the spatial location of the spatial patch, so as to guide the user to perform targeted supplementary sampling actions.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 8.