A visual mapping system
By coupling the hardware and software of the visual mapping reference pad, and utilizing the micro-convex and concave structure and asymmetric marking sequence, the problem of separation between the physical splicing reference and the visual positioning reference is solved, realizing the accuracy and automation of large-scale mapping, eliminating cumulative errors, and improving the mapping accuracy and automation level.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MCGARVEY PROJECT MANAGEMENT SERVICES LLC
- Filing Date
- 2026-04-29
- Publication Date
- 2026-07-03
AI Technical Summary
In industrial maintenance sites, the separation of physical stitching reference and visual positioning reference in existing visual mapping systems leads to accumulated errors in large-scale mapping and difficulty in direction determination. Furthermore, the lack of standardized physical reference backgrounds affects mapping accuracy.
A visual mapping reference pad is used, including a matte coating with a micro-uneven structure and embedded ring-shaped positioning anchors and visually encoded marks with asymmetric identification sequences. Through the coupling of hardware and software, a unified physical and visual reference is established to eliminate accumulated errors and automatically determine the direction.
It achieves precision and automation in large-scale surveying, eliminates accumulated errors, improves surveying accuracy and automation, reduces interference from ambient light, and ensures the stability of physical scales and the standardization of surveying.
Smart Images

Figure CN122329264A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial visual surveying and recognition, and more particularly to a visual surveying system. Background Technology
[0002] Currently, in complex environments such as industrial maintenance (e.g., heavy truck repair), visual mapping and identification of parts typically rely on operators using handheld imaging devices. However, existing solutions have significant drawbacks: firstly, the lack of a standardized physical reference background means the image acquisition process lacks a unified physical reference system, preventing the backend algorithm from determining the true physical dimensions of the parts; secondly, the operator's non-perpendicular shooting angle leads to severe perspective distortion, causing the parts to appear deformed in the image, significantly impacting matching and mapping accuracy. Especially when mapping ultra-large parts, existing technologies often employ multi-pad stitching, but the physical stitching reference and visual positioning reference are often separate. This separation results in accumulated errors at the stitching points, and the algorithm struggles to quickly and automatically determine the absolute orientation of the image, requiring manual intervention for correction, severely hindering the standardization and automation of industrial visual mapping. Therefore, there is an urgent need to provide a visual mapping system to address these issues. Summary of the Invention
[0003] The purpose of this invention is to provide a visual mapping system to improve the problems of large-scale mapping cumulative errors and difficulty in direction determination caused by the separation of physical stitching reference and visual positioning reference in the prior art.
[0004] To achieve the above objectives, the present invention adopts the following technical solution: This invention provides a visual mapping system, comprising: a visual mapping reference pad and a processing module, wherein: The visual mapping reference pad includes a pad body; a surface layer disposed on the surface of the pad body; the surface layer is a matte coating with a microscopic uneven structure; a positioning anchor disposed on the pad body, the positioning anchor being used for positioning and alignment during splicing; a visual coding mark disposed on the surface layer; wherein the center of the visual coding mark coincides with the center of the positioning anchor in the vertical projection direction; and a multi-level size reference calibration area disposed on the surface layer. The processing module is configured to calculate the homography matrix by recognizing the visual coding marks of the visual mapping reference pad, reconstruct the orthophoto plane coordinate system, establish a mapping model between pixel coordinates and physical dimensions, and output the key geometric parameters of the part to be measured.
[0005] The above scheme couples the hardware visual mapping reference pad with the software processing module. The processing module directly establishes a physical coordinate system based on the visual coding marks on the reference pad, realizing the inseparable overall linkage between the hardware physical reference and the software correction algorithm, and achieving the technical effect of accurate mapping and recognition.
[0006] Furthermore, the above scheme achieves collaborative unification of physical stitching reference and visual positioning reference by aligning the center of the positioning anchor with the center of the visual coding mark in the vertical projection direction. When performing multi-pad stitching, the alignment of the physical positioning anchor and the visual coding mark helps to transform it into the coordinate system reconstruction reference of the visual mapping algorithm, thereby eliminating the cumulative error under large-scale mapping. At the same time, the surface layer is a matte coating with a micro-uneven structure, which can effectively absorb direct light, diffracted light and reflected light, helping to eliminate diffraction and reflection ghosting at the edges of the photographed object. The multi-level size reference calibration area provides a reliable physical conversion basis for the visual mapping algorithm.
[0007] As an optional implementation, the processing module is specifically configured as follows: The visually encoded markers are identified through image segmentation, and the corner points of the visually encoded markers are extracted to sub-pixel accuracy. Based on the known physical size ratio of the corner points of the visual coding mark to the visual mapping reference pad, the homography matrix is calculated, and the image area with perspective distortion is mapped to the orthophoto plane to reconstruct the orthophoto plane coordinate system. Based on the pixel spacing of the multi-level size reference calibration area on the visual mapping reference pad, a mapping model between pixel coordinates and physical dimensions is established. Based on the coordinates of the central positioning area of the visual mapping reference pad, the region of interest of the part to be measured is locked, and the key geometric parameters of the part to be measured are extracted within the region of interest.
[0008] As an optional implementation, after locating the visual coding mark, the processing module is specifically configured as follows: A gray-scale weighted centroid algorithm is used to extract the corner points of the visually encoded marker at the sub-pixel level. The gray-scale weighted centroid algorithm calculates the centroid coordinates by using the gray values of each pixel in the neighborhood of the corner point as weights, thereby controlling the positioning deviation of the corner points of the visually encoded marker within a set pixel value. The corner points of the visually encoded marker extracted by the gray-scale weighted centroid algorithm are used as the absolute base points for subsequent homography matrix calculation and orthophoto plane coordinate system reconstruction.
[0009] The visual mapping system also includes an image acquisition module. The image acquisition module calls its own attitude sensor data. When the angle between the plane of the image acquisition module and the plane of the visual mapping reference pad exceeds a preset threshold, the image acquisition module is configured to automatically disable the image acquisition function and output a visual guidance prompt.
[0010] As an optional implementation, the processing module is specifically configured as follows: Automatically scan the grid pixel spacing of the multi-level size reference calibration area to obtain the pixel distribution data of the surface grid of the visual mapping reference pad in the original image; The radial distortion of the image acquisition device is nonlinearly corrected in real time using grid curvature, where grid curvature reflects the degree of bending of the grid in the distorted image. The radial distortion parameters of the image acquisition device are calculated based on the distribution characteristics of the grid curvature and inverse compensation is performed. Based on the corrected grid pixel spacing, the mapping model is constructed as a mapping function with pixel coordinates as input and physical size as output. The mapping function is used to convert between pixel coordinates and physical size at any position on the visual mapping reference pad.
[0011] As an optional implementation, the positioning anchor is an embedded ring structure, and the visual coding mark is an asymmetric identifier sequence.
[0012] The above scheme provides a stable physical positioning connection point through an embedded ring structure, while the asymmetric label sequence gives the pose and orientation attributes of the visual coding mark, enabling the back-end visual mapping algorithm to automatically determine the absolute orientation of the photo in a very short time and complete the orthophoto correction without manual intervention.
[0013] As an optional implementation, the multi-level size reference calibration area includes at least one of the following: a subdivided grid, an edge scale, or a central concentric circle positioning area.
[0014] The above scheme provides a high-precision local pixel conversion reference through subdivided grid, provides macroscopic size calibration through edge scale, and forces the photographer to center the object in the center through the central concentric circle positioning area. The combination of these three elements provides a multi-level size reference for visual mapping algorithms, which significantly improves feature extraction efficiency and mapping accuracy.
[0015] As an optional implementation, the embedded ring structure is an embedded metal ring, and the asymmetric identifier sequence is an ArUco encoding array.
[0016] The above scheme ensures the physical accuracy of long-term splicing and positioning through the high strength and wear resistance of the metal ring, while the ArUco encoding array provides directional information for the detection of the visual mapping algorithm, ensuring the reliable operation of the hardware and software coupled system.
[0017] As an optional implementation, the gloss of the surface layer is less than or equal to 5 GU.
[0018] The above solution controls the gloss level to an extremely low level, which can absorb most of the direct light, fundamentally eliminating the diffraction and reflection ghosting of metal parts edges under strong light in industrial settings, and improving the accuracy of edge extraction by the algorithm.
[0019] As an optional implementation, the pad body includes a damping anti-slip layer and a fiber reinforcement layer covering the surface of the damping anti-slip layer.
[0020] The above solution uses a composite structure of a damping and anti-slip layer and a fiber reinforcement layer covering the surface of the damping and anti-slip layer, and then adds a surface layer. The damping and anti-slip layer ensures that the pad body does not shift with the ground when heavy metal parts are placed on it, the fiber reinforcement layer prevents physical deformation after long-term use or folding, and the matte coating with micro-concave and convex structure ensures the quality of visual acquisition. Together, they ensure the absolute stability and accuracy of the physical scale.
[0021] As an optional implementation, the surface of the bottom high-damping anti-slip layer is embossed with a cross-shaped anti-slip pattern.
[0022] The above solution further increases the frictional damping between the bottom layer and the contact surface through the cross-shaped anti-slip texture, which can ensure that the position of the reference pad is absolutely fixed even when heavy industrial parts weighing tens of kilograms are placed on it, thus ensuring the stability of the mapping reference.
[0023] As an optional implementation, the edge of the pad body is provided with a detachable splicing connection structure.
[0024] The above solution achieves modular and flexible splicing of the background pad through the combination of a detachable splicing connection structure and positioning anchors. The positioning anchors provide physical-level seamless alignment during splicing, enabling the system to adapt to the surveying needs of ultra-large parts, and the overall coordinate system remains unified after splicing.
[0025] In addition, the detachable splicing connection structure is Velcro, and the surface roughness Ra of the surface layer falls within [0.3μm, 3μm].
[0026] The beneficial effects of this invention are as follows: 1. Unified collaboration between physical and visual references to eliminate cumulative errors: By designing the center of the positioning anchor and the visual coding mark to coincide in the vertical projection direction, the positioning reference during physical splicing and the coordinate system reconstruction reference recognized by the visual algorithm are completely of the same origin. When performing multi-pad splicing to map ultra-large parts, the seamless alignment at the physical level is directly transformed into the unification of the algorithm coordinate system, which completely solves the problem of cumulative errors in large-size mapping caused by the separation of physical splicing and visual positioning in traditional solutions.
[0027] 2. Automatic determination of absolute direction, improving correction efficiency: Asymmetric label sequences are used as visual encoding markers, which give the markers themselves directional attributes. The backend visual mapping algorithm can automatically determine the absolute direction of the image in a very short time by recognizing the asymmetric features of the markers. Orthophoto correction can be completed without manual intervention, which greatly improves the automation level and processing efficiency of the system.
[0028] 3. Eliminate ambient light interference and improve edge extraction accuracy: The surface layer with micro-uneven structure (especially the matte coating with gloss controlled below 5GU) can absorb most of the direct light, effectively balancing the contrast between the highlights and dark edges of the metal parts, fundamentally eliminating diffraction and reflection ghosting at the edges of the photographed object, significantly reducing the interference of oil stains and reflections on the algorithm, and improving the accuracy of edge extraction.
[0029] 4. Composite anti-deformation structure to ensure physical scale stability: The pad adopts a three-layer composite structure with a bottom layer of high-damping anti-slip, a middle layer of fiber reinforcement, and a top layer of matte coating. This ensures that the pad does not shift or deform when heavy industrial parts are placed on it, and guarantees the absolute physical accuracy of multi-level size reference calibration areas (subdivision grid, scale, etc.), providing a reliable hardware foundation for pixel-physical size mapping.
[0030] 5. Achieve high-precision mapping at extremely low cost: Through the coupling system of hardware background pad and software correction algorithm, the collection environment standards of different regions are forcibly unified. No expensive 3D scanning equipment is required. Ordinary shooting equipment can be used to obtain physical data of parts with millimeter-level accuracy, which solves the most difficult problem of "environmental standardization" in industrial big data collection.
[0031] 6. Four-step algorithm pipeline for fully automated and accurate mapping: Through a four-step algorithm pipeline of sub-pixel corner extraction, homography matrix spatial reconstruction, dynamic pixel-physical size mapping compensation, and automatic feature extraction of region of interest, a fully automated processing flow from raw distorted image to millimeter-level physical parameter output is achieved. Orthorectification and size calculation can be completed without manual intervention, completely eliminating the influence of perspective distortion caused by human shooting angle and radial distortion deviation of different device lenses.
[0032] 7. Acquisition end attitude locking ensures standardized input data: By calling the device attitude sensor data through the acquisition module, when the angle between the acquisition device and the reference pad plane exceeds the preset threshold, the shutter is automatically disabled and a visual guidance reminder is given. This ensures that the perspective distortion of the input image is controlled within a very small range from the source, providing standardized high-quality input data for the back-end algorithm and fundamentally improving the overall surveying accuracy and stability of the system. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of the cross-sectional structure of the visual mapping reference pad according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the overall planar structure of the visual mapping reference pad according to an embodiment of the present invention; Figure 3 This is a schematic diagram showing the detailed structure of a local calibration area of the visual mapping reference pad according to an embodiment of the present invention; Figure 4 This is a schematic diagram comparing the visual correction and coordinate system reconstruction logic in an embodiment of the present invention; Figure 5 This is a cross-sectional schematic diagram of the modular splicing array and connection structure according to an embodiment of the present invention.
[0034] Figure 6 This is a schematic diagram of the structure of the visual mapping system according to an embodiment of the present invention.
[0035] Figure 7 This is a schematic diagram of the visual mapping method flow according to an embodiment of the present invention. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0038] Example 1: As Figure 1 , Figure 2 and Figure 3 As shown, this embodiment provides a visual mapping reference pad. The visual mapping reference pad includes: a pad body 10 (… Figure 1(as shown in the figure); surface layer 20, disposed on the surface of the pad body, the surface layer being a matte coating with a micro-uneven structure on the surface; positioning anchor 30, disposed on the pad body, the positioning anchor 30 being used for positioning and alignment during splicing; visual coding mark 40, disposed on the surface layer; the center of the positioning anchor 30 and the center of the visual coding mark coincide in the vertical projection direction; multi-level size reference calibration area 50, disposed on the surface layer.
[0039] Specifically, the base body 10 serves as the physical foundation for the entire reference pad, and its shape is typically rectangular or square planar to provide a sufficiently large background area for photographing industrial parts. However, it should be understood that the reference pad can also be designed as circular or other irregular shapes according to specific mapping requirements, with the aim of stably supporting subsequent functional layers and components. The surface layer 20 covers the top surface of the base body 10, and its core physical function is to effectively absorb and diffuse ambient direct light.
[0040] Under complex lighting conditions in industrial settings, especially when faced with strong reflections and diffraction phenomena that are easily generated on the surface of metal parts, this surface layer 20 can fundamentally eliminate the "ghosting" or highlight glare formed by these interfering lights in the image, thereby providing a clean, low-noise edge extraction background for the backend visual algorithm and ensuring the stability of the mapping accuracy.
[0041] Positioning anchors 30 are disposed on the pad body 10, typically located at the corners or key edge nodes of the pad body 10, serving as physical-level hard positioning and connection benchmarks. Visual coding marks 40 are printed or embedded in corresponding positions on the surface layer, serving as soft reference benchmarks for visual algorithm recognition and coordinate system construction. The core spatial relationship in this embodiment is that the center of the positioning anchor 30 and the center of the visual coding mark 40 coincide in the vertical projection direction. This coaxial coincidence design has profound physical and algorithmic synergy significance. In traditional solutions, the hardware anchors used for physical splicing and the coding patterns used for visual positioning are often set separately. This separation leads to a spatial offset between the alignment benchmark of physical splicing and the coordinate reconstruction benchmark of the visual algorithm. When multiple benchmark pads are spliced to map ultra-large parts, this offset will generate serious cumulative errors as the number of splices increases, ultimately causing the overall mapping data to fail. This embodiment achieves the homogeneity and unification of the physical splicing benchmark and the visual positioning benchmark by forcing the centers of the two to coincide in the vertical projection direction. This means that when the operator aligns multiple reference pads seamlessly on a physical level using the positioning anchor 30, this physical alignment relationship will be directly and without deviation converted into the coordinate system reconstruction reference in the visual algorithm, completely eliminating the problem of accumulated error in large-scale surveying scenarios.
[0042] It should be understood that, although Figure 2 and Figure 3 The image shows that the positioning anchor 30 is a ring-shaped metal structure and the visual coding mark 40 is a square array. However, in other embodiments, the positioning anchor can also be a columnar pin hole, a block-shaped buckle, or other structures that can provide physical positioning functions. The visual coding mark 40 can also be a circular code, a cross code, or other patterns that can be recognized by the algorithm and locate the center and pose direction. When the two satisfy the core spatial condition that the centers of the two coincide in the vertical projection direction, the surveying requirements can be met.
[0043] A multi-level size reference calibration area 50 is set on the surface layer 20. Its function is to provide the visual mapping algorithm with a multi-level physical size conversion basis from macro to micro. With a single reference calibration, the algorithm often struggles to simultaneously capture the overall contour of large-sized parts and accurately extract small-sized local features (such as minute apertures or fine tooth counts), easily leading to accuracy fluctuations at different mapping scales. This embodiment, by setting a multi-level size reference calibration area, enables the algorithm to consistently find a matching physical reference benchmark at different scaling ratios and feature scales, thereby achieving accurate pixel-physical size mapping across the entire scale range. It should be understood that the multi-level size reference calibration area is not limited to... Figure 3 The specific partitioning pattern shown can be a combination of various scale lines, grids, and geometric patterns with different precision, as long as it can provide at least two levels of size reference gradient.
[0044] Example 2: Based on Example 1, in order to further enhance the homogeneity and unity of the physical splicing reference and the visual positioning reference, and to give the system the ability to automatically determine the image pose direction, this example refines the design of the intermediate and bottom layers of the specific shapes of the positioning anchors and visual coding marks.
[0045] First, the positioning anchor 30 in this embodiment can be an embedded ring structure, with visual coding markings as an asymmetrical identification sequence. Specifically, compared to convex or surface-adhesive positioning elements, the embedded ring structure, with its structure embedded inside the pad, provides a more stable physical positioning connection point. Under frequent splicing and disassembly and the pressure of heavy industrial parts, it is less likely to fall off or shift position, thus ensuring the long-term stability of the physical splicing reference.
[0046] The visual coding mark 40 employs an asymmetrical mark sequence, and its core error-proofing logic lies in endowing the mark itself with inherent directional attributes. In actual industrial field photography, operators often find it difficult to ensure that each shot strictly adheres to a fixed up, down, left, and right orientation. If a symmetrical mark sequence (such as a common square or circular symmetrical pattern) is used, the algorithm cannot distinguish the rotational state of the pattern during recognition, which can easily lead to misjudgment of orientation. For example, it might mistake an inverted part image for an upright one, thus rendering subsequent dimensioning and feature matching completely ineffective. The visual coding mark 40, however, uses an asymmetrical mark sequence, which breaks this rotational symmetry. This allows the algorithm to automatically and uniquely determine the absolute orientation of the photo within a very short time (e.g., within 0.1 seconds) simply by considering the asymmetrical feature distribution of the sequence pattern upon mark recognition, without any manual intervention, thus completing the preparation for orthophoto correction. It should be understood that the asymmetrical mark sequence is not limited to a specific geometric arrangement. When the pattern cannot coincide with its original shape after being rotated by 90, 180, or 270 degrees, the error-proofing requirement for orientation determination can be met.
[0047] This embodiment also includes: the embedded ring structure is an embedded metal ring, such as an embedded copper ring, and the asymmetric identifier sequence can be an ArUco (Augmented Reality University of Cordoba) coding array. For example, the choice of a copper ring as the concrete realization of the embedded ring structure is based on the high strength and excellent wear resistance of copper in industrial environments. It can withstand repeated insertion and removal of locating pins and lateral shear forces when heavy components are placed, ensuring that the physical anchor points do not deform or wear during long-term use. As a mature visual benchmark marker library, the ArUco coding array has extremely high corner detection robustness and fast decoding capabilities. Even in complex industrial backgrounds with oil stains or uneven lighting, the algorithm can still stably extract corner features.
[0048] It is worth noting that, such as Figure 4 and Figure 5As shown, when the embedded copper ring and ArUco encoding array are used in combination, the design that their centers coincide in the vertical projection direction plays an important role in splicing alignment: during multi-pad modular splicing, the operator achieves a seamless physical hard connection by passing the positioning pin through the copper ring. At this time, the physical alignment center of the copper ring is directly equivalent to the visual coordinate reconstruction center of the ArUco encoding array. The accuracy of the physical splicing is transferred to the coordinate system of the visual algorithm without loss, completely eliminating the cumulative error caused by the separation of soft and hard references. As a specific example of an asymmetric sequence, the ArUco encoding array at the four corners in this embodiment can use a combination of ID 10 to ID 13. These four different IDs are uniquely distributed at the four corners of the pad. The algorithm can instantly lock the absolute position of the pad and the four corner points by recognizing these four different IDs. However, it should be understood that ID 10-13 is only a preferred example of asymmetric error prevention logic and is by no means a limitation of the present invention. Any ID combination or encoding format that can provide asymmetric directional guidance should fall within the protection scope of the present invention.
[0049] Example 3: Based on Example 1, this example further refines the internal functional partitioning of the multi-level size reference calibration area 50. Specifically, the multi-level size reference calibration area 50 includes at least one of the following: a subdivided grid, an edge scale, or a central concentric circle positioning area.
[0050] The three partitions—the subdivision mesh, the edge scale, and the central concentric circle positioning area—are not isolated visual elements but rather constitute an inseparable whole, with a rigorous combination and synergistic logic among them. The subdivision mesh covers most or all of the surface layer 20, its core function being to provide a high-precision local pixel conversion reference for the visual algorithm. When the algorithm needs to identify minute local features of a part, such as small thread diameters or dense bolt hole spacing, the subdivision mesh provides dense and uniform pixel anchor points, resulting in extremely high resolution and stability in the pixel-physical size mapping at the microscale. The edge scale is set at the edge of the pad, complementing the subdivision mesh in both macroscopic and microscopic aspects, providing macroscopic dimensional calibration for the algorithm. When measuring the overall outline dimensions of a large part, the edge scale can span a larger physical span, providing a long-distance reference scale for the algorithm and avoiding error amplification caused by relying solely on local mesh accumulation calculations. The central concentric circle positioning area is set at the geometric center of the pad, its core mechanism being to force the operator to place the part to be measured in the center through the geometric constraint property of the circle. Centering the parts not only maximizes the use of the effective visual acquisition area of the pad and reduces the impact of edge perspective distortion, but more importantly, the centered position allows the algorithm to obtain the most complete contour data when extracting the overall features of the parts, significantly improving the efficiency and completeness of feature extraction.
[0051] It must be emphasized why these three zones must exist in combination. Without a finer mesh, the algorithm loses its local high-precision physical reference when dealing with minute features, resulting in insufficient resolution and a significant drop in accuracy for microscopic mapping. Without an edge scale, the algorithm can only rely on the step-by-step accumulation of local meshes when calculating large contours. This accumulation is prone to numerical drift and cumulative errors over long distances, rendering macroscopic dimension calibration ineffective. Without a central concentric circle positioning area, operators are prone to arbitrarily offset placement of parts, causing parts of the part's contour to extend beyond the effective acquisition area of the pad or be located in areas of strong perspective distortion at the edges. The algorithm not only struggles to extract features completely but may also experience severe perspective distortion due to the part's deviation from the center, making subsequent orthorectification and dimension mapping extremely difficult. Therefore, the combination of these three zones is indispensable, jointly constructing a comprehensive collaborative defense system from micro to macro, from positioning to calibration.
[0052] It should be understood that this invention does not absolutely limit the specific parameters and shapes of the above-mentioned partitions, but can be flexibly adjusted according to actual surveying needs. For example, the size of the subdivision grid is not limited to 10mm. In scenarios requiring higher surveying accuracy for small parts, it can be a 5mm or even smaller grid, while in scenarios requiring only coarse calibration for large parts, it can be a 20mm grid. The range of the edge scale is not limited to 0-30cm, and can be extended to 0-50cm or longer depending on the overall size of the pad. The central concentric circle positioning area is not limited to a double-circle structure. To provide stronger centering guidance constraints, it can be a triple-circle structure, or even have cross lines superimposed inside the concentric circles, which can force the part to be centered through geometric symmetry. All of the above alternative solutions should fall within the protection scope of the multi-level size reference calibration area of this invention. Figure 2 As shown, this embodiment demonstrates a preferred layout of the subdivided grid, edge scale, and central concentric circle positioning area. This layout has demonstrated optimal collaborative mapping accuracy and feature extraction efficiency in long-term industrial field testing.
[0053] Example 4: Based on Example 1, in order to further enhance the ability to eliminate reflective interference in industrial settings, this example specifies the material and key optical parameters of the surface layer. Specifically, the surface layer is a matte coating with a micro-uneven structure, the gloss of the matte coating is less than or equal to 5 GU, and the surface roughness Ra of the surface layer falls within the range of [0.3 μm, 3 μm].
[0054] The microscopic physical mechanism of the matte coating lies in its dense nanoscale micropores and undulating particle structure. When strong light from an industrial environment (such as direct light from overhead lights or flashlight illumination) shines on the surface of the mat, these nanoscale structures disperse the direct light beams, which would otherwise form specular reflections, into countless diffusely reflected microbeams that scatter uniformly in all directions. Through this microscopic optical dispersion and absorption mechanism, the nanoscale diffusely reflected matte coating can absorb more than 90% of the direct light, greatly reducing the intensity of reflected light entering the camera lens. This fundamentally blocks the optical path that causes highlight glare and diffraction ghosting at the edges of metal parts in the image, thus providing a clean, low-noise background for edge extraction in the backend visual algorithm and ensuring the stability of mapping accuracy.
[0055] It is worth noting that the parameter limit of controlling the gloss level below 5 GU is a critical threshold for resolving reflections and ghosting, determined by researchers through extensive industrial field experiments. Gloss level, or gloss unit, is a quantitative indicator that measures the ability of a coating surface to reflect light. In this embodiment, specific examples of gloss levels include an upper limit of 5 GU, a middle value of 3 GU, and a lower limit of 1 GU. When the gloss level is within these ranges, the coating can effectively suppress diffraction phenomena under strong industrial light. However, if the gloss level exceeds the critical threshold of 5 GU, the anti-reflective ability of the coating will experience a sudden and drastic decrease. As an example of out-of-bounds comparison, if a matte coating with a gloss level of 8 GU or 10 GU (such as commercially available matte paint or ordinary silicone pads) is used, although it may not appear to be very reflective to the naked eye, in industrial high-light environments, the edges of metal parts will still produce obvious diffraction and reflection ghosting. These ghostings appear as bright noise and false edges in the image, causing severe noise in the algorithm when performing edge extraction and contour recognition. It is very easy to misjudge the ghostings as the real physical edges of the parts, resulting in millimeter-level or even centimeter-level deviations in the mapping results. Therefore, 5 GU is not an arbitrarily chosen empirical value, but a rigid optical barrier to ensure stable operation of the algorithm under complex lighting conditions.
[0056] Furthermore, as a preferred example, the color value of the matte coating can be selected as industrial gray (such as Cool Gray 10C). This specific color value has been experimentally verified to maximize the balance of edge contrast between the highlight areas of the metal parts and the shadow areas of the dark rubber parts. It avoids obscuring the outline of the light-colored metal parts due to an overly bright background, nor makes the edges of the dark parts difficult to discern due to an overly dark background. However, it should be understood that Cool Gray 10C is merely a preferred color configuration example, designed to improve contrast adaptability in specific industrial scenarios, and is by no means a limitation on the scope of protection of this invention. In other embodiments, depending on the typical color characteristics of the parts under test, the coating color value can also be adjusted to other neutral gray levels or specific color schemes, as long as the core optical parameter of gloss control below 5 GU is met.
[0057] Example 5: Based on Example 1, this example further refines the internal composite structure of the pad body 10 and the details of the bottom anti-slip layer to establish a defensive depth with two layers of composite functions. Specifically, as follows... Figure 1 As shown, the pad body 10 includes a bottom damping and anti-slip layer 101 and a middle fiber reinforcement layer 102, combined with a top surface layer 20, forming a three-layer composite structure. These three layers are not simply a stack of materials, but rather constitute an inseparable composite anti-deformation and anti-displacement whole. The bottom high-damping anti-slip layer directly contacts the ground or workbench; its high damping physical properties ensure that the pad will not experience any relative displacement with the contact surface when heavy industrial metal parts are placed on it, providing the most basic physical stability for the entire surveying process. The middle fiber reinforcement layer is embedded between the bottom and top layers; it has a high tensile modulus, and its core function is to prevent irreversible physical deformation of the pad after long-term use or repeated folding, thereby ensuring the absolute accuracy of the grid and scale in the multi-level dimensional reference calibration area on the surface. The top surface layer 20 directly faces the visual acquisition equipment, undertaking the key tasks of absorbing direct light, eliminating ghosting interference, and ensuring the quality of visual acquisition. The indivisibility of these three composite layers is reflected in the deep coupling of their functions: if the middle fiber reinforcement layer is missing, relying solely on the soft material combination of the bottom and top layers, the surface subdivision grid (e.g., a 10mm reference grid) of the reference pad will undergo irreversible tensile or compressive physical deformation after long-term rolling and storage. This deformation will directly cause the mapping model between pixels and physical dimensions in the visual algorithm to completely fail, because the physical reference datum on which the algorithm relies has undergone unpredictable distortion. If the bottom high-damping anti-slip layer is missing, even a small displacement when placing heavy objects will cause the coordinate system to shift between two consecutive shots, making orthorectification meaningless. Therefore, the composite structure of the reference pad is the cornerstone for ensuring the stability of the entire chain from physical placement to visual extraction of the mapping datum.
[0058] Furthermore, to enhance the anti-slip damping effect of the bottom layer, a cross-shaped anti-slip pattern is embossed on the surface of the high-damping anti-slip layer. The physical mechanism of the cross-shaped anti-slip pattern lies in the fact that by forming a dense cross-shaped raised and recessed structure on the bottom surface, the friction damping coefficient between the bottom layer and a smooth ground or workbench is greatly increased. When heavy industrial parts are placed on the pad, these cross-shaped patterns can firmly grip the ground like miniature claws, effectively resisting the lateral slippage tendency caused by the weight of the parts. For example, when placing a heavy truck brake drum or pressure plate weighing 20 to 50 kg, even if the operator pushes it forcefully, the position of the pad can remain absolutely fixed, thereby ensuring the absolute stability of the spatial coordinates of the visual coding mark and the multi-level size reference calibration area during the shooting process.
[0059] It should be understood that the shape of the anti-slip texture is not limited to a cross shape; it can also be a rhombus, a wave shape, or other geometric patterns that can increase surface roughness. However, after extensive testing on heavy-duty components, the cross-shaped texture has been shown to have the most uniform stress distribution and provide consistent frictional damping in all directions when subjected to the pressure and lateral shear forces of heavy components, exhibiting the best performance. Therefore, this embodiment discloses it as a preferred solution. However, this is by no means an absolute limitation on the shape of the anti-slip texture; any surface texture that can achieve an equivalent increase in frictional damping should fall within the protection scope of this invention. Figure 5 The AA edge connection structure cross section clearly shows the three-layer composite structure of the bottom high-damping anti-slip layer and its surface cross anti-slip texture, the middle fiber reinforcement layer and the top surface layer. This structure provides a physical bearing basis for the surveying of heavy parts in industrial sites.
[0060] Example 6: Based on Example 1, in order to enable the visual mapping reference pad to adapt to the mapping needs of ultra-large industrial parts, this example further refines the modular splicing and expansion capability of the reference pad. Specifically, the edge of the reference pad is provided with a detachable splicing connection structure.
[0061] The detachable splicing connection structure is located around the perimeter or on specific sides of the pad. Its core function is to allow operators to quickly connect multiple reference pads at physical boundaries to create a larger shooting background when the area of a single pad is insufficient to cover an oversized part under test. However, simple edge physical connection only achieves stacking in terms of area. Without precise positioning references, relative misalignment between multiple pads is very likely to occur. Therefore, this embodiment deeply couples the splicing connection structure with the positioning anchor: when multiple pads are spliced, the operator first achieves preliminary physical connection between adjacent pads through the detachable splicing connection structure at the edges, and then achieves seamless and precise physical alignment through the positioning anchor (e.g., with a positioning pin passing through an embedded copper ring). This process fully utilizes the core spatial relationship established in Embodiments 1 and 2—the center of the positioning anchor coincides with the center of the visual coding mark in the vertical projection direction. Due to this coaxial overlapping design, when two or more pads are physically aligned seamlessly through positioning anchors, this absolute physical alignment is directly and effortlessly translated into a unified coordinate system in the visual algorithm. The visually encoded marks on the spliced pads will appear in the same seamlessly connected orthophoto plane coordinate system to the algorithm, completely eliminating the large-scale measurement errors caused by physical gaps and visual coordinate misalignment in traditional multi-pad splicing.
[0062] It should be understood that the specific implementation forms of detachable splicing connection structures are diverse and not limited to... Figure 5The cross-section shows the hook and loop fasteners. For example, in scenarios requiring faster assembly and disassembly, the connection structure can be a magnetic structure, achieving automatic adsorption and bonding of multiple pads through strong magnetic strips embedded in the edges of the pads; in scenarios requiring higher connection strength, the connection structure can also be a snap-on structure or a zipper structure. The key is to achieve a detachable physical bond between the edges of adjacent pads.
[0063] like Figure 5 The diagram illustrates a preferred embodiment of the modular splicing array and connection structure cross-section. In this embodiment, reference pad A (shown as pad fabric in the diagram) and reference pad B (shown as pad fabric in the diagram) achieve a flat, high-strength physical connection through embedded snap fasteners at the edges. Simultaneously, the embedded copper metal rings at the four corners, in conjunction with positioning pins, ensure absolute uniformity of the overall visual coordinate system after splicing, providing a solid hardware foundation for 1:1 millimeter-level precise visual mapping of ultra-large components such as heavy-duty truck gearboxes.
[0064] Example 7: As Figure 6 As shown, this embodiment provides a visual mapping system. The visual mapping system includes: a visual mapping reference pad as described in any of the foregoing embodiments of the present invention, and a processing module.
[0065] The visual mapping reference pad, serving as the hardware carrier and physical foundation of the entire system, has had its specific structure and spatial layout described in detail in Examples 1 to 6, and will not be repeated here. The processing module, as the software and algorithm hub of the system, has the following core functions: calculating the homography matrix by recognizing the visual encoding marks on the visual mapping reference pad, reconstructing the orthophoto plane coordinate system, establishing a mapping model between pixel coordinates and physical dimensions, and outputting the key geometric parameters of the part to be measured.
[0066] Specifically, this embodiment emphasizes the inseparable hardware and software coupling between the visual mapping reference pad and the processing module. Traditional industrial visual mapping solutions often rely solely on pure software algorithms to forcibly extract features from a chaotic field background, or simply use ordinary background boards without any physical reference definition. This results in algorithms lacking a unified reference origin and scale basis, making them highly susceptible to environmental noise and outputting unreliable mapping data. In the system of this invention, the hardware reference pad is not merely a simple shooting backdrop. It eliminates reflective ghosting interference through a low-reflection matte surface layer, provides a standardized physical reference background through multi-level dimensional reference calibration areas, and, more importantly, establishes an absolutely precise mapping bridge between physical space and visual image space through the coaxial coincidence design of the positioning anchors and visual encoding marks. During operation, the processing module does not blindly address within a void pixel array, but directly uses the visual encoding marks on the reference pad as anchor points to instantly lock the physical origin and coordinate axis directions, thereby establishing a physical coordinate system with millimeter-level precision. The hardware reference pad provides an indispensable standardized input source and absolute spatial constraints for the software algorithm, while the software processing module endows the hardware markers with the vitality of dynamic calculation and coordinate reconstruction. Both are indispensable. Without the visual anchor points and low-noise background provided by the hardware reference pad, the coordinate system establishment of the processing module will lose its physical foundation, making accurate mapping impossible. Without the coordinate reconstruction calculations of the processing module, the markers and grids on the hardware reference pad are merely static patterns, unable to automatically output the true physical parameters of the parts.
[0067] It should be understood that the processing module can receive the data stream uploaded by the image acquisition module through a standardized application programming interface (API). The specific hardware implementation of the processing module is highly diverse. It can be a high-performance computing cluster deployed on a cloud server, receiving image data transmitted from the front-end imaging device via a wireless network for remote coordinate reconstruction and size calculation; it can also be a local processor integrated into an industrial tablet or smartphone terminal, utilizing the device's own computing power to instantly establish the physical coordinate system and output the mapping on-site; or it can even be a separately designed dedicated embedded vision processing box, directly connected to the imaging device via a wired interface. Regardless of the hardware form, as long as it is configured to establish a physical coordinate system based on the visual coding marks on the visual mapping reference pad described in this invention, it should fall within the protection scope of the visual mapping system of this invention. Details regarding the specific homography matrix calculation, orthophoto plane reconstruction, and pixel mapping algorithms within the processing module will be further elaborated in subsequent embodiments.
[0068] Example 8: Based on Example 7, this example further refines the internal algorithm logic of the processing module. Specifically, the processing module is configured to execute the following four-step processing flow: The first step is sub-pixel level corner extraction and preprocessing. The processing module identifies the visually encoded markers through image segmentation and extracts the corner points of the visually encoded markers to sub-pixel accuracy. Combined with... Figure 4 The left-right comparison diagram illustrates that in the original image, due to the tilt of the operator's handheld shooting angle, the visually encoded marker appears as an irregular quadrilateral with a smaller distance and a larger distance. The processing module first locates the region of the visually encoded marker in the complex background using image segmentation technology, and then extracts the four corner points of the marker, improving the corner point localization accuracy to the sub-pixel level (within 0.1 pixel deviation), which serves as the absolute base point for subsequent spatial reconstruction. The significance of sub-pixel level corner point extraction is that the corner points, as input data for homography matrix calculation, directly determine the upper limit of accuracy for all subsequent spatial reconstruction and size mapping. If corner point localization only stays at the pixel level (integer pixels), the optimal accuracy of subsequent calculations is also locked at a granularity of 1 pixel, making it impossible to achieve millimeter-level physical mapping.
[0069] The second step is spatial reconstruction driven by the homography matrix. Based on the known physical size ratio of the corner points extracted in the first step and the visual mapping reference pad, the processing module calculates the homography matrix, mapping the image region with perspective distortion to the orthophoto plane to reconstruct the orthophoto plane coordinate system. The specific execution process of the algorithm is as follows: The processing module detects four corner points of the visually encoded markers in the original image, labeled P1, P2, P3, and P4 respectively. These four corner points are the physical projection positions of the visually encoded markers in the distorted image. Based on the detected four corner points P1-P4, combined with the known physical aspect ratio parameters of the visual mapping reference pad at the time of manufacture, the processing module calculates the homography matrix H. This matrix is a 3×3 transformation matrix, and its mathematical essence is to describe the perspective mapping relationship between two planes. The processing module uses the calculated homography matrix... Matrix H maps the perspective-distorted corner points P1-P4 to the standard corner points P1', P2', P3', and P4' of the orthographic view. These mapped corner points form a perfect standard rectangle, completely eliminating the "smaller objects appear larger than closer objects" perspective error caused by the non-perpendicular shooting posture. Based on the reconstructed standard corner points P1'-P4', the processing module reconstructs a millimeter-level physical coordinate system, establishing mutually perpendicular X and Y axes on the orthographic plane and anchoring the origin to the marked geometric center, providing an absolute spatial reference for subsequent dimensional measurements. Even if the shooting angle deviates from the vertical axis (by tens of degrees), the algorithm can force the distorted quadrilateral region back to the orthographic projection plane through linear transformation of the homography matrix, completely eliminating perspective errors.
[0070] The third step is dynamic pixel-to-physical size mapping compensation. The processing module establishes a mapping model between pixel coordinates and physical dimensions based on the pixel spacing of the multi-level dimensional reference calibration area on the visual mapping reference pad. The module automatically scans the pixel spacing of the subdivided grid in the multi-level dimensional reference calibration area on the visual mapping reference pad, and combines this with the physical parameters known at the time of manufacture (e.g., the actual physical distance between the centers of adjacent marks is precisely 1000 mm), establishing a mapping model between pixel coordinates and physical dimensions through proportional conversion. For example, if the algorithm measures the pixel distance between two marks in the orthographic view as 2000px, and the actual physical distance is known to be 1000mm, then the mapping model is 1px = 0.5mm. Afterward, the algorithm only needs to extract the pixel coordinates and pixel span of the part's edge in the orthographic view, directly multiply it by the conversion factor, and automatically and accurately output the core physical parameters of the part, such as the thread diameter and bolt hole spacing, without any manual intervention in reading the data.
[0071] The fourth step involves automatic feature extraction based on the Region of Interest (ROI). The processing module locks the ROI of the part under test based on the coordinates of the central positioning area of the visual mapping reference pad, and extracts the key geometric parameters of the part within the ROI. The algorithm automatically locks the area where the part body is located based on the coordinates of the central positioning area (such as a concentric circle positioning area), shields the area from ambient noise outside the pad, and automatically identifies and labels the key geometric parameters of the part (such as bolt hole center distance, gear module, maximum outer diameter, etc.) within the locked area. Combining this with a mapping model, these parameters are converted from the pixel domain to the physical dimension domain, outputting millimeter-level physical mapping data.
[0072] It should be understood that the corner detection and homography matrix calculation in the above four-step algorithm process are illustrated using visual encoding markers as a preferred example. However, in other embodiments, the visual encoding markers can also be other asymmetric identifier sequences with corner extraction characteristics, as long as they can provide at least four non-collinear feature points for calculating the homography matrix. This invention does not impose an absolute limitation on the specific algorithm library type. More specific algorithm details in each step (such as dynamic threshold segmentation parameters, the specific calculation process of the gray-scale weighted centroid algorithm, and the grid curvature nonlinear correction method, etc.) will be further elaborated in subsequent embodiments.
[0073] Example 9: Based on Example 8, this example refines the key steps in the four-step algorithm flow of the processing module at a lower level.
[0074] First, regarding the underlying refinement of sub-pixel-level corner extraction: Specifically, the processing module uses dynamic thresholding to locate visually encoded markers in complex backgrounds containing oil contamination. In industrial maintenance sites, the surface of the background mat is often covered with contaminants such as oil and dust. These contaminants appear as irregular dark and bright spots in the image, severely interfering with marker detection. The core mechanism of dynamic thresholding is that the processing module adaptively adjusts the segmentation threshold based on the grayscale distribution of local regions in the original image, rather than using a globally fixed threshold. In well-lit areas, the local threshold automatically increases to avoid misclassifying reflective areas as markers; in shadowed areas, the local threshold automatically decreases to ensure that dark markers are not missed. This adaptive mechanism ensures stable detection of visually encoded markers under various industrial lighting conditions.
[0075] After locating the visually encoded markers, the processing module employs a grayscale weighted centroid algorithm to refine the corner points of the markers at the sub-pixel level. The specific calculation process of the grayscale weighted centroid algorithm is as follows: For each initially detected pixel-level corner point, the processing module extracts the grayscale values of all pixels within its neighborhood (e.g., a 5×5 pixel window). Using the grayscale values of each pixel as weights, it calculates the weighted centroid of the coordinates of all pixels in the neighborhood. Through this grayscale weighted calculation, the positioning accuracy of the corner points is improved from the integer pixel level to the sub-pixel level, with the positioning deviation controlled within 0.1 pixels. The corner points extracted by the grayscale weighted centroid algorithm serve as the absolute base points for subsequent homography matrix calculations and orthophoto plane coordinate system reconstruction. Their sub-pixel-level positioning accuracy directly determines the overall upper limit of accuracy for spatial reconstruction and size mapping.
[0076] Secondly, regarding the underlying refinement of dynamic pixel-physical size mapping compensation: Specifically, the processing module automatically scans the grid pixel spacing of the multi-level dimensional reference calibration area to obtain the pixel distribution data of the grid on the surface of the visual mapping reference pad in the original image. Under ideal distortion-free conditions, a 10mm × 10mm subdivided grid on the reference pad should appear as a uniformly distributed square pixel grid in the orthographic view. However, due to the radial distortion (fisheye effect) of lenses from different brands of acquisition equipment, the grid will bend and stretch in the image edge area, manifesting as changes in grid curvature.
[0077] The processing module utilizes grid curvature to perform nonlinear real-time correction of radial distortion in the acquisition device. Grid curvature reflects the degree of bending of the grid in the distorted image: in the central region of the image, the grid curvature is close to zero (the grid is approximately square); in the edge region of the image, the grid curvature increases significantly (the grid bends into a barrel or pincushion shape). By analyzing the distribution characteristics of the grid curvature, the processing module calculates the radial distortion parameters of the acquisition device (including the distortion center and distortion coefficients) and performs reverse compensation correction on the original image, making the pixel spacing of the corrected grid tend to be uniform across the entire image.
[0078] The processing module constructs a mapping model based on the corrected grid pixel spacing, using pixel coordinates as input and physical dimensions as output. This mapping function takes the form f(x_pixel, y_pixel) = (x_physical, y_physical), enabling accurate conversion between pixel coordinates and physical dimensions at any position on the visual mapping reference pad. This eliminates the dimensional deviations caused by uncorrected radial distortion in edge regions, a problem inherent in traditional linear mapping.
[0079] Finally, regarding the low-level refinement of automatic feature extraction based on the region of interest (ROI), specifically, the processing module automatically locks the ROI of the part under test based on the coordinates of the central positioning region. The central positioning region (such as a concentric circle positioning region) has a clear coordinate position and geometric features in the image. The processing module uses the coordinates of this region as the reference origin and extends outwards to the effective acquisition range boundary of the reference pad, forming the ROI that defines the body of the part under test. The effectiveness of this ROI lies in the fact that it is confined within the effective acquisition range of the reference pad, automatically shielding it from cluttered environmental noise outside the ROI. This cluttered environmental noise includes background objects outside the pad, oil stains on the ground, and irregular shadows, etc., which are easily misjudged as the true edges of the part by algorithms in traditional methods, leading to serious distortion of the measurement results.
[0080] The processing module automatically identifies and labels the key geometric parameters of the part under test within the region of interest (ROI). These key geometric parameters include at least one of the following: bolt hole center distance, gear module, and maximum outer diameter. The module uses edge detection and contour extraction algorithms to identify the geometric features of the part within the ROI and, combined with a mapping model, converts these parameters from the pixel domain to the physical dimension domain, outputting millimeter-level physical measurement data of the part. For example, for extracting the bolt hole center distance, the module first detects all circular contours within the ROI as candidate bolt holes, calculates the pixel coordinates of the center of each circular contour, then selects the target bolt hole pair and calculates the pixel distance between their centers. Finally, it converts the pixel distance to the physical distance using a mapping function, outputting bolt hole center distance data with millimeter-level accuracy.
[0081] Example 10: Based on Example 7, this example further refines the acquisition end of the visual mapping system by introducing an image acquisition module and a sensor fusion mechanism to ensure the standardization of input data from the source.
[0082] Specifically, the visual mapping system also includes a data acquisition module, which calls upon the attitude sensor data of the acquisition device. The acquisition module is typically integrated as software within the acquisition device (such as a smartphone or industrial tablet), interacting directly with the device's operating system and sensor hardware. Attitude sensors (such as gyroscopes and accelerometers) can sense the attitude angle of the acquisition device in three-dimensional space in real time, including the tilt angle of the device's plane relative to the horizontal plane.
[0083] When the angle between the plane of the acquisition device and the plane of the visual mapping reference pad exceeds a preset threshold, the acquisition module automatically disables the image acquisition function and outputs a visual guidance prompt. The preset threshold is usually set to a small angle (e.g., 2 degrees) to ensure that the acquisition device is in a near-absolute vertical overhead position. When the attitude sensor detects that the tilt angle of the device exceeds this threshold, the acquisition module immediately disables the shutter button or shooting trigger interface to prevent the operator from acquiring images in an unqualified posture. At the same time, it outputs a visual guidance prompt on the screen of the acquisition device (e.g., displaying a tilt direction indicator arrow, attitude correction animation, or text reminder) to guide the operator to adjust the device attitude until the tilt angle falls back to within the threshold range. At this time, the acquisition module automatically resumes the image acquisition function, allowing the operator to trigger the shutter.
[0084] The technical effect of this sensor fusion mechanism is that it significantly reduces the degree of perspective distortion at the source, making the original image input to the processing module closer to the orthophoto plane, thereby improving the stability of homography matrix calculation and the accuracy of pixel mapping. In traditional solutions, operators can freely take pictures at any tilt angle, resulting in uncontrollable perspective distortion of the input image. Even after homography matrix correction, the mapping accuracy of edge areas in images taken at extreme tilt angles will still decrease significantly. This embodiment, however, controls the distortion of the input image to a very small range by locking the pose of the acquisition module, providing standardized, high-quality input data for the backend algorithm, fundamentally improving the overall mapping accuracy and stability of the system.
[0085] It should be understood that the specific type of attitude sensor is not limited to a gyroscope; it can also be an accelerometer, magnetometer, or a combination thereof (such as an IMU inertial measurement unit), as long as it can sense the tilt angle of the acquisition device relative to the reference plane. The specific value of the preset threshold is not limited to 2 degrees; it can be adjusted according to the actual mapping accuracy requirements. For example, it can be set to 1 degree in scenarios requiring higher accuracy, and to 5 degrees in scenarios where slight distortion is acceptable. Furthermore, the specific form of visual guidance is not limited to screen display; it can also be voice prompts, vibration feedback, or flashing LED indicators, as long as it effectively guides the operator to adjust their posture.
[0086] Example 11: To more clearly illustrate the actual operating effect and feasibility of the visual mapping reference pad and visual mapping system of the present invention in extreme industrial environments, the following example uses the visual mapping scenario of a super-large component (gearbox assembly) at a heavy truck repair site to demonstrate the hardware splicing and software correction features in the aforementioned examples.
[0087] In this scenario, the repair site environment is extremely harsh, with severe oil contamination on the ground and complex industrial lighting from the workshop ceiling, including direct light from multiple angles. The object to be mapped is a heavy-duty truck gearbox assembly, whose dimensions exceed the effective acquisition area of a single reference pad, and its surface is covered with metallic high-reflection areas and dark oil-stained shadow areas, posing a significant challenge to traditional visual mapping solutions. For this scenario, such as... Figure 7 As shown, the surveying method operation procedure is as follows: Step S710: Modular splicing and physical benchmark construction. The operator modularly splices three visual mapping benchmark pads using a detachable splicing connection structure at the edges to form an ultra-large shooting background. Specifically, adjacent pads are initially physically connected via hook-and-loop fasteners at the edges. Then, at the overlapping corners, positioning pins are passed through embedded copper rings to achieve seamless physical alignment. Because the center of the positioning anchor coincides with the center of the visual coding mark in the vertical projection direction, the absolute physical alignment of these three pads is directly and losslessly converted into a unified benchmark for the coordinate system in the visual algorithm, completely eliminating the large-scale mapping cumulative errors caused by physical gaps and visual coordinate misalignment in traditional multi-pad splicing.
[0088] Step S720: Heavy-duty component placement and anti-slip deformation locking. The operator hoists and places the heavy-duty truck gearbox assembly into the central concentric circle positioning area of the splicing pad. The geometric constraints of the concentric circles forcefully guide the gearbox's centered placement, maximizing the effective visual acquisition area of the extra-large splicing pad and reducing the impact of edge perspective distortion. At the moment of placement, the weight of the gearbox, weighing tens of kilograms, ensures that the cross-shaped anti-slip texture on the surface of the high-damping anti-slip layer of the pad firmly grips the ground, ensuring no relative displacement between the pad and the ground. Simultaneously, the middle fiber reinforcement layer effectively resists localized tensile deformation caused by the pressure of the heavy component, ensuring that the absolute physical accuracy of the surface subdivision grid and edge scale is not compromised.
[0089] Step S730: Image Acquisition and Optical Interference Elimination. The operator uses a smartphone to take a top-down shot. During the acquisition phase, the phone's built-in gyroscope performs attitude locking, allowing the shutter to be triggered only when the device's tilt angle is less than 2 degrees, reducing perspective distortion at the source. More importantly, facing the strong sunlight in the workshop and the intense reflection from the gearbox's metal surface, the nanoscale diffuse-reflective matte coating on the top layer of the pad plays a crucial optical barrier role. Its micro-nano structure, with a gloss level controlled below 5 GU, disperses and absorbs most of the direct light into diffuse-reflective microbeams, fundamentally eliminating diffraction and reflection ghosting caused by the edges of metal parts in the image, providing a clean, low-noise background for edge extraction in the backend algorithm.
[0090] Step S740: Algorithm correction and millimeter-level parameter output. After the image is transmitted to the processing module, the system enters a fully automatic correction and mapping process. The processing module first identifies the asymmetric ArUco encoded arrays at the four corners of the stitched ultra-large background. Based on its asymmetric characteristics, it automatically determines the absolute orientation of the image in a very short time, confirming the true vertical position of the gearbox in the image without manual intervention. Subsequently, the processing module executes a four-step algorithm: First, it locates visually encoded markers against an oil-polluted background using dynamic threshold segmentation and extracts corner points to sub-pixel accuracy using a gray-scale weighted centroid algorithm; second, it calculates the homography matrix based on the ratio of sub-pixel corner points to the known physical dimensions of the pad, mapping the perspective-distorted image region to the orthographic projection plane to reconstruct a millimeter-level physical coordinate system; next, it scans the grid pixel spacing of the multi-level size reference calibration area, uses grid curvature to perform nonlinear correction on the radial distortion of the lens, and constructs a mapping function between pixel coordinates and physical dimensions; finally, it locks the region of interest where the gearbox is located based on the coordinates of the centrally located area, and automatically extracts and outputs millimeter-level physical parameters such as the gearbox bolt hole spacing and outer contour dimensions under the condition of shielding environmental noise outside the pad.
[0091] As can be seen from the complete operation process of the above extreme scenarios, the present invention achieves the same origin and unity of physical splicing benchmark and visual positioning benchmark through the coaxial layout design of physical positioning anchor and visual coding mark. Combined with multi-layer composite anti-deformation structure and nano-level anti-reflective coating, it can achieve 1:1 accurate mapping without manual intervention to correct deviation in industrial sites with serious oil pollution, complex lighting and huge heavy parts. This fully verifies the core value of the technical solution of the present invention in eliminating cumulative errors and resisting environmental interference.
[0092] It should be understood that the above-mentioned mapping scenario for heavy-duty truck gearbox assemblies is merely a typical example among the many industrial applications of this invention, and is not intended to limit the scope of protection of this invention. In other embodiments, the visual mapping reference pad and system of this invention can also be applied to mapping scenarios for other ultra-large or highly reflective metal parts such as aero-engine components, large wind turbine gearboxes, and hydraulic valve blocks for mining machinery. As long as they adopt a core structure in which physical anchor points and visual marks are coaxially aligned, and an optical design with a low-reflection matte surface layer, they should all fall within the scope of protection of this invention.
[0093] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer, implements the method described in the above-described method embodiments. Specific beneficial effects can be found in the above-described method embodiments.
[0094] The present invention also provides a computer program product that, when executed by a computer, implements the method described in the above-described method embodiments. Specific beneficial effects can be found in the above-described method embodiments.
[0095] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0096] In the various embodiments of this invention, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0097] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.
[0098] The above description is merely a specific implementation of the embodiments of the present invention, but the protection scope of the embodiments of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of the present invention should be covered within the protection scope of the embodiments of the present invention. Therefore, the protection scope of the embodiments of the present invention should be determined by the protection scope of the claims.
Claims
1. A visual mapping system, characterized in that, include: Visual mapping reference pad and processing module, wherein: The visual mapping reference pad includes a pad body; a surface layer disposed on the surface of the pad body; the surface layer is a matte coating with a microscopic uneven structure; a positioning anchor disposed on the pad body, the positioning anchor being used for positioning and alignment during splicing; a visual coding mark disposed on the surface layer; wherein the center of the visual coding mark coincides with the center of the positioning anchor in the vertical projection direction; and a multi-level size reference calibration area disposed on the surface layer. The processing module is configured to calculate the homography matrix by recognizing the visual coding marks of the visual mapping reference pad, reconstruct the orthophoto plane coordinate system, establish a mapping model between pixel coordinates and physical dimensions, and output the key geometric parameters of the part to be measured.
2. The visual mapping system according to claim 1, characterized in that, The processing module is specifically configured as follows: The visually encoded markers are identified through image segmentation, and the corner points of the visually encoded markers are extracted to sub-pixel accuracy. Based on the known physical size ratio of the corner points of the visual coding mark to the visual mapping reference pad, the homography matrix is calculated, and the image area with perspective distortion is mapped to the orthophoto plane to reconstruct the orthophoto plane coordinate system. Based on the pixel spacing of the multi-level size reference calibration area on the visual mapping reference pad, a mapping model between pixel coordinates and physical dimensions is established. Based on the coordinates of the central positioning area of the visual mapping reference pad, the region of interest of the part to be measured is locked, and the key geometric parameters of the part to be measured are extracted within the region of interest.
3. The visual mapping system according to claim 1, characterized in that, After locating the visual encoding mark, the processing module is specifically configured as follows: A gray-scale weighted centroid algorithm is used to extract the corner points of the visually encoded marker at the sub-pixel level. The gray-scale weighted centroid algorithm calculates the centroid coordinates by using the gray values of each pixel in the neighborhood of the corner point as weights, thereby controlling the positioning deviation of the corner points of the visually encoded marker within a set pixel value. The corner points of the visually encoded marker extracted by the gray-scale weighted centroid algorithm are used as the absolute base points for subsequent homography matrix calculation and orthophoto plane coordinate system reconstruction.
4. The visual mapping system according to any one of claims 1 to 3, characterized in that, It also includes an image acquisition module, which calls its own attitude sensor data. When the angle between the plane of the image acquisition module and the plane of the visual mapping reference pad exceeds a preset threshold, the image acquisition module is configured to automatically disable the image acquisition function and output a visual guidance prompt.
5. The visual mapping system according to claim 4, characterized in that, The processing module is specifically configured as follows: Automatically scan the grid pixel spacing of the multi-level size reference calibration area to obtain the pixel distribution data of the surface grid of the visual mapping reference pad in the original image; The radial distortion of the image acquisition device is nonlinearly corrected in real time using grid curvature, where grid curvature reflects the degree of bending of the grid in the distorted image. The radial distortion parameters of the image acquisition device are calculated based on the distribution characteristics of the grid curvature and inverse compensation is performed. Based on the corrected grid pixel spacing, the mapping model is constructed as a mapping function with pixel coordinates as input and physical size as output. The mapping function is used to convert between pixel coordinates and physical size at any position on the visual mapping reference pad.
6. The visual mapping system according to any one of claims 1 to 3, characterized in that, The positioning anchor is an embedded ring structure, and the visual coding mark is an asymmetric identifier sequence.
7. The visual mapping system according to claim 6, characterized in that, The embedded ring structure is an embedded metal ring, and the asymmetric identifier sequence is an ArUco encoding array.
8. The visual mapping system according to any one of claims 1 to 3, characterized in that, The multi-level size reference calibration area includes at least one of the following: subdivision grid, edge scale, or center concentric circle positioning area.
9. The visual mapping system according to any one of claims 1 to 3, characterized in that, The pad body includes a damping and anti-slip layer and a fiber reinforcement layer covering the surface of the damping and anti-slip layer.
10. The visual mapping system according to any one of claims 1 to 3, characterized in that, The edge of the pad body is provided with a detachable splicing connection structure.