Visual measurement positioning system and method based on environment structured information
By utilizing the inherent structure of the measurement field and intelligent scheduling algorithms in large-scale industrial environments, a visual measurement and positioning system based on structured environmental information is constructed. This solves the problems of high cost, complex calibration, weak anti-occlusion, and slow response of traditional visual measurement systems, and achieves efficient, accurate, and fast measurement and positioning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies for visual measurement and positioning systems in large-scale, dynamically occluded, and unstructured industrial environments suffer from problems such as high cost, complex calibration processes, weak anti-occlusion capabilities, and insufficient response speed, making it difficult to meet the high-precision and real-time measurement requirements of high-end equipment manufacturing.
A visual measurement and positioning system based on environmental structured information is adopted. The inherent structure of the measurement field is used as visual feature points. Combined with intelligent scheduling algorithms and multi-camera collaborative perception, a collaborative perception architecture of "laser coarse positioning - visual fine positioning" is realized. Through dynamic hybrid calibration and intelligent resource scheduling, occlusion is predicted in real time and cameras are scheduled to perform efficient and continuous measurement.
It significantly reduces system costs, achieves high-precision measurement, has strong anti-obstruction capabilities, rapid response, supports modular deployment, and meets the real-time measurement needs of industrial sites.
Smart Images

Figure CN121739983A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of large-size industrial measurement and positioning technology, specifically relating to a visual measurement and positioning system and method based on structured environmental information. Background Technology
[0002] In the manufacturing, assembly, and maintenance of large industrial equipment such as aircraft, high-precision measurement and positioning are crucial for ensuring product quality and improving process efficiency. Traditionally, these tasks have relied primarily on high-precision optical measurement equipment such as laser trackers and LiDAR. While these devices can achieve high single-point measurement accuracy (typically down to the micrometer level), they exhibit several significant limitations when dealing with large-scale, multi-point, and dynamic measurement tasks: High cost: Laser trackers, lidar and other equipment are expensive, and usually require professional software, calibration tools and highly skilled operators. The overall deployment and maintenance costs are extremely high, making it difficult to promote on a large scale in small and medium-sized manufacturing enterprises or scenarios that require multi-regional coverage.
[0003] Inefficient and complex in operation: For large objects such as aircraft, traditional equipment often requires multiple relocations, that is, re-establishing stations at different locations and measuring common points to achieve full-area coverage. This process is not only time-consuming (a single full-aircraft measurement can take several hours), but also introduces cumulative errors due to alignment errors and coordinate fitting errors caused by each station setting, affecting the consistency of the final measurement accuracy.
[0004] Poor environmental adaptability: These types of equipment have high requirements for the measurement environment and are easily affected by common factors in industrial sites such as vibration, temperature and humidity changes, and dust. Furthermore, their continuous measurement capability is limited in scenarios with frequent dynamic obstructions (such as AGVs, robotic arms, and personnel movement).
[0005] On the other hand, vision-based measurement methods, due to their advantages such as non-contact operation, rich information, and relatively low hardware costs, have gradually become a research hotspot in the field of industrial measurement. In close-range, fixed-viewpoint, structured lighting environments such as laboratories or small-scale scenarios, vision measurement technology has matured, achieving sub-millimeter or even higher positioning accuracy. However, when facing large-scale measurement fields (such as 30m × 15m and above), frequent dynamic occlusion, and complex industrial environments with non-fixed viewpoints, existing vision measurement solutions still face significant challenges. High hardware costs: To cover the entire large measurement area, using fixed-view cameras would require deploying a large number of sensors (typically hundreds), causing a surge in system construction costs. Furthermore, the large number of cameras also introduces complexities in wiring, power supply, network communication, and subsequent maintenance.
[0006] The calibration process is complex and lacks stability: In dynamic industrial environments, cameras are prone to parameter drift due to factors such as mechanical vibration, temperature changes, and equipment displacement. In particular, the internal and external parameters of gimbal zoom cameras change after zooming and rotation. Traditional one-time calibration cannot be used for a long time and frequent recalibration is required, which seriously affects system availability and maintenance costs.
[0007] Weak resistance to obstruction: Once the line of sight of a fixed-view camera is obstructed by a mobile device (such as an AGV or robotic arm) or a temporary obstacle, it cannot continue to complete the measurement, resulting in data interruption or loss. In processes such as aircraft assembly and painting, obstruction is extremely common, and the system must have dynamic adaptation and compensation capabilities.
[0008] Insufficient system response speed: Traditional vision measurement methods often rely on global high-resolution image processing, complex feature matching and optimization algorithms, which have a large computational load and are difficult to meet the industrial field's demand for second-level real-time response, and cannot support high-speed continuous operation and real-time closed-loop control.
[0009] Therefore, how to build a cost-effective, calibrated, anti-interference, and fast-response visual measurement and positioning system in large-scale, dynamically occluded, and unstructured industrial environments has become a technical bottleneck that urgently needs to be overcome in the field of intelligent manufacturing of high-end equipment. Summary of the Invention
[0010] To address the problems existing in the prior art, this invention provides a visual measurement and positioning system and method based on environmental structured information, aiming to transform the environmental structured information of the measurement field from "background" into a usable "perceptual asset" to achieve high-precision, high-efficiency, and high-robustness visual measurement and positioning in large-scale dynamic industrial scenarios.
[0011] To achieve the above objectives, the present invention provides the following solution: A visual measurement and positioning system based on structured environmental information, comprising: an intelligent sensing hardware layer, an environmental structured feature library, and a core processing and control unit; The intelligent sensing hardware layer includes cameras equipped with high-resolution sensors and motorized gimbals deployed in the measurement field, and a multi-source sensor fusion module; the camera is used to achieve visual fine positioning; the multi-source sensor fusion module is used to form a collaborative sensing architecture of "laser coarse positioning - visual fine positioning"; The environmental structured feature library includes natural feature points and auxiliary calibration objects; The core processing and control unit includes a data processing module and an intelligent scheduling and calibration module. The data processing module is used to perform image processing, deep learning recognition, feature extraction and matching, and multi-view geometric calculation based on the positioning results and natural feature points of the intelligent sensing hardware layer and auxiliary calibration objects, so as to obtain the three-dimensional coordinates of the target feature points of the target under test in the global coordinate system. The intelligent scheduling and calibration module is used to perform dynamic calibration and intelligent scheduling of the camera to realize continuous measurement of the target under test.
[0012] Preferably, the intelligent scheduling and calibration module includes a dynamic hybrid calibration engine and a resource intelligent scheduler; A dynamic hybrid calibration engine is used to manage camera parameters, including offline calibration based on calibration boards and online PnP real-time calibration based on environmental feature points, and outputs camera parameters through offline and online calibration. Resource intelligent scheduler: It receives the target position and occlusion prediction information of the target under test in real time, dynamically calculates and directs the most suitable subset of cameras to adjust the gimbal angle and focal length, and performs the measurement task with the best perspective, so as to realize continuous measurement of the target under test under moving occlusion.
[0013] This invention also provides a visual measurement and localization method based on structured environmental information, the method being implemented using the aforementioned system, the method comprising: Collect the three-dimensional coordinates of pre-set structured feature points within the measurement field to construct a feature map in the global coordinate system of the measurement field; Based on feature points with known coordinates observed by several cameras, determine the initial intrinsic and extrinsic parameters of each camera, and unify all cameras to the same world coordinate system; Based on the preset offline calibration benchmark and online PnP calibration, the camera parameters are updated online through self-calibration. Based on the feature map and the camera updated online by self-calibration, combined with deep learning and image processing technology, the target to be measured in the measurement field is identified, and the three-dimensional coordinates of the target feature points in the global coordinate system are obtained. Based on the three-dimensional coordinates of the target feature points in the global coordinate system, a geometric model is used to obtain occlusion prediction information, and the camera is scheduled to achieve continuous measurement of the target.
[0014] Preferably, the method for online self-calibration and updating of camera parameters based on a preset offline calibration benchmark and online PnP calibration includes: The PnP algorithm is called to perform online real-time calibration and obtain online calibration results; The online calibration results are combined with the offline calibration results in the offline calibration benchmark to obtain the final calibration result, thereby realizing online self-calibration and updating of camera parameters.
[0015] Preferably, the method for obtaining online calibration results by calling the PnP algorithm for online real-time calibration includes: Points in the known world coordinate system and its corresponding image plane projection point The projection relationship is as follows: pi = KR(Pi - t); Where i is the number of points in the known world coordinate system. This is the intrinsic parameter matrix of the camera. For extrinsic rotation matrix, This is the extrinsic translation matrix; Construct a quartic polynomial equation: ; in, a , b、 、 All are defined vectors. k 1. k 2 represents the scaling factor; Solve k 1. k 2. Select the solution that minimizes the projection error of all known world coordinate system points from the solution results as the final extrinsic rotation matrix and extrinsic translation matrix.
[0016] Preferably, the method of integrating the online calibration results with the offline calibration results in the offline calibration benchmark to obtain the final calibration result and realize the online self-calibration update of camera parameters includes: ; ; in, For the final calibration result, For online calibration results, For offline calibration results, and These are the time-varying calibration result weighting coefficients, satisfying the condition that the sum of the two is 1, a constant. and constant These are used to control the initial weights and the rate of change of the weights, respectively. t For time.
[0017] Preferably, a method for identifying the target in the measurement field and obtaining the three-dimensional coordinates of the target feature points in the global coordinate system based on feature maps and an online self-calibrated camera, combined with deep learning and image processing techniques, includes: When the target enters the measurement field, a lidar is used to scan the target to achieve coarse positioning and obtain the region of interest; Based on the region of interest, the cameras are scheduled after online self-calibration updates to obtain the optimal camera combination; The optimal camera combination is used to acquire images of the target object. Based on the image of the target to be tested, deep learning models and image processing techniques are used to obtain the target feature points of the target to be tested; The target feature points are solved, and the solution results are fused to obtain the three-dimensional coordinates of the target feature points in the global coordinate system.
[0018] Preferably, the method for scheduling cameras after online self-calibration updates based on the region of interest to obtain the optimal camera combination includes: Establish a multi-objective optimization decision model and construct the cost function: ; in, These are hyperparameter weighting coefficients, which can be adjusted independently according to the working environment. The angle of rotation of the gimbal. This is the maximum rotatable angle of the gimbal. It is the distance between the camera's imaging point and the point to be measured. This is the camera's maximum working distance. It is the standard deviation of the measurement error. It is the system calibration constant. It's the camera's focal length. It refers to the camera pixel size.
[0019] Preferably, the method for obtaining occlusion prediction information using a geometric model based on the three-dimensional coordinates of the target feature points in the global coordinate system includes: The MES system based on the measurement field acquires real-time coordinates and path planning data of the AGV and robotic arm scheduling system; Based on the real-time coordinates and path planning data of the AGV and robotic arm scheduling system, a bounding box model of the target to be tested is constructed. Based on the three-dimensional coordinates of the target feature points in the global coordinate system, and combined with the bounding box model of the target, the dot product of the normal vector of the bounding box surface of the target and the camera viewing direction vector is obtained. Based on the dot product, determine whether occlusion exists and obtain occlusion prediction information.
[0020] Preferably, methods for determining the existence of occlusion and obtaining occlusion prediction information based on dot product include: If dot product It is 0, and satisfies If so, it is considered an obstruction; If dot product Not zero, and satisfies the parameter If it is occluded, it is considered occluded; otherwise, it is considered unoccluded. The parameter is... t The calculation formula is as follows: ; in, A , B , C , D For plane equation parameters, Define variables.
[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: The visual measurement and positioning system and method proposed in this invention exhibit the following significant advantages in terms of technical economy, performance indicators, and engineering feasibility compared to traditional large-size industrial measurement solutions: 1. Significantly reduce overall costs This invention leverages the inherent structure of the measurement field (factory building) (such as columns, window frames, and steel frame connection points) as natural and stable visual feature points, significantly reducing the number of high-precision calibration plates requiring specialized processing, installation, and maintenance, directly lowering material and construction costs. Simultaneously, based on a gimbal camera (zoom camera) and intelligent scheduling algorithms, it can dynamically cover a large measurement area with fewer sensors (after optimization, the number of cameras can be reduced by approximately 70% compared to a fixed-viewpoint solution), saving not only hardware procurement costs but also subsequent power, network, and maintenance costs, achieving effective control over the entire lifecycle cost.
[0022] 2. Achieve and maintain high-precision measurements This invention employs a hybrid calibration strategy of "offline high-precision calibration benchmark + online dynamic real-time correction". During the initialization phase, feature point coordinates are acquired using precision instruments, laying a solid foundation for the system. During operation, continuous self-calibration is performed using environmental feature points, effectively compensating for parameter changes caused by factors such as equipment mechanical wear and temperature drift. By combining observation data from multiple cameras at the same target point from different perspectives and performing fusion calculations, single-view errors are eliminated through geometric redundancy. Thus, even in complex industrial environments, the end-point positioning accuracy can be stably controlled within 2 millimeters, meeting the stringent requirements of high-end equipment manufacturing and assembly.
[0023] 3. Highly adaptable to dynamic occlusion environments. The core intelligent scheduling algorithm of this invention can predict and respond to occlusion events in real time. By accessing motion data from devices such as AGVs and robotic arms, it proactively assesses the risk of occlusion in the measurement line of sight. When occlusion is predicted, the scheduler can dynamically wake up or adjust another suitable set of gimbal cameras to take over the measurement task within seconds, achieving seamless switching and continuous operation of the measurement process. This capability ensures that the measurement data stream is uninterrupted in operation processes with frequent occlusion, such as spraying and grinding, and the system exhibits extremely high robustness.
[0024] 4. Achieve rapid system response. From triggering the measurement command to outputting the positioning result, this invention reduces the overall response time to less than 2 seconds through parallel processing and pipeline optimization. This is thanks to the efficient "coarse-fine" two-stage positioning process: the lidar or wide-angle camera first performs rapid coarse positioning, guiding the high-resolution gimbal camera to aim precisely; at the same time, the intelligent scheduling algorithm quickly selects the optimal camera resources based on spatial indexing, greatly reducing unnecessary image acquisition and calculation time, thereby meeting the speed requirements of modern intelligent manufacturing for real-time feedback and closed-loop control.
[0025] 5. Supports modular, rapid deployment and expansion. The hardware layout and software architecture of the entire system of this invention adopt a highly modular design. Sensor nodes, computing units, and network devices can all be installed and configured independently according to functional modules. When building a new measurement field or expanding an existing field, the deployment mode can be quickly reproduced according to the layout planning method described in this invention. Newly connected cameras can be seamlessly integrated into the existing system through a unified calibration and scheduling protocol, which greatly improves the efficiency and flexibility of engineering implementation and supports rapid ramp-up and flexible expansion of production capacity. Attached Figure Description
[0026] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a schematic diagram of the visual measurement and positioning method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the measurement field voxel network according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the dynamic occlusion probability model in an embodiment of the present invention. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0030] Example 1 This invention provides a visual measurement and positioning system based on structured environmental information, comprising: an intelligent sensing hardware layer, an environmental structured feature library, and a core processing and control unit; The intelligent sensing hardware layer includes cameras equipped with high-resolution sensors and motorized gimbals deployed in the measurement field, and a multi-source sensor fusion module; the camera is used to achieve visual fine positioning; the multi-source sensor fusion module is used to form a collaborative sensing architecture of "laser coarse positioning - visual fine positioning"; The environmental structured feature library includes natural feature points and auxiliary calibration objects; The core processing and control unit includes a data processing module and an intelligent scheduling and calibration module. The data processing module is used to perform image processing, deep learning recognition, feature extraction and matching, and multi-view geometric calculation based on the positioning results and natural feature points of the intelligent sensing hardware layer and auxiliary calibration objects, so as to obtain the three-dimensional coordinates of the target feature points of the target under test in the global coordinate system. The intelligent scheduling and calibration module is used to perform dynamic calibration and intelligent scheduling of the camera to realize continuous measurement of the target under test.
[0031] Furthermore, the intelligent scheduling and calibration module includes a dynamic hybrid calibration engine and a resource intelligent scheduler; A dynamic hybrid calibration engine is used to manage camera parameters, including offline calibration based on calibration boards and online PnP real-time calibration based on environmental feature points, and outputs camera parameters through offline and online calibration. Resource intelligent scheduler: It receives the target position and occlusion prediction information of the target under test in real time, dynamically calculates and directs the most suitable subset of cameras to adjust the gimbal angle and focal length, and performs the measurement task with the best perspective, so as to realize continuous measurement of the target under test under moving occlusion.
[0032] Example 2 Based on the same inventive concept, the present invention also provides a visual measurement and localization method based on structured environmental information, implemented using the system described in the foregoing embodiments, the method comprising: Collect the three-dimensional coordinates of pre-set structured feature points within the measurement field to construct a feature map in the global coordinate system of the measurement field; Based on feature points with known coordinates observed by several cameras, determine the initial intrinsic and extrinsic parameters of each camera, and unify all cameras to the same world coordinate system; Based on the preset offline calibration benchmark and online PnP calibration, the camera parameters are updated online through self-calibration. Based on the feature map and the camera updated online by self-calibration, combined with deep learning and image processing technology, the target to be measured in the measurement field is identified, and the three-dimensional coordinates of the target feature points in the global coordinate system are obtained. Based on the three-dimensional coordinates of the target feature points in the global coordinate system, a geometric model is used to obtain occlusion prediction information, and the camera is scheduled to achieve continuous measurement of the target.
[0033] like Figure 1 As shown, the implementation process of this embodiment is as follows: Initialization phase: From hardware deployment to system readiness, this phase involves preparatory work that can be completed in one go or periodically.
[0034] Run the measurement loop: the core working loop from target triggering to result output.
[0035] Dynamic adjustment loop: This reflects the system's adaptive ability to adjust the scheduling strategy in real time according to changes in the environment during task execution, which is the key to the system's robustness.
[0036] This process design ensures clear and operable guidance throughout the entire lifecycle of the system, from installation and calibration to operation and maintenance. Furthermore, the core measurement cycle is highly efficient and adaptive, enabling reliable applications in complex industrial scenarios.
[0037] The specific implementation process is as follows: Step 1: System Planning and Hardware Deployment Site surveying and planning: A 3D laser scan of the target factory building (surveying area) is performed to create a digital model containing all major structures (columns, beams, walls, equipment foundations). Based on this model, the optimal installation location, number, and initial viewing angle of the pan-tilt zoom camera are determined.
[0038] Feature point selection and measurement: In the digital model, a series of environmentally structured feature points that are not easily moved and are easy to image are selected (such as the intersection of the edges of specific columns, the fixing bolts of window frames, specific welding points on steel frames, etc.). At the same time, the installation positions of auxiliary calibration plates are planned in areas with sparse natural features. Using high-precision measuring equipment such as laser trackers, the three-dimensional global coordinates of all selected feature points are accurately measured on-site, and a "global feature point coordinate database" is established.
[0039] The term "sparsely ...
[0040] A region is typically classified as "sparse" based on one or more of the following conditions: Lack of geometric features: The area consists mostly of large, smooth, and textureless flat surfaces (such as large, flat walls, ceilings, and floors), lacking features with clear geometric meaning, such as corners and edge intersections.
[0041] Low feature distribution density: The few natural feature points in the region (such as an isolated bolt) are too far apart to form an effective "feature point set" that can be used for high-precision calculation from the same camera view (usually the P3P algorithm requires at least 3 non-collinear points).
[0042] Poor feature stability: There may be temporary piles of materials or areas where mobile devices frequently pass through, or the surface material may be reflective or easily contaminated, causing the appearance of feature points to be changeable or appear and disappear intermittently, making them unreliable.
[0043] Limited viewing angle: Due to the permanent obstruction of the factory structure (such as beams and large equipment), most cameras cannot observe the natural features in the area from multiple angles, which cannot meet the requirements of multi-view intersection measurement.
[0044] Hardware installation and networking: Install the pan-tilt zoom camera, necessary LiDAR (for coarse positioning), computing server, network switch, and other equipment according to the plan. Complete the power supply and communication network (usually industrial Ethernet) connection and initial debugging of all equipment.
[0045] Step 2: System Initialization and Offline Calibration Camera Fixing and Parameter Presetting: Ensure all cameras are physically stable. Record the approximate coordinates and initial orientation of each camera's mounting position.
[0046] Offline calibration and system modeling: Each camera is controlled to sequentially observe multiple feature points with known global coordinates (including calibration boards and natural structure points), acquiring multiple images. Classic algorithms such as the Zhang Zhengyou calibration method are employed to calculate and store high-precision initial intrinsic parameters (focal length, distortion, etc.) and extrinsic parameters (position and attitude) for each camera offline. Simultaneously, by using common points in the overlapping fields of view, the relative positional relationships between all cameras are calculated and solidified, constructing a system-level measurement network model.
[0047] Specifically, control requires establishing a group of related cameras (e.g., camera A and camera B) to ensure that they can both capture three points in a unified coordinate system.
[0048] For each image captured by a camera, the pixel coordinates (u, v) of common points are extracted. A correspondence is established between image points, common points, and their corresponding ID codes (QR codes). That is, it determines which known 3D coordinate point P_w = (X, Y, Z) in the database corresponds to a specific pixel in the image. Since all world coordinate system points use a unified world coordinate system, the derived extrinsic parameters for each camera are obtained based on the same world coordinate system. Therefore, any pair of cameras can be combined to form a measurement system.
[0049] Step 3: Online self-calibration and system readiness Routine startup and self-test: When the system starts up, each camera automatically rotates to the preset "calibration pose" (i.e., the position that can capture at least three three-dimensional coordinate points in the known world coordinate system), so that its field of view can cover several known environmental feature points.
[0050] Online calibration and parameter fusion: The system automatically captures the current image, identifies environmental feature points within the field of view (using Hough circle detection to identify target spheres in the environment and YOLOv8 to detect key points in the image), and calls algorithms such as PnP for online real-time calibration to obtain extrinsic parameters under the current pose, including: The PnP algorithm is called to perform online real-time calibration and obtain online calibration results; The online calibration results are combined with the offline calibration results in the offline calibration benchmark to obtain the final calibration result, thereby realizing online self-calibration and updating of camera parameters.
[0051] Specifically, it includes: Assume there are 3 points in a known world coordinate system. and its corresponding image plane projection point The intrinsic parameter matrix of the camera is known to be... The extrinsic rotation matrix is The extrinsic translation matrix is The projection relationship can be expressed as: pi = KR(Pi - t); Where i represents the number of points in the known world coordinate system.
[0052] Expanding the above equation, we get: ; in,( u i , v i () represents the image coordinates. f x , f y For the camera focal length, ( u 0, v 0) represents the coordinates of the camera's imaging origin. r ij (i,j=1,2,3) are elements in the extrinsic rotation matrix. t x , t y , t z The elements are in the extrinsic translation matrix. It is a scaling factor, used for solving... The equation can be rewritten as: ; but: ; ; Assuming the world coordinates of three points are known and the image coordinates of its three corresponding points Vectors can be defined: ; Among them, a, b, , These are all defined vectors used for convenient subsequent calculations and have no actual meaning.
[0053] According to the similarity of triangles, we have: ; in, k 1. k 2 represents the scaling factor.
[0054] Using the above relationships, construct a quartic polynomial equation: ; Solve k 1. k 2. A maximum of four solutions will be obtained. From the solution results, the solution that minimizes the projection error of all known world coordinate system points will be selected as the final extrinsic rotation matrix and extrinsic translation matrix.
[0055] The result is then combined with the stored offline calibration benchmark to generate the camera parameters for the current moment. This process enables the system's "warm start" and parameter self-calibration without manual intervention. Specifically: Assuming the final calibration result is The online calibration results are (Varies over time), offline calibration results The final result of the mixed calibration is: ; ; in, For the final calibration result, For online calibration results, For offline calibration results, and These are the time-varying calibration result weighting coefficients, satisfying the condition that the sum of the two is 1, a constant. and constant These are used to control the initial weights and the rate of weight change, respectively. Considering the system maintenance cycle is one quarter, the following is selected: at the same time With this parameter configuration, offline calibration accounted for 90% in the initial stage, and online calibration accounted for 90% after 90 days. t For time.
[0056] Furthermore, based on feature maps and online self-calibrated updated cameras, combined with deep learning and image processing techniques, methods for identifying the target within the measurement field and obtaining the three-dimensional coordinates of the target's feature points in the global coordinate system include: When the target enters the measurement field, a lidar is used to scan the target to achieve coarse positioning and obtain the region of interest; Based on the region of interest, the cameras are scheduled after online self-calibration updates to obtain the optimal camera combination; The optimal camera combination is used to acquire images of the target object. Based on the image of the target to be tested, deep learning models and image processing techniques are used to obtain the target feature points of the target to be tested; The target feature points are solved, and the solution results are fused to obtain the three-dimensional coordinates of the target feature points in the global coordinate system.
[0057] Specifically, it includes steps four and five: Step 4: Target Coarse Guidance and Resource Scheduling Target entry and coarse localization: When targets such as aircraft parts or AGVs enter the measurement field, the lidar network deployed above the field area quickly scans and obtains the approximate outline and position information of the target within 1 second, achieving coarse localization (accuracy of about 50mm), and defining this area as the region of interest (ROI).
[0058] Intelligent camera scheduling: Based on the ROI provided by coarse localization and combined with a real-time updated "observability map" (including camera position, view frustum, current attitude, and occlusion prediction), the central scheduler first constructs a device capability matrix based on the installation coordinates and initial pitch angles of all gimbal cameras. This matrix includes information such as angle limitations, maximum working distance, and gimbal rotation speed. Simultaneously, the measurement field is divided into a voxel network with sides no longer than half the field of view when the camera is at its maximum working distance. Figure 2 As shown. The set of visible cameras for each voxel network is calculated, and spatial relationships are stored using an octree. The hierarchical recursive spatial partitioning mechanism of the octree and the pruning ability of the tree structure can significantly reduce access to invalid data, achieving... The time complexity of proximity device retrieval.
[0059] Specifically, a sphere centered on the lidar positioning coordinates and with the error as its radius is used as the query area. Candidate cameras within the octree are examined, and the distance from the target point (x, y, z) to the imaging point of each camera is determined. Filter by distance: ; in, d min This represents the minimum distance.
[0060] At the same time, verify the horizontal angle of the camera in the camera coordinate system. and vertical angle Is it within the mechanical limit of the gimbal? ; ; , , ; in, , , The difference is the sum of the two values.
[0061] Based on this, a set of voxel-visible cameras is constructed.
[0062] Dynamically calculate and select the 3 optimal gimbal cameras: After removing cameras with obstructed views based on occlusion conditions, establish a multi-objective optimization decision model and construct the cost function: ; in, These are hyperparameter weighting coefficients, which can be adjusted independently according to the working environment. The angle of rotation of the gimbal. This is the maximum rotatable angle of the gimbal. It is the distance between the camera's imaging point and the point to be measured. This is the camera's maximum working distance. It is the standard deviation of the measurement error. It is the system calibration constant. It's the camera's focal length. It refers to the camera pixel size.
[0063] The tournament algorithm is used to calculate the cost function of the candidate cameras, quickly select the top three cameras, and convert the calculated rotation matrix into a quaternion to avoid gimbal deadlock.
[0064] The dispatch instructions are sent through a high-speed network, commanding these cameras to rotate and zoom their gimbals so that their high-resolution field of view centers are aligned with the ROI, ready for precise measurement.
[0065] Step 5: Visual Precision Measurement and Data Fusion Target Feature Recognition: After the scheduled camera focuses, it simultaneously acquires high-resolution images. The image processing unit runs deep learning models (such as BiRefNet) and computer vision algorithms such as YOLOv8 and Hough circle detection to robustly identify key feature points on the target surface, such as the center of the cooperative target or specific landmarks, and extracts their sub-pixel image coordinates.
[0066] Multi-view triangulation and solution: For the same target feature point, using images captured by at least two cameras with different viewpoints and camera parameters optimized through dynamic hybrid calibration, 3D coordinates are obtained through triangulation. The solution results from multiple camera groups are further optimized through data fusion algorithms (such as weighted averaging), ultimately outputting the 3D coordinates of the target feature point in the global coordinate system. The installation position of the measured camera group is used as the standard, with cameras closer to the target having higher weights. Three cameras can form three pairwise measurement and positioning results, with the weights of the three measurement positions being 1 / 2, 1 / 3, and 1 / 6 respectively, according to the camera distance from smallest to largest.
[0067] Step Six: Dynamic Monitoring and Adaptive Adjustment Continuous status monitoring: Throughout the measurement task, the system continuously monitors the working status of each camera, image quality, and real-time position information of the AGV / robotic arm from the factory MES.
[0068] Occlusion prediction and strategy adjustment: Based on the trajectory of moving objects, the scheduler proactively predicts possible field-of-view occlusions within the next 1-2 seconds.
[0069] Methods for obtaining occlusion prediction information using geometric models based on the three-dimensional coordinates of target feature points in the global coordinate system include: The MES system based on the measurement field acquires real-time coordinates and path planning data of the AGV and robotic arm scheduling system; Based on the real-time coordinates and path planning data of the AGV and robotic arm scheduling system, a bounding box model of the target to be tested is constructed. Based on the three-dimensional coordinates of the target feature points in the global coordinate system, and combined with the bounding box model of the target, the dot product of the normal vector of the bounding box surface of the target and the camera viewing direction vector is obtained. Based on the dot product, determine whether occlusion exists and obtain occlusion prediction information.
[0070] Specifically, such as Figure 3 As shown, this invention uses a dynamic occlusion probability model for occlusion prediction, integrates real-time coordinates and path planning data from the AGV and robotic arm scheduling system, abstracts the size of the moving object (target) into a cuboid bounding box to obtain the bounding box model (geometric model), and establishes the planar equations for each face of the bounding box: ; in, A , B , C , D The parameters of the plane equation represent the plane's position in space.
[0071] The field-of-view ray equation is established based on the line connecting the camera's imaging position and the point to be measured (the three-dimensional coordinates of the target feature point in the global coordinate system): ; in, P ( t )for t The ray equation at time t, It is the origin of the ray, i.e., the coordinates of the camera's imaging position. It is the vector of the camera's line of sight.
[0072] The camera's line-of-sight vector is calculated based on the camera's imaging position and the point to be measured. Given the coordinates of the camera's imaging position P and the three-dimensional spatial coordinates of the point to be measured M, let the coordinates of P be (x...). p ,y p ,z p The coordinates of M are (x m ,y m ,z m If the direction vector from point M to point P is..., then... v For: (x m -x p ,y m -y p ,z m -z p ).
[0073] Calculate the dot product of the normal vector of the bounding box of the AGV or robotic arm and the camera's line-of-sight vector. : ; if If the value is 0, then the ray is parallel to the plane. If the following conditions are also met, then the line of sight and the plane overlap, which is considered occlusion: ; if If the value is not 0, the parameter is calculated using the following formula. ,if If it is, it is considered an obstruction; otherwise, it is considered unobstructed. .
[0074] in, Defined as a variable, without actual meaning, used to calculate parameters.t When it is used as a molecule.
[0075] Meanwhile, to ensure the continuity of system measurements, the system will calculate the possible occlusions in advance based on the movement status of the AGV and robotic arm within the next 2 seconds, and schedule the required cameras in advance.
[0076] Once it is predicted that the current working camera will be blocked, the system immediately activates the backup plan, which involves scheduling other available cameras that can observe the measurement point to intervene in advance, ensuring the continuity and integrity of the measurement data stream. At the same time, the system continuously performs periodic online calibration to maintain long-term operational accuracy.
[0077] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A visual measurement and positioning system based on structured environmental information, characterized in that, The system includes: an intelligent sensing hardware layer, an environmental structured feature library, and a core processing and control unit; The intelligent sensing hardware layer includes cameras equipped with high-resolution sensors and motorized gimbals deployed in the measurement field, and a multi-source sensor fusion module; the camera is used to achieve visual fine positioning; the multi-source sensor fusion module is used to form a collaborative sensing architecture of "laser coarse positioning - visual fine positioning"; The environmental structured feature library includes natural feature points and auxiliary calibration objects; The core processing and control unit includes a data processing module and an intelligent scheduling and calibration module. The data processing module is used to perform image processing, deep learning recognition, feature extraction and matching, and multi-view geometric calculation based on the positioning results and natural feature points of the intelligent sensing hardware layer and auxiliary calibration objects, so as to obtain the three-dimensional coordinates of the target feature points of the target under test in the global coordinate system. The intelligent scheduling and calibration module is used to perform dynamic calibration and intelligent scheduling of the camera to realize continuous measurement of the target under test.
2. The system according to claim 1, characterized in that, The intelligent scheduling and calibration module includes a dynamic hybrid calibration engine and a resource intelligent scheduler; A dynamic hybrid calibration engine is used to manage camera parameters, including offline calibration based on calibration boards and online PnP real-time calibration based on environmental feature points, and outputs camera parameters through offline and online calibration. Resource intelligent scheduler: It receives the target position and occlusion prediction information of the target under test in real time, dynamically calculates and directs the most suitable subset of cameras to adjust the gimbal angle and focal length, and performs the measurement task with the best perspective, so as to realize continuous measurement of the target under test under moving occlusion.
3. A visual measurement and localization method based on structured environmental information, wherein the method is implemented using the system described in claims 1-2, characterized in that, The method includes: Collect the three-dimensional coordinates of pre-set structured feature points within the measurement field to construct a feature map in the global coordinate system of the measurement field; Based on feature points with known coordinates observed by several cameras, determine the initial intrinsic and extrinsic parameters of each camera, and unify all cameras to the same world coordinate system; Based on the preset offline calibration benchmark and online PnP calibration, the camera parameters are updated online through self-calibration. Based on the feature map and the camera updated online by self-calibration, combined with deep learning and image processing technology, the target to be measured in the measurement field is identified, and the three-dimensional coordinates of the target feature points in the global coordinate system are obtained. Based on the three-dimensional coordinates of the target feature points in the global coordinate system, a geometric model is used to obtain occlusion prediction information, and the camera is scheduled to achieve continuous measurement of the target.
4. The method according to claim 3, characterized in that, Methods for online self-calibration and updating of camera parameters based on preset offline calibration benchmarks and online PnP calibration include: The PnP algorithm is called to perform online real-time calibration and obtain online calibration results; The online calibration results are combined with the offline calibration results in the offline calibration benchmark to obtain the final calibration result, thereby realizing online self-calibration and updating of camera parameters.
5. The method according to claim 4, characterized in that, Methods for using the PnP algorithm for online real-time calibration and obtaining online calibration results include: Points in the known world coordinate system and its corresponding image plane projection point The projection relationship is as follows: pi = KR(Pi - t); Where i is the number of points in the known world coordinate system. This is the intrinsic parameter matrix of the camera. For extrinsic rotation matrix, This is the extrinsic translation matrix; Construct a quartic polynomial equation: ; in, a , b、 、 All are defined vectors. k 1. k 2 represents the scaling factor; Solve k 1. k 2. Select the solution that minimizes the projection error of all known world coordinate system points from the solution results as the final extrinsic rotation matrix and extrinsic translation matrix.
6. The method according to claim 4, characterized in that, The method of combining online calibration results with offline calibration results from an offline calibration benchmark to obtain the final calibration result and achieving online self-calibration update of camera parameters includes: ; ; in, For the final calibration result, For online calibration results, For offline calibration results, and These are the time-varying calibration result weighting coefficients, satisfying the condition that the sum of the two is 1, a constant. and constant These are used to control the initial weights and the rate of change of the weights, respectively. t For time.
7. The method according to claim 3, characterized in that, Based on feature maps and online self-calibrated updated cameras, combined with deep learning and image processing techniques, methods for identifying targets within the measurement field and obtaining the three-dimensional coordinates of target feature points in the global coordinate system include: When the target enters the measurement field, a lidar is used to scan the target to achieve coarse positioning and obtain the region of interest; Based on the region of interest, the cameras are scheduled after online self-calibration updates to obtain the optimal camera combination; The optimal camera combination is used to acquire images of the target object. Based on the image of the target to be tested, deep learning models and image processing techniques are used to obtain the target feature points of the target to be tested; The target feature points are solved, and the solution results are fused to obtain the three-dimensional coordinates of the target feature points in the global coordinate system.
8. The method according to claim 7, characterized in that, Methods for scheduling cameras based on regions of interest after online self-calibration updates to obtain the optimal camera combination include: Establish a multi-objective optimization decision model and construct the cost function: ; in, These are hyperparameter weighting coefficients, which can be adjusted independently according to the working environment. The angle of rotation of the gimbal. This is the maximum rotatable angle of the gimbal. It is the distance between the camera's imaging point and the point to be measured. This is the camera's maximum working distance. It is the standard deviation of the measurement error. It is the system calibration constant. It's the camera's focal length. It refers to the camera pixel size.
9. The method according to claim 3, characterized in that, Methods for obtaining occlusion prediction information using geometric models based on the three-dimensional coordinates of target feature points in the global coordinate system include: The MES system based on the measurement field acquires real-time coordinates and path planning data of the AGV and robotic arm scheduling system; Based on the real-time coordinates and path planning data of the AGV and robotic arm scheduling system, a bounding box model of the target to be tested is constructed. Based on the three-dimensional coordinates of the target feature points in the global coordinate system, and combined with the bounding box model of the target, the dot product of the normal vector of the bounding box surface of the target and the camera viewing direction vector is obtained. Based on the dot product, determine whether occlusion exists and obtain occlusion prediction information.
10. The method according to claim 9, characterized in that, Methods for determining whether occlusion exists and obtaining occlusion prediction information based on dot product include: If dot product It is 0, and satisfies If so, it is considered an obstruction; If dot product Not zero, and satisfies the parameter If it is occluded, it is considered occluded; otherwise, it is considered unoccluded. The parameter is... t The calculation formula is as follows: ; in, A , B , C , D For plane equation parameters, Define variables.