Operation member state inference method and system

By setting up non-reflective markers in low-pressure sealed equipment to construct a local coordinate system and perform spatial transformation, the instability problem of inferring the state of the operating component under low-frequency vision is solved, achieving stable operation in low-pressure environments and reducing the load on the vision system.

CN122492832APending Publication Date: 2026-07-31HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS
Filing Date
2026-07-03
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In low-pressure enclosed equipment, vision systems struggle to stably infer the state of operating components under low-frequency or low-refresh-rate conditions, resulting in poor stability and reproducibility of the inference of operating component states, and increasing the power consumption and thermal load of the vision system.

Method used

By setting non-reflective markers within the operating area, a real-time local coordinate system is constructed and a spatial transformation relationship is established with the reference local coordinate system. The reference pose state data is mapped to the real-time local coordinate system using a rigid body transformation matrix, thereby achieving stable inference of the operating component's state and avoiding the imaging instability problem of directly identifying the operating component's body.

Benefits of technology

Stable inference of the state of the operating component was achieved under low-frequency vision conditions, reducing the workload and heat dissipation pressure of the vision system and improving the long-term operational stability and execution reliability in a closed low-pressure environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492832A_ABST
    Figure CN122492832A_ABST
Patent Text Reader

Abstract

This application discloses a method and system for inferring the state of an operating component, applied in the field of machine vision technology. It includes: acquiring a visual image of the target operating area where the target operating component is located; detecting non-reflective markers in the visual image to obtain the real-time pose state of each non-reflective marker; constructing a real-time local coordinate system based on the real-time pose states of each non-reflective marker; determining a real-time spatial transformation relationship based on the real-time local coordinate system and a pre-constructed reference local coordinate system; and inferring the real-time pose state of the target operating component based on the pre-calibrated reference pose state of the target operating component using the real-time spatial transformation relationship. By setting non-reflective markers within the target operating area, the real-time pose state of the operating component can be inferred from the real-time pose state of the non-reflective markers under low-frequency visual input conditions, without directly identifying the operating component itself, thus avoiding recognition instability caused by reflection, texture loss, or occlusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine vision technology, and in particular to a method and system for inferring the state of an operating component. Background Technology

[0002] In applications such as aviation, medical, and industrial applications, robots often need to perform remote or automated operations on buttons, switches, knobs, and other actuators within low-pressure, enclosed equipment. These low-pressure, enclosed devices are typically sealed chambers with limited heat dissipation, making it difficult for vision systems to maintain high refresh rates or continuous high loads for extended periods. They may even be limited to single-frame or low-frequency operation, thus affecting the stability of visual input when inferring the state of the actuators.

[0003] To improve the stability of visual input during operator state inference, compensation is usually achieved through multi-angle, multi-frame continuous imaging or by increasing the visual refresh rate. However, this approach still relies on the vision system continuously acquiring visual information about the operator itself. In cases of low-frequency imaging or limited imaging quality, the stability and reproducibility of operator state inference cannot be guaranteed, and it will also significantly increase the power consumption and thermal load of the vision system. Summary of the Invention

[0004] This application provides a method and system for inferring the state of an operating component, thereby addressing the problems of poor stability and reproducibility in inferring the state of an operating component. The technical solution provided by this application is as follows: On the one hand, this application provides a method for inferring the state of an operating component, including: A visual image of the target operating area where the target operating component is located is acquired through a vision system; wherein the visual image contains at least two non-reflective markers set within the target operating area; The visual image is subjected to marker detection and coordinate system transformation to obtain the real-time physical space coordinates of each non-reflective marker, and a real-time local coordinate system of the target operation area is constructed based on the real-time physical space coordinates of each non-reflective marker. Based on the real-time local coordinate system and the reference local coordinate system of the target operating area, the real-time spatial transformation relationship of the target operating area is determined; wherein, the reference local coordinate system is the local coordinate system of the target operating area that is pre-constructed based on the reference physical space coordinates of each non-reflective marker. Based on the pre-calibrated reference pose state data of the target manipulator in the reference local coordinate system, the real-time pose state data of the target manipulator in the real-time local coordinate system is inferred by using real-time spatial transformation relationship.

[0005] The above scheme sets up non-reflective markers in the target operating area, constructs a real-time local coordinate system and establishes a spatial transformation relationship with the reference local coordinate system, thereby mapping the pre-calibrated reference pose state data to the real-time space, realizing stable inference of the state of the operating component under low-frequency vision conditions, and avoiding the imaging instability problem caused by directly identifying the operating component itself.

[0006] Optionally, before acquiring a visual image of the target operating area where the target operating component is located through the vision system, the method further includes: For each operating area within a low-pressure confined space, a reference visual image of the operating area is acquired using a vision system; wherein, the reference visual image contains at least two non-reflective markers placed within the operating area; The reference physical space coordinates of each non-reflective marker are obtained by performing marker detection and coordinate system transformation on the reference visual image. Determine the origin spatial coordinates based on the reference physical spatial coordinates of each non-reflective marker. The X-axis direction is determined based on the preset coding order of two non-reflective markers among each non-reflective marker. The Z-axis direction is determined based on the normal vector of the plane containing each non-reflective marker; The Y-axis direction is determined using the right-hand rule based on the X-axis and Z-axis directions. Using the origin spatial coordinates as the origin, a reference local coordinate system for the operating area is constructed based on the X-axis, Y-axis, and Z-axis directions.

[0007] The above scheme defines the origin and axis of the local coordinate system by pre-setting the coding order and geometric relationship, so that each operating area has a unique and stable spatial reference benchmark, providing a reliable foundation for subsequent geometric consistency mapping.

[0008] Optionally, before acquiring a visual image of the target operating area where the target operating component is located through the vision system, the method further includes: For each manipulator in different operating areas, the actuator at the end of the teaching robot arm is controlled to perform different key operations on the manipulator. According to the execution sequence of the key operations, the relative pose state data of the manipulator in the reference local coordinate system under different key operations is recorded as the reference pose state data. Consistency verification is performed on the reference pose state data of each operand in different operating areas; After confirming that the consistency check has passed, save the reference pose state data of each operand in different operating areas.

[0009] The above scheme records the pose state of the operating component in the reference local coordinate system through teaching and performs consistency verification, ensuring the physical rationality and geometric continuity of the reference data, and providing accurate data support for precise state inference during the operation phase.

[0010] Optionally, a real-time local coordinate system for the target operation area is constructed based on the real-time physical space coordinates of each non-reflective marker, including: Based on the real-time physical spatial coordinates of each non-reflective marker, determine the real-time origin spatial coordinates; The real-time X-axis direction is determined based on the preset coding order of two non-reflective markers among each non-reflective marker. The real-time Z-axis direction is determined based on the normal vector of the plane where each non-reflective marker is located; The real-time Y-axis direction is determined using the right-hand rule based on the real-time X-axis and real-time Z-axis directions. Using the real-time origin spatial coordinates as the origin, and based on the real-time X-axis, real-time Y-axis and real-time Z-axis directions, a real-time local coordinate system for the operating area is constructed.

[0011] The above scheme uses the same construction logic as the reference local coordinate system to restore the real-time local coordinate system, ensuring the consistency of coordinate system definition between the teaching and running phases, thus making mapping restoration based on geometric relationships possible.

[0012] Optionally, based on the real-time local coordinate system and the reference local coordinate system of the target operating region, the real-time spatial transformation relationship of the target operating region is determined, including: The real-time physical space coordinates of each non-reflective marker are converted into real-time three-dimensional space coordinates in a real-time local coordinate system. Based on the real-time three-dimensional spatial coordinates of each non-reflective marker in the real-time local coordinate system and the reference three-dimensional spatial coordinates in the reference local coordinate system, the real-time rigid transformation matrix from the reference local coordinate system to the real-time local coordinate system is solved as the real-time spatial transformation relationship; wherein, the reference three-dimensional spatial coordinates of each non-reflective marker in the reference local coordinate system are obtained by performing coordinate system transformation in advance based on the reference physical spatial coordinates of each non-reflective marker.

[0013] The above scheme describes the spatial transformation relationship between two coordinate systems by solving the rigid transformation matrix, which can accurately compensate for the spatial offset caused by changes in camera viewpoint or overall device attitude, and realize the accurate mapping of the teaching geometric relationship model to the real-time spatial state.

[0014] Optionally, based on the pre-calibrated reference pose state data of the target manipulator in the reference local coordinate system, a real-time spatial transformation relationship is used to infer the real-time pose state data of the target manipulator in the real-time local coordinate system, including: Based on the three-dimensional rotation matrix and three-dimensional translation vector in the real-time spatial transformation relationship, the reference pose state data of the target manipulator in the reference local coordinate system is mapped to the real-time local coordinate system to obtain the real-time pose state data of the target manipulator in the real-time local coordinate system.

[0015] The above scheme uses rigid body transformation to directly spatially map the geometric relationship parameters recorded in the teaching phase, without the need to re-identify the manipulator body. Thus, it can still stably recover the real-time pose of the manipulator in low-frequency visual input and strong reflective environments.

[0016] Optionally, if a non-reflective marker is disposed on the target operating component, the visual image includes at least one non-reflective marker disposed on the target operating component; After acquiring a visual image of the target operating area where the target operating component is located through the vision system, the following is also included: Real-time pose data of each non-reflective marker is obtained by performing marker detection and coordinate system transformation on the visual image. Based on the real-time pose data of each non-reflective marker and the pre-calibrated baseline pose data of each non-reflective marker under different key operations, the real-time pose data of the target manipulator under different key operations is inferred.

[0017] The above scheme is designed for operating parts with independent markers. It directly uses the attitude changes of the markers to infer the state, so that the gear position judgment of rotating or pulling operating parts such as knobs is independent of the regional point cloud, which further improves the robustness and accuracy of state inference.

[0018] Optionally, the method for inferring the state of an operating component provided in this application further includes: If the real-time pose state data of the target manipulator in the real-time local coordinate system includes the real-time pose state data of the target manipulator under a critical operation, then the critical operation is performed on the target manipulator based on the real-time pose state data of the target manipulator under the critical operation. If the real-time pose state data of the target manipulator in the real-time local coordinate system includes the real-time pose state data of the target manipulator under at least two key operations, then according to the preset key operation execution order, based on the real-time pose state data of the target manipulator under each key operation, each key operation is executed sequentially on the target manipulator.

[0019] The above scheme directly uses the state inference results to guide the robotic arm in performing operations, forming a complete closed loop from visual perception to action execution, which enables automated operation tasks under low-frequency vision conditions to proceed in an orderly manner.

[0020] Optionally, when executing each key operation sequentially on the target operator based on the real-time pose state data of the target operator under each key operation according to a preset key operation execution order, the method further includes: The actual pose state data of the target manipulator after the current key operation is performed is compared with the reference pose state data; If the comparison result is determined to be satisfactory, the next critical operation is executed; if the comparison result is determined to be unsatisfactory, the real-time pose state data of the target manipulator under the current critical operation is adjusted based on the visual image until the comparison result between the actual pose state data of the target manipulator and the reference pose state data after the current critical operation is executed is satisfactory, and then the next critical operation is executed.

[0021] The above solution, by introducing a status comparison and feedback adjustment mechanism after operation execution, ensures that each key operation achieves the expected goal before proceeding to the next node, significantly improving the execution reliability and success rate of multi-step operation tasks under low-frequency vision conditions.

[0022] On the other hand, this application provides an operating element state inference system, including: The image acquisition unit is used to acquire a visual image of the target operating area where the target operating component is located through a vision system; wherein the visual image contains at least two non-reflective markers set within the target operating area; The real-time construction unit is used to perform marker detection and coordinate system transformation on the visual image to obtain the real-time physical space coordinates of each non-reflective marker, and to construct the real-time local coordinate system of the target operation area based on the real-time physical space coordinates of each non-reflective marker. The spatial transformation unit is used to determine the real-time spatial transformation relationship of the target operating area based on the real-time local coordinate system and the reference local coordinate system of the target operating area; wherein, the reference local coordinate system is a local coordinate system of the target operating area that is pre-constructed based on the reference physical spatial coordinates of each non-reflective marker. The state inference unit is used to infer the real-time pose state data of the target manipulator in the real-time local coordinate system based on the pre-calibrated reference pose state data of the target manipulator in the reference local coordinate system and by adopting the real-time spatial transformation relationship.

[0023] The above system achieves stable inference of the state of the operating component under low-frequency visual input conditions through the collaborative work of each unit, avoiding the imaging instability problem caused by directly recognizing the operating component itself. It is suitable for automated operation scenarios that operate for a long time in a closed low-pressure environment with limited heat dissipation.

[0024] The beneficial effects of this application are as follows: (1) By setting at least two non-reflective markers in the target operating area, a local geometric reference frame based on non-reflective markers can be constructed, so that the real-time pose state inference of the target operating part is based on a stable spatial reference. This fundamentally eliminates the dependence on the texture and appearance features of the target operating part such as buttons, switches, and knobs. Since non-reflective markers are not affected by metal reflection, they can maintain stable recognition even in strong reflection environments, thereby avoiding the problems of structured light stripe distortion and missing depth data.

[0025] (2) By constructing a local geometric reference frame, under single-frame or low-frequency visual input conditions, only non-reflective markers need to be detected to recover the real-time local coordinate system. The real-time pose state of the target operation is geometrically mapped by solving the real-time spatial transformation relationship between the target and the reference local coordinate system. There is no need for continuous high-frequency visual refresh. This not only significantly reduces the workload and heat dissipation pressure of the vision system and improves the long-term operational stability of the vision system in a closed low-pressure environment, but also is not affected by changes in camera posture, robotic arm movement or local mechanism action. This enables the robotic arm to stably perform multi-step, multi-node equipment operation tasks under low visual frequency conditions, significantly improving the execution reliability and system adaptability in a closed low-pressure environment.

[0026] Other features and advantages of this application will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0027] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram outlining the operation state inference method in the embodiments of this application; Figure 2 This is a schematic diagram illustrating the construction process of the reference local coordinate system in an embodiment of this application; Figure 3 This is a schematic diagram of the operator calibration process within each operating area in the embodiments of this application; Figure 4 This is a schematic diagram illustrating the construction process of the real-time local coordinate system in an embodiment of this application; Figure 5 This is a schematic diagram illustrating the process of determining real-time spatial transformation relationships in the embodiments of this application; Figure 6 This is a schematic diagram illustrating the inference process of real-time pose state data in an embodiment of this application; Figure 7 This is a schematic diagram of the composition structure of the operator state inference system in the embodiments of this application; Figure 8 This is a schematic diagram of the hardware structure of the robot control device in the embodiments of this application. Detailed Implementation

[0028] To make the objectives, technical solutions, and beneficial effects of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] This application provides a method for inferring the state of a manipulator, applied to a manipulator state inference system. This system can be mounted in a robot to provide the robot with real-time manipulator state data of the target manipulator, guiding the actuator at the end effector of the robot's robotic arm to perform remote or automated operations on the target manipulator based on the real-time manipulator state data. (See also...) Figure 1 As shown, the general flow of the operation component state inference method provided in this application embodiment is as follows: Step 101: Acquire a visual image of the target operating area where the target operating component is located using a vision system; wherein the visual image contains at least two non-reflective markers set within the target operating area.

[0030] Specifically, the target operating area refers to the operating area inside the low-pressure sealed equipment where the target operating component is located. This target operating area includes a rigid structure that does not deform or shift when the target operating component is subjected to critical operations, such as a separate metal panel or a fixed equipment compartment shell. This rigid structure serves as a stable spatial reference, and at least two non-reflective markers are set on the rigid structure.

[0031] At least two non-reflective markers are pre-installed on a rigid structure within the target operating area. "Non-reflective markers" refer to planar or three-dimensional markers made of matte, diffuse-reflective materials with high-contrast patterns. Examples include sandblasted or frosted aluminum plates, ceramic sheets coated with matte ink, or patches made of special diffuse-reflective polymers, printed with easily visually detectable patterns such as checkerboard patterns, QR codes, concentric rings, or coded dot matrices. The emphasis on "non-reflective" is crucial because in aerospace and industrial settings, low-pressure sealed equipment panels and housings are typically made of metal or polished materials. Strong reflectivity in these materials can cause distortion of structured light stripes, loss of depth data, and unstable edge contours. Using ordinary reflective markers would create new reflective sources, failing to address the root cause of the problem. Non-reflective markers, however, maintain stable imaging quality even against strong reflective backgrounds, ensuring high marker detection success rates and accurate spatial coordinates.

[0032] The vision system can be a structured light camera, an RGB-D camera, a binocular camera, or any imaging device capable of simultaneously acquiring two-dimensional texture information and three-dimensional depth information. The visual image acquired by the vision system includes at least two non-reflective markers set on a rigid structure within the target operating area, and may also include at least one target operating component such as a button, toggle switch, or knob.

[0033] Step 102: Perform marker detection and coordinate system transformation on the visual image to obtain the real-time physical space coordinates of each non-reflective marker, and construct the real-time local coordinate system of the target operation area based on the real-time physical space coordinates of each non-reflective marker.

[0034] Specifically, the construction of the real-time local coordinate system comprises two closely interconnected sub-steps. First is marker detection and spatial coordinate acquisition: After the vision system acquires the visual image, it uses a marker detection algorithm to locate the two-dimensional pixel position of each non-reflective marker in the visual image. Combining this with depth information or structured light calculation results, the two-dimensional pixel coordinates are converted into three-dimensional physical space coordinates referenced to the camera coordinate system using camera intrinsics and the imaging model. For example, for coded markers, the algorithm can not only locate their center point but also obtain their unique identifier through decoding, thus distinguishing different markers. Second is the construction of the local coordinate system: Based on the real-time physical space coordinates of each detected non-reflective marker, a local orthogonal coordinate system corresponding to the target operation area is established in three-dimensional space, called the "real-time local coordinate system." The origin, X-axis, Y-axis, and Z-axis of this coordinate system are defined by the spatial position and geometric relationship of the markers. Therefore, this real-time local coordinate system changes with the pose of the target operation area within the camera's field of view. This can be understood as follows: The real-time local coordinate system is a mathematical description framework of the target operating area in three-dimensional space at the "current moment". It is dynamically updated as the camera viewpoint changes or the overall posture of the device changes, but it always maintains a fixed geometric relationship with the rigid structure of the target operating area.

[0035] Step 103: Determine the real-time spatial transformation relationship of the target operating area based on the real-time local coordinate system and the reference local coordinate system of the target operating area; wherein, the reference local coordinate system is the local coordinate system of the target operating area pre-constructed based on the reference physical space coordinates of each non-reflective marker.

[0036] Specifically, the reference local coordinate system is a local coordinate system established during the teaching or initialization phase based on the physical spatial coordinates of each non-reflective marker at a certain reference time (e.g., when the equipment is first installed or during initial calibration). The reference local coordinate system represents the spatial reference frame of the target operating area under "standard conditions" and is permanently stored as a reference for subsequent comparisons.

[0037] The difference between the real-time local coordinate system and the reference local coordinate system lies in the fact that the former describes the spatial pose of the "current" target operating area, while the latter describes the spatial pose of the "reference" target operating area. The connection between the two is that they are constructed on the same physical rigid structure using the same set of non-reflective markers according to the same geometric rules; therefore, there must be a rigid body transformation relationship between them—that is, a rotation and translation transformation in three-dimensional space. Using the physical spatial coordinates of each non-reflective marker in both coordinate systems, this rigid body transformation relationship can be solved, thus serving as a real-time spatial transformation relationship. This real-time spatial transformation relationship quantitatively describes the spatial displacement and attitude rotation of the target operating area from the reference state to the current state.

[0038] Step 104: Based on the pre-calibrated reference pose state data of the target manipulator in the reference local coordinate system, the real-time pose state data of the target manipulator in the real-time local coordinate system is inferred by using the real-time spatial transformation relationship.

[0039] Specifically, during the teaching phase, the pose state data of the target actuator (such as a button, switch, or knob) in the reference local coordinate system has been recorded. For example, the three-dimensional coordinates of the center point of the end of the button when it is in the "pressed" state, and the quaternion of the pose of the marker of the knob when it is in "position 1". These data are called "reference pose state data", which describe the relative spatial relationship between the actuator and the reference local coordinate system.

[0040] During operation, real-time spatial transformation relationships—typically a 3×3 rotation matrix R and a 3×1 translation vector T—are used to map the reference pose state data from the reference local coordinate system to the real-time local coordinate system. For example, if the pressed state coordinates of a button in the reference local coordinate system are... Then its corresponding coordinates in the real-time local coordinate system This mapping yields This refers to the inferred real-time pose state data. It's called "inference" because we don't currently "see" the target actuator and "identify" its real-time physical spatial position. Instead, it's based on a geometric assumption: the relative spatial relationship between the actuator and the rigid structure of the target operating area remains unchanged during the teaching and running phases. As long as this assumption holds—which is usually true in actual rigid panel structures—then by recovering the real-time spatial pose of the target operating area, the real-time spatial pose of the actuator can be indirectly derived. This "inference" mechanism allows the system to completely avoid the need for visual recognition of the actuator itself, thus fundamentally eliminating reliance on the actuator's texture, contour, reflective properties, and other appearance features.

[0041] Through the above four steps, a complete technical path of "marker detection → real-time coordinate system construction → spatial transformation determination → pose inference" is established. Under low-frequency or even single-frame visual input conditions, as long as non-reflective markers within the target operating area are successfully detected, the real-time local coordinate system of that area can be recovered. Then, through rigid body transformation, all the pose knowledge of the operating parts accumulated during the teaching phase can be transferred to the current spatial state, achieving stable recovery of the operating point position and reliable inference of the state. This mechanism significantly reduces the requirements for visual refresh rate, enabling the vision system to operate stably for a long time in a closed, low-pressure environment with limited heat dissipation.

[0042] In this embodiment, the low-pressure sealed equipment typically possesses a low-pressure sealed space. This low-pressure sealed space refers to the internal environment of low-pressure sealed equipment in scenarios such as aviation, medical, and industrial applications, where the internal air pressure is lower than standard atmospheric pressure and the structure is a closed or semi-closed cavity. Examples include aircraft avionics equipment bays, sealed air chambers of high-voltage switchgear, and vacuum coating chambers. The low-pressure sealed space contains the operating panel of the low-pressure sealed equipment. The principle for dividing the operating area is to divide the rigid components of the operating panel that do not deform or shift when performing critical operations on the target operating component into multiple operating areas. For example, an aluminum alloy panel fixed to the equipment frame with bolts can serve as an operating area; a stainless steel bracket welded to the cabin structure can also serve as an operating area. At least two non-reflective markers are pre-set on the rigid structure of each operating area. These non-reflective markers maintain a fixed geometric relationship with the rigid structure during subsequent operation of the operating component and will not shift due to button pressing, switch toggling, or knob rotation.

[0043] The construction of the reference local coordinate system for each operating area is typically completed during the teaching phase to provide a stable and unique spatial reference for solving real-time spatial transformation relationships in the subsequent operation phase. The construction process of the reference local coordinate system for each operating area in a low-pressure confined space is described in detail below. (See also...) Figure 2 As shown, during the teaching phase, the process of constructing the reference local coordinate system for each operating area in the low-pressure confined space is as follows: Step 201: Acquire a reference visual image of the operating area using a vision system; wherein the reference visual image contains at least two non-reflective markers placed within the operating area.

[0044] Specifically, the reference visual image acquired by the vision system includes at least two non-reflective markers set on a rigid structure within the operating area, and may also include at least one operating element such as a button, toggle switch, or knob.

[0045] Step 202: Perform marker detection and coordinate system transformation on the reference visual image to obtain the reference physical space coordinates of each non-reflective marker.

[0046] Specifically, after acquiring a reference visual image, the vision system uses a marker detection algorithm to locate the two-dimensional pixel coordinates of each non-reflective marker in the reference visual image. Combining this with depth information or structured light calculation results, the two-dimensional pixel coordinates are converted into three-dimensional physical space coordinates referenced to the camera coordinate system using a camera intrinsic parameter model. The three-dimensional physical space coordinates of each non-reflective marker are called the "reference physical space coordinates," recording the precise position of each non-reflective marker in three-dimensional space at the teaching reference moment. Because the non-reflective markers are made of matte, diffuse reflective materials, they do not produce specular reflections even against metallic or polished backgrounds; therefore, the detected center point coordinates have sub-millimeter repeatability accuracy.

[0047] Step 203: Determine the origin spatial coordinates based on the reference physical space coordinates of each non-reflective marker.

[0048] Specifically, the methods for determining the spatial coordinates of the origin include the direct selection method and the geometric relationship definition method, which can be flexibly selected according to actual engineering needs. Among them: The direct selection method involves choosing the physical coordinates of a single non-reflective marker from all non-reflective markers within the operating area, according to preset rules, as the origin. For example, it can be agreed that the center point of the marker with the smallest ID is always chosen as the origin; or, if the operating area is a rectangular panel, the marker located at the lower left corner of the panel can be chosen as the origin. Assuming the reference physical coordinates of marker A are (10mm, 20mm, 0mm), then the origin O is determined to be (10mm, 20mm, 0mm). The advantage of this method is its extreme simplicity of calculation, making it suitable for scenarios with a small number of markers and a regular arrangement.

[0049] The geometric relationship definition method uses the physical spatial coordinates of multiple non-reflective markers to define a single physical spatial coordinate as the origin through geometric operations. For example, the geometric operation is to take the midpoint: when there are two non-reflective markers A and B in the operating area, the midpoint between the coordinates of A and B is taken as the origin, i.e., O = (A + B) / 2. When there are three or more non-collinear non-reflective markers in the operating area, the centroid of the polygon formed by these markers can be taken as the origin. For example, if three non-reflective markers form a triangle, the centroid coordinates are the arithmetic mean of the physical spatial coordinates of the three vertices. The advantage of the geometric relationship definition method is that the origin does not depend on the absolute positional accuracy of a single marker, but is determined by multiple non-reflective markers together, resulting in better fault tolerance and stability. Even if a non-reflective marker experiences slight physical displacement or dirt due to long-term use, the origin defined by the geometric relationship is far less affected than that of the direct selection method.

[0050] It should be understood that the above two methods are only illustrative examples. In actual deployment, other geometric rules can also be used to define the origin. For example, the weighted average of the physical spatial coordinates of the non-reflective sign can be taken, or a feature point of the least squares fitted plane can be taken, as long as the rule can uniquely determine a spatial point based on the physical spatial coordinates of the sign.

[0051] Step 204: Determine the X-axis direction based on the preset coding order of two non-reflective markers among the various non-reflective markers.

[0052] Specifically, each non-reflective marker has a unique code, such as a unique QR code pattern or coded dot matrix printed on it. When the vision system detects the location of the non-reflective marker, it can decode it to obtain its identification. These codes are usually represented by sequences of numbers or letters, such as ID=1, ID=2, ID=3, etc.

[0053] The preset encoding order refers to a manually agreed-upon encoding arrangement rule during system initialization. For example, it can be agreed that the direction is determined by the ascending order of the encoding IDs. Specifically, from all non-reflective markers within the operating area, two non-reflective markers are selected. Following the ascending order of their encoding IDs, the direction vector pointing from the center point of the marker with the smallest ID to the center point of the marker with the largest ID is taken as the positive direction of the X-axis. Assuming there are two non-reflective markers within the operating area with IDs 1 and 2, and their reference physical space coordinates P1 and P2 respectively, then the X-axis direction vector is the normalized vector of (P2-P1). If there are three non-reflective markers within the operating area with IDs 1, 2, and 3, it can be agreed that IDs 1 and 2 are selected to determine the X-axis, while ID 3 is used for subsequent verification or Z-axis calculation.

[0054] By employing a preset encoding order to determine the X-axis direction, it is ensured that regardless of the camera's viewing angle during the teaching or operational phases, as long as the marker is successfully detected and decoded, the positive direction of the X-axis remains uniquely determined. This "encoding-driven" direction definition method fundamentally avoids directional ambiguity caused by accidental symmetry in the spatial position of the marker or uncertainty in the visual detection order. Without the constraint of an encoding order, relying solely on the spatial position of the marker to define the X-axis could lead to a serious error: the X-axis points to the left during the teaching phase and to the right during the operational phase due to camera angle rotation. This would result in a mirror flip of the entire coordinate system, rendering the operation point inference completely ineffective.

[0055] Step 205: Determine the Z-axis direction based on the normal vector of the plane where each non-reflective marker is located.

[0056] Specifically, all non-reflective markers within the operating area can be placed on the same plane, such as the outer surface of the panel. By performing plane fitting on the reference physical space coordinates of these non-reflective markers, the normal vector of the plane can be determined. Plane fitting can employ methods such as least squares or principal component analysis. In practical applications, the positive Z-axis direction can be defined as the normal vector direction pointing outwards from the operating area, i.e., towards the side of the robotic arm's operating space. This design gives the Z-axis direction a clear physical meaning: perpendicular to the panel surface, pointing in the direction in which the manipulator is operated.

[0057] By defining the Z-axis using a planar normal vector, the spatial orientation of the rigid structure within the operating area can be directly reflected. When the low-pressure sealed equipment tilts or the camera's viewing angle changes, the normal vector of the plane containing the non-reflective marker changes accordingly, ensuring that the Z-axis remains perpendicular to the panel surface. This "structure-driven" Z-axis definition method guarantees the inherent consistency between the local coordinate system and the physical geometry of the operating area.

[0058] Step 206: Determine the Y-axis direction based on the X-axis and Z-axis directions using the right-hand rule.

[0059] After determining the X and Z axes, the direction of the Y axis is uniquely determined using the right-hand rule. Specifically, the right-hand rule states: if the index finger of your right hand points in the positive direction of the X-axis and the middle finger points in the positive direction of the Z-axis, then the direction pointed to by your thumb is the positive direction of the Y-axis. This can be represented using the cross product of vectors as: Y = Z × X.

[0060] It's important to note that the Y-axis is not defined independently; it is derived from the X and Z axes using the right-hand rule. This is to ensure the orthogonality and right-hand property of the coordinate system. Orthogonality means that the X, Y, and Z axes are pairwise perpendicular, and right-hand property means that the coordinate system conforms to the right-hand rule. In an orthogonal and right-handed coordinate system undergoing rigid body transformations, the determinant of the rotation matrix is ​​+1, ensuring the invariance of volume and orientation during the transformation. If a left-handed system or a non-orthogonal coordinate axis definition is used, the subsequent solution for the rigid transformation matrix becomes complex and prone to errors.

[0061] Step 207: Using the origin spatial coordinates as the origin, construct a reference local coordinate system for the operating area based on the X-axis, Y-axis, and Z-axis directions.

[0062] By using the origin spatial coordinates as the origin and based on the X-axis, Y-axis and Z-axis directions, a complete three-dimensional orthogonal coordinate system {O,X,Y,Z} with geometric uniqueness and directional consistency can be constructed for the operating region, namely the reference local coordinate system.

[0063] The aforementioned logic for constructing the local reference coordinate system ensures the geometric uniqueness of the local reference coordinate system for each operating area. Geometric uniqueness means that for the same operating area, as long as the physical position and encoding of the non-reflective markers remain unchanged, the resulting coordinate system will be mathematically identical regardless of when, by what camera, or from what angle the system is constructed following the above steps. This is because the origin is uniquely determined by the spatial coordinates of the markers, the X-axis by the encoding order, the Z-axis by the plane normal vector, and the Y-axis by the right-hand rule; there is no subjective choice or random factor involved in the entire process.

[0064] Simultaneously, the aforementioned logic for constructing the reference local coordinate system ensures the directional consistency between the reference local coordinate system and the real-time local coordinate system for each operating area. Directional consistency means that the reference local coordinate system constructed during the teaching phase and the real-time local coordinate system constructed during the running phase (described in subsequent embodiments) use the exact same axis definition rules. Therefore, the X, Y, and Z axes of the two coordinate systems are physically corresponding, and there will be no reversal or exchange of axis directions. This directional consistency is a prerequisite for obtaining a correct and stable rotation matrix when subsequently solving the rigid transformation matrix. If the coordinate axis definition rules are different between the teaching and running phases, even if the origin positions completely coincide, the solved rotation matrix may contain a 180-degree directional flip, causing the operating point to be mapped to a completely incorrect position.

[0065] Furthermore, by constructing the aforementioned local coordinate system, a local coordinate system is built for each operating area. This enables the establishment of a stable, unique, and oriented reference frame for each operating area. This reference frame is independent of the state of the operating component and does not change with the actions of the operating component, such as pressing a button or rotating a knob. This provides a solid geometric foundation for subsequent teaching recording, spatial transformation solving, and pose inference.

[0066] Furthermore, during the teaching phase, after establishing the reference local coordinate system for each operating area, the reference pose state data for each operating component can be calibrated, and these data can be systematically verified for consistency to ensure the accuracy and reliability of state inference in subsequent operation phases. (See also...) Figure 3 As shown, the calibration process for the operating components within each operating area of ​​a low-pressure confined space is as follows: Step 301: Control the actuator at the end of the teaching robot arm to perform different key operations on the manipulator, and record the relative pose state data of the manipulator in the reference local coordinate system under different key operations as the reference pose state data according to the execution sequence of the key operations.

[0067] Specifically, a teach pendant robotic arm refers to a precisely calibrated robotic arm with high repeatability and positioning accuracy. Its end effector is equipped with an actuator that matches the manipulated object, such as a cylindrical pressure head for pressing buttons, a fork-shaped lever for toggling switches, or a rotating gripper for holding knobs. The teach pendant robotic arm is typically guided by an operator via a teach pendant or automatically performs a series of actions according to a preset program.

[0068] A critical operation refers to a representative state-switching action of an actuator within its normal operating range. The definition of a critical operation differs depending on the type of actuator. For button-type actuators, critical operations typically include two states: "pressed" and "released." During teaching, the pressure head at the end of the teaching arm is first suspended at a safe height above the button, then slowly lowered along the Z-axis until the button touches the bottom. The three-dimensional coordinates of the end effector center point (or the pressure head end face center point) in the reference local coordinate system are recorded as the reference pose data for the "pressed" state. Subsequently, the robotic arm is slowly raised until the button is fully released and reset. The three-dimensional coordinates of the end effector center point are recorded as the reference pose data for the "released" state. For lever-type actuators, critical operations typically include two end states: "on" and "off." During teaching, the fork-shaped lever at the end of the robotic arm engages the switch handle, which is then moved to its two extreme positions. The three-dimensional coordinates of the end effector center point and the direction vector of the switch handle are recorded for each end state. For knob-type actuators, critical operations typically include rotating to various positions. During teaching, the rotating gripper at the end of the robotic arm is controlled to hold the knob and rotate it to each preset position in sequence. The three-dimensional coordinates of the center point of the end effector and the attitude parameters (such as quaternions or Euler angles) of the knob marker are recorded at each position.

[0069] Recording reference pose state data according to the execution sequence of critical operations means that if an operation task contains multiple consecutive critical operations, such as first rotating the knob to position 2, then pressing the button, and finally toggling the switch to the "on" position, then during the teaching phase, the reference pose state data corresponding to each critical operation must be recorded in strict order, and the sequence information must be saved together so that the operation process can be reproduced in the same order during the operation phase, ensuring the logical correctness of the task execution.

[0070] Step 302: Perform consistency verification on the reference pose state data of each operand in different operating areas.

[0071] Specifically, consistency verification is a crucial step that distinguishes it from simple data recording. If data is recorded without verification, human error, robotic arm positioning deviations, or structural abnormalities in the manipulator itself that may be introduced during the teaching process will not be detected in time, resulting in defects in the baseline data and consequently affecting the accuracy of state inference during operation. Consistency verification includes at least the following three types: The first type is physical range rationality verification. Each operating component has its inherent physical travel range. For example, a standard industrial button typically has a pressing stroke between 0 and 10 millimeters; a toggle switch typically has a swing angle within ±30 degrees; and a multi-position knob typically has a rotation angle range of 0 to 300 degrees, with fixed angular intervals between each position. During verification, the system reads the reference pose data recorded during the teaching phase, extracts the displacement or rotation of the operating component, and compares it with a preset physical range threshold. For example, if the Euclidean distance between the "pressed" and "released" coordinates of a button is 15 millimeters, exceeding the reasonable range of 0 to 10 millimeters, the system will determine the data is abnormal and issue an alarm prompting the operator to re-teach. This abnormality may be due to the robotic arm not being fully pressed during teaching or a jamming malfunction in the button itself. Through physical range rationality verification, erroneous data that clearly does not conform to common sense can be eliminated at the source.

[0072] The second type is path geometric continuity verification. For operable components involving multiple consecutive key operations, their teaching paths must meet certain geometric continuity constraints. For example, for a linear sliding switch, its multiple teaching points should be approximately collinear; for a lever swinging in the same plane, its multiple teaching points should be approximately coplanar. During verification, the system performs least-squares linear or planar fitting on the coordinates of multiple points in the teaching record, calculating the deviation of each point from the fitted line or plane. If the maximum deviation exceeds a preset threshold (e.g., 1 mm), the path geometric continuity is deemed unsatisfactory. This anomaly may indicate a deviation in the robot arm's trajectory during teaching, or problems such as looseness or deformation of the operable component itself. Path geometric continuity verification ensures that the teaching path is geometrically smooth and reasonable, preventing the robot arm from performing operations along a distorted path during operation, which could lead to jamming or equipment damage.

[0073] The third type is the uniqueness verification of knob position mapping. This verification is specifically for knob-type operating devices. Knobs typically have multiple discrete positions, each corresponding to a specific rotation angle. The data recorded during the teaching phase must ensure a one-to-one correspondence between each position and a unique attitude parameter (such as the rotation angle around the Z-axis), and situations where two different positions correspond to the same or too close attitude angles are not allowed. During verification, the system calculates the angle difference between adjacent positions. If the angle difference is less than the preset minimum resolution angle (e.g., 5 degrees), the position mapping is considered ambiguous, and re-teaching is required. In addition, the system also verifies whether the rotation direction of the knob is consistent with the direction of position increment to prevent the position order from being reversed during teaching. Through the uniqueness verification of knob position mapping, it can be ensured that the system can uniquely determine the current position based on the real-time attitude of the knob marker during the operation phase, avoiding ambiguity in state judgment.

[0074] Step 303: After confirming that the consistency check has passed, save the reference pose state data of each operand in different operating areas.

[0075] The reference pose state data that passes all the above consistency checks is saved to the system's non-volatile memory, associated with the corresponding operation area ID, operation item ID, and key operation type. This data constitutes the "knowledge base" for state inference in subsequent operation phases. If the reference data of an operation item fails the check, the system will refuse to save it and prompt the operator to re-teach the operation item until the data meets all the check conditions.

[0076] By recording and verifying the consistency of the reference pose state data for each manipulator within different operating areas during the teaching phase, a rigorously validated, physically sound, and geometrically consistent reference pose state dataset can be established for each manipulator before formal operation. This dataset is tightly bound to the reference local coordinate system, forming a reliable foundation for the "perception-inference-execution" closed loop during the operational phase. Moreover, by introducing a multi-dimensional automatic verification mechanism, the risk of unreliable reference data due to human teaching errors or equipment malfunctions is significantly reduced, thereby improving the long-term operational stability and success rate of the entire system under low-frequency vision conditions.

[0077] The following details the construction of the real-time local coordinate system of the target operating area based on the real-time physical spatial coordinates of each non-reflective marker during the operational phase. During the operational phase, after the vision system acquires a visual image of the target operating area and detects the real-time physical spatial coordinates of each non-reflective marker, the construction of the real-time local coordinate system of the target operating area based on these coordinates is described in detail below. Figure 4 As shown, the system can adopt, but is not limited to, the following methods: Step 401: Determine the real-time origin spatial coordinates based on the real-time physical spatial coordinates of each non-reflective marker.

[0078] Similar to the method for determining the origin of the aforementioned reference local coordinate system, the origin of the real-time local coordinate system can also be determined using either the direct selection method or the geometric relationship definition method, but the same rules must be followed as in the teaching phase. For example, if the midpoint of the coordinates of two markers A and B is used as the origin when constructing the reference local coordinate system in the teaching phase, i.e., O = (A + B) / 2, then the same midpoint rule must be used to determine the real-time origin O' when constructing the real-time local coordinate system in the runtime phase. Assuming the real-time physical space coordinates of markers A and B in the current frame are A' and B' respectively, then the real-time origin O' = (A' + B') / 2. If the coordinates of the marker with the smallest encoded ID are directly selected as the origin in the teaching phase, then the real-time coordinates of the marker with the smallest encoded ID are also selected as the real-time origin in the runtime phase. This consistency of rules ensures that the real-time origin and the reference origin physically correspond to the same geometric feature point on the rigid structure of the target operating area—even though the absolute coordinates of this feature point in the camera coordinate system have changed due to changes in camera viewpoint or overall device displacement.

[0079] It is important to emphasize that consistency in the rules for determining the origin is a prerequisite for ensuring that only rigid body transformations (rotation plus translation) exist between the real-time local coordinate system and the reference local coordinate system, without scaling, shearing, or nonlinear distortion. If the midpoint method is used during the teaching phase but the direct selection method is used during the runtime phase, the origins of the two coordinate systems will correspond to different physical points on the rigid structure. The transformation matrix subsequently solved will contain an additional offset, leading to systematic errors in the mapping of the operation points.

[0080] Step 402: Determine the real-time X-axis direction based on the preset coding order of two non-reflective markers among the various non-reflective markers.

[0081] The method for determining the X-axis direction is the same as that for the aforementioned reference local coordinate system. When determining the X-axis direction of the real-time local coordinate system, two markers are selected from the detected non-reflective markers according to a preset coding order—for example, always selecting the two markers with the smallest and second smallest coding IDs—and then the direction vector from the real-time physical space coordinates of the marker with the smallest ID to the real-time physical space coordinates of the marker with the largest ID is calculated, normalized, and used as the positive direction of the real-time X-axis.

[0082] Using a preset encoding order to determine the real-time X-axis direction has technical significance far beyond simply "defining a direction." During operation, the camera may capture images of the target operating area from any angle, and the relative positions of markers in the image may be rotated, mirrored, or even inverted. If the X-axis is defined solely based on the spatial position of the markers—for example, "taking the direction from the leftmost marker to the rightmost marker"—then when the camera shoots from the back of the panel, the definitions of "left" and "right" will be completely reversed, causing the real-time X-axis to be opposite to the reference X-axis. However, the encoding order is an inherent property of the markers and does not change with the camera's viewing angle. Therefore, the rule of "from small ID to large ID" ensures that the real-time X-axis and the reference X-axis physically point in the same direction from any viewing angle, thus avoiding the catastrophic error of coordinate system orientation reversal.

[0083] Step 403: Determine the real-time Z-axis direction based on the normal vector of the plane where each non-reflective marker is located.

[0084] Specifically, after performing plane fitting on the real-time physical space coordinates of each non-reflective marker and solving for the real-time normal vector of the plane where each non-reflective marker is located, the positive direction of the real-time Z-axis of the real-time local coordinate system is determined to be the direction of the normal vector pointing to the outside of the target operation area, i.e. towards the side of the robotic arm operation space, in the same way as the aforementioned reference local coordinate system.

[0085] It's important to note that the accuracy of plane fitting directly affects the accuracy of the Z-axis direction, and consequently, the attitude accuracy of the entire real-time local coordinate system. In practical engineering, if only two non-reflective markers are placed within the target operating area, a plane cannot theoretically be uniquely determined. In this case, the system can use point cloud data of the area surrounding the markers in the depth information to assist in plane fitting, or assume that the plane containing the markers is approximately perpendicular to the camera's optical axis to provide an initial estimate. Of course, a more robust approach is to place three or more non-collinear non-reflective markers within each operating area. This way, the plane can be accurately fitted using only the spatial coordinates of the markers themselves, without relying on additional point cloud data.

[0086] Step 404: Determine the real-time Y-axis direction based on the real-time X-axis and Z-axis directions using the right-hand rule.

[0087] With the real-time X-axis and real-time Z-axis already determined, the direction of the real-time Y-axis is uniquely determined using the right-hand rule, i.e., Y' = Z' × X'. This step is exactly the same as the method for determining the reference Y-axis of the aforementioned local coordinate system, ensuring that the real-time coordinate system and the reference coordinate system are both right-handed orthogonal coordinate systems.

[0088] Step 405: Using the real-time origin spatial coordinates as the origin, construct a real-time local coordinate system for the operating area based on the real-time X-axis direction, real-time Y-axis direction, and real-time Z-axis direction.

[0089] Through the aforementioned logic for constructing the real-time local coordinate system, a complete real-time local coordinate system {O',X',Y',Z'} is built for the target operating region. This real-time local coordinate system is mathematically isomorphic to the aforementioned reference local coordinate system {O,X,Y,Z}—both use the same origin definition rules, the same axis definition rules, and the same right-hand rule. The only difference is that the reference local coordinate system describes the spatial pose of the target operating region at the teaching reference time, while the real-time local coordinate system describes the spatial pose of the target operating region at the current time.

[0090] Because of the complete symmetry in the construction logic between the real-time local coordinate system and the reference local coordinate system, the spatial transformation relationship between them must be a strict rigid body transformation—that is, it only includes rotation and translation, and does not include scaling, shearing, or twisting. This property mathematically guarantees the uniqueness and stability of the rigid transformation matrix solved in subsequent steps. Specifically, if different coordinate system construction rules are used in the teaching and running phases, even if the origins of the two coordinate systems happen to coincide and their axes happen to be consistent, the transformation relationship between them may still contain non-rigid body components, leading to unpredictable errors when mapping the operation points. However, in the embodiments of this application, by forcibly maintaining the symmetry of the construction logic, this source of error is fundamentally eliminated.

[0091] Through the aforementioned real-time local coordinate system construction logic, the system can reconstruct a real-time reference frame that is completely consistent with the definition of the baseline local coordinate system for each frame of visual image input. This frame provides a precise geometric basis for subsequent spatial transformation relationship solving and operation point mapping, and is the core guarantee for achieving high-precision operation point inference under low-frequency visual conditions.

[0092] The following section describes in detail how, during the operational phase, the real-time spatial transformation relationships of the target operational area are determined based on the real-time local coordinate system and the reference local coordinate system. (See also...) Figure 5 As shown, the process for determining the real-time spatial transformation relationship of the target operating area is as follows: Step 501: Convert the real-time physical space coordinates of each non-reflective marker into real-time three-dimensional space coordinates in the real-time local coordinate system.

[0093] Specifically, for each non-reflective marker, assuming its real-time physical space coordinates detected and calculated by the vision system are: (This is a 3D vector referenced to the camera coordinate system.) The origin of the real-time local coordinate system is O' in the camera coordinate system. The unit direction vectors of the X', Y', and Z' axes in the camera coordinate system are respectively... Then, the real-time three-dimensional spatial coordinates of the non-reflective marker in the real-time local coordinate system are... It can be calculated using the following formula: .in,[ [] is a 3×3 rotation matrix. The three columns are the unit direction vectors of the three coordinate axes of the real-time local coordinate system in the camera coordinate system, and T represents the transpose. That is, the origin of the camera coordinate system is first translated to the origin O' of the real-time local coordinate system, and then the translated vector is projected onto the three coordinate axes of the real-time local coordinate system to obtain the three coordinate components of the non-reflective marker in the real-time local coordinate system.

[0094] Performing the above transformation operation on each non-reflective marker within the target operating area yields a set of real-time three-dimensional spatial coordinates in the real-time local coordinate system. The real-time three-dimensional spatial coordinates of each non-reflective marker within the target operating area form a one-to-one correspondence with the pre-stored reference three-dimensional spatial coordinates of each non-reflective marker in the reference local coordinate system. It should be noted that the reference three-dimensional spatial coordinates of each non-reflective marker in the reference local coordinate system are obtained during the teaching phase using the exact same transformation logic as described above, based on the reference physical spatial coordinates of each non-reflective marker. These coordinates are pre-calculated and stored in the system's non-volatile memory for use during runtime.

[0095] Step 502: Based on the real-time three-dimensional spatial coordinates of each non-reflective marker in the real-time local coordinate system and the reference three-dimensional spatial coordinates in the reference local coordinate system, solve for the real-time rigid transformation matrix from the reference local coordinate system to the real-time local coordinate system as the real-time spatial transformation relationship.

[0096] Specifically, after obtaining the real-time and reference 3D spatial coordinates of each non-reflective marker within the target operating area, the rigid transformation matrix between the real-time local coordinate system and the reference local coordinate system can be solved based on these coordinates. The rigid transformation matrix is ​​a 4×4 homogeneous transformation matrix consisting of two parts: a 3×3 rotation matrix R and a 3×1 translation vector T. This matrix can transform any 3D point in the reference local coordinate system... Through formula Precisely mapped to the corresponding point in the real-time local coordinate system .

[0097] Solving for the rotation matrix R and translation vector T essentially involves solving a "point set registration" problem. This means finding the optimal rigid body transformation that maximizes the overlap of the two point sets in the least squares sense, given a known one-to-one correspondence between two point sets (the point set in the reference coordinate system and the point set in the real coordinate system). The specific method is as follows: The first method: an analytical solution based on Singular Value Decomposition (SVD). First, the centroids of the set of marker points in the reference local coordinate system are calculated. centroid of the set of marker points in real-time local coordinate system The centroid is calculated as the arithmetic mean of the coordinates of all marker points. Then, both point sets are decentroided by subtracting the centroid of the corresponding point set from the coordinates of each point, resulting in a decentroided point set. The purpose of decentroiding is to eliminate translation components, simplifying the problem to solving only the rotation matrix. Next, a 3×3 covariance matrix H is constructed, where H is equal to the decentroided reference point set matrix multiplied by the transpose of the decentroided real-time point set matrix. Then, the covariance matrix H is decomposed using SVD to obtain... Where U and V are orthogonal matrices, Σ is a diagonal matrix, and the rotation matrix is... Finally, after calculating the rotation matrix R, the translation vector T can be obtained through... It can be obtained directly.

[0098] It's important to note that after SVD decomposition, it's necessary to check if the determinant of the rotation matrix R is +1. If the determinant is -1, it indicates that the solved transformation includes mirror reflections, which should not occur in actual rigid body transformations. This situation usually arises from collinear or coplanar markers causing ambiguity in the SVD decomposition, or from significant errors in marker detection. In this case, it can be corrected by inverting the last column of V, i.e., letting... This is done to ensure that the determinant of the rotation matrix is ​​+1.

[0099] The second approach is an iterative solution based on the least squares method. When there are many markers (e.g., more than five) and some detection noise exists, a variant of the iterative nearest point approach can be used to further optimize the solution accuracy: first, an initial rotation matrix is ​​obtained using the SVD method. Translation vector Then, with the initial rotation matrix Translation vector Starting with the Levenberg-Marquardt or Gauss-Newton nonlinear optimization algorithms, we use them to minimize the reprojection error between all pairs of marker points—that is, to minimize the objective function. To achieve the objective, the rotation matrix R and translation vector T are solved iteratively. The iterative solution method can achieve higher accuracy than the pure SVD method even in the presence of noise.

[0100] It is important to emphasize that, regardless of whether the SVD or least squares method is used, a fundamental prerequisite for solving the rigid transformation matrix is ​​that at least three non-collinear markers are needed to provide coordinate correspondences. This is because the pose of a rigid body in three-dimensional space has six degrees of freedom—three translational degrees of freedom and three rotational degrees of freedom. Each marker point provides three coordinate constraints (x, y, z), so theoretically, two marker points provide six constraints, which seems sufficient to solve for six degrees of freedom. However, the line connecting two marker points defines an axis, and rotation around this axis cannot be constrained by two points, leading to one degree of freedom being undetermined. Only three non-collinear marker points can provide sufficient geometric constraints to uniquely determine the rotation matrix and translation vector. If the markers happen to be collinear, the same problem arises where the degree of freedom for rotation around this line cannot be determined. Therefore, at least three non-collinear, non-reflective markers should be placed in each operating region to ensure the geometric uniqueness and numerical stability of the rigid transformation matrix solution.

[0101] Specifically, based on the pre-calibrated reference pose state data of the target manipulator in the reference local coordinate system, the real-time pose state data of the target manipulator in the real-time local coordinate system is inferred by using the real-time spatial transformation relationship. This includes mapping the reference pose state data of the target manipulator in the reference local coordinate system to the real-time local coordinate system based on the three-dimensional rotation matrix and three-dimensional translation vector in the real-time spatial transformation relationship, thereby obtaining the real-time pose state data of the target manipulator in the real-time local coordinate system.

[0102] The "three-dimensional rotation matrix" mentioned here refers to the 3×3 matrix R obtained through SVD decomposition or least squares method, which describes the rotational transformation relationship between the reference local coordinate system and the real-time local coordinate system. The "three-dimensional translation vector" refers to the 3×1 vector T obtained through SVD decomposition or least squares method, which describes the translational offset of the origin of the reference local coordinate system relative to the origin of the real-time local coordinate system. R and T together constitute the complete rigid transformation matrix.

[0103] The mapping formula for mapping the reference pose state data of the target manipulator in the reference local coordinate system to the real-time local coordinate system is: .in, It is a 3×1 column vector representing the reference pose state data of the target operation part in the reference local coordinate system - such as the three-dimensional coordinates of the center point of the end of the button when it is in the "pressed" state, or the three-dimensional coordinates of the center point of the marker of the knob when it is in "gear 1". This is a 3×1 column vector obtained after mapping, representing the real-time pose state data of the operator in the real-time local coordinate system. The physical meaning of this mapping formula is: first, the points in the reference coordinate system are mapped... The system is rotated according to the rotation relationship between the reference coordinate system and the real-time coordinate system to align its direction with the real-time coordinate system. Then, it is translated according to the translation relationship to align its position with the origin of the real-time coordinate system. Through this combined transformation of rotation and translation, the operation point originally defined in the reference spatial frame is accurately migrated to the real-time spatial frame at the current moment.

[0104] To illustrate this mapping process more intuitively, a specific numerical example is given below. Suppose that during the teaching phase, for a certain button, the system records the three-dimensional coordinates of its end-point center point in the reference local coordinate system when it is in the "pressed" state. The unit is millimeters. This coordinate means that in the reference local coordinate system, the target point when the button is pressed is located at 10 mm in the X-axis direction, 20 mm in the Y-axis direction, and 0 mm in the Z-axis direction. During the runtime phase, the rotation matrix R and translation vector T, obtained through SVD decomposition or least squares method, are as follows: R is a 3×3 identity matrix (indicating that there is no relative rotation between the reference coordinate system and the real-time coordinate system, only translation), and T = (2, -2, 1), with units in millimeters. Therefore, the mapping formula is used to calculate... The result (12,18,1) is the real-time pose data of the button in the real-time local coordinate system. It means that at the current moment, the robotic arm needs to move the end effector to a position of X=12mm, Y=18mm, Z=1mm in the real-time local coordinate system in order to accurately press the button.

[0105] Of course, using the identity matrix R in the above example is a simplification. In actual operation, due to changes in camera viewpoint or overall device posture, there is usually both rotation and translation between the reference local coordinate system and the real-time local coordinate system. For example, assuming the device panel is tilted by 30 degrees, the solved rotation matrix R will contain the corresponding trigonometric function values, and the translation vector T will reflect the spatial displacement of the panel. Regardless of the specific values ​​of R and T, the mapping formula... It is always applicable and can accurately compensate for spatial offsets caused by arbitrary rigid body transformations.

[0106] It should be noted that the reference pose state data is not limited to the three-dimensional spatial coordinates of the actuator. For different types of actuators, the recorded reference pose state data may take different forms. For buttons and toggle switches, the reference pose state data is typically represented by the three-dimensional coordinates of key feature points of the actuator, such as the coordinates of the center point of the button end or the coordinates of the end point of the switch handle. For knob-type actuators, in addition to the three-dimensional coordinates of the knob's center point, the reference pose state data may also include the knob's attitude parameters, such as the knob's orientation represented by quaternions or rotation matrices. When the reference pose state data contains attitude information, the mapping process needs to process the position and attitude separately. The position part is still processed through... Mapping is then performed. The attitude part requires mapping through a composite transformation of the rotation matrix: if the reference attitude uses a rotation matrix... The real-time attitude is indicated by (describing the orientation of the manipulator in the reference local coordinate system). In this way, the complete pose of the manipulator in the real-time local coordinate system—including position and orientation—can be derived from the reference data through rigid body transformation.

[0107] By employing a rigid body transformation-based mapping mechanism, the real-time pose state data of the target actuator in the real-time local coordinate system is inferred. This allows all actuator pose knowledge accumulated during the teaching phase—whether it's the pressed coordinates of a button or the multiple gear positions of a knob—to be batch-mapped into the real-time space using the same transformation matrices R and T, without requiring re-identification or re-positioning of each actuator individually. In other words, during the runtime phase, only the correspondence between the reference local coordinate system of the teaching phase and the real-time local coordinate system of the runtime phase needs to be solved. This allows the reference pose state data of each key operation of the target actuator in the reference local coordinate system during the teaching phase to be mapped all at once to the real-time pose state data of each key operation of the target actuator in the real-time local coordinate system, without needing to re-explore each one individually. This batch mapping mechanism can quickly recover the real-time pose of all actuators within the entire operating area under single-frame or low-frequency visual input conditions.

[0108] The above describes how, by placing non-reflective markers on a rigid structure (such as a panel) within the target operating area, the real-time pose state data of the target operating component is indirectly inferred by solving the rigid transformation matrix between the reference local coordinate system and the real-time local coordinate system of the target operating area. The following explains the method for inferring the real-time pose state data of the target operating component when non-reflective markers are placed on it. If non-reflective markers are placed on the target operating component, the visual image contains at least one non-reflective marker placed on the target operating component. In this scenario, the inference of the real-time pose state data of the target operating component no longer relies on the construction and spatial transformation of the local region coordinate system, but is directly accomplished by detecting and comparing the pose of the non-reflective markers carried by the target operating component itself.

[0109] This method is suitable for actuators that have sufficient space for installation and undergo significant posture changes during operation, such as knobs. Knobs typically have a cylindrical or prismatic head for gripping, with an end or side surface area large enough to attach or embed a non-reflective marker. Of course, it's not limited to knobs; it can also be applied to some larger toggle switches or pull-type actuators, as long as at least one non-reflective marker can be securely mounted on its body.

[0110] During the teaching phase, the system controls the actuator at the end of the teaching robot arm to perform different key operations on a manipulator with a non-reflective marker. The system records the pose data of the non-reflective marker under each key operation as reference pose data. This "pose data" includes both position and orientation information. Position information typically refers to the three-dimensional coordinates of the center point of the non-reflective marker in a reference coordinate system (which could be the camera coordinate system or a reference local coordinate system for the area where the manipulator is located). Orientation information refers to the orientation of the marker in three-dimensional space, which can be represented by rotation matrices, quaternions, or Euler angles. For example, for a three-position knob, during teaching, the knob is rotated sequentially to position 1, position 2, and position 3. At each position, the position coordinates and orientation quaternions of the non-reflective marker on the knob are recorded, and the correspondence between the position number and the pose data is stored in the teaching database.

[0111] During the operation phase, after acquiring visual images of the target operating area where the target operating component is located through the vision system, refer to... Figure 6 As shown, the real-time pose state data of the target manipulator can be inferred in the following ways: Step 601: Perform marker detection and coordinate system transformation on the visual image to obtain the real-time pose data of each non-reflective marker.

[0112] After the vision system acquires a visual image of the target operation area where the target operation device is located, it uses a marker detection algorithm to locate the non-reflective markers placed on the target operation device in the visual image. The non-reflective markers on the target operation device undergo spatial pose changes as the target operation device moves. Therefore, the marker detection algorithm not only needs to output the two-dimensional pixel coordinates (center point pixel coordinates) of the non-reflective markers, but also needs to decode to obtain the identifiers of the non-reflective markers. Combined with depth information or structured light calculation results, the two-dimensional pixel coordinates are converted into three-dimensional physical space coordinates, and the pose parameters of the non-reflective markers in the current camera coordinate system are calculated. The pose parameters can be calculated based on the geometric relationships of multiple feature points on the non-reflective markers. For example, if the non-reflective markers use a square checkerboard pattern, the rotation matrix and translation vector of the non-reflective markers relative to the camera can be solved using the PnP algorithm by detecting the four corner points of the checkerboard. If a circular dot matrix coded pattern is used, the spatial pose of the non-reflective markers can be recovered through ellipse fitting and coded point matching.

[0113] Step 602: Based on the real-time pose state data of each non-reflective marker and the pre-calibrated reference pose state data of each non-reflective marker under different key operations, infer the real-time pose state data of the target manipulator under different key operations.

[0114] After obtaining the real-time pose data of the non-reflective marker on the target manipulator, the system compares it with the reference pose data of the non-reflective marker under different key operations pre-calibrated and stored during the teaching phase, thereby inferring the current key operation state of the target manipulator. Specifically, the system calculates the distance between the real-time pose data and each reference pose data, and then selects the reference pose data with the smallest distance as the inference result. The "distance" mentioned here is a comprehensive metric that includes both positional and orientational differences. Positional differences can be measured by the Euclidean distance between two 3D coordinate points; orientational differences can be measured by the angle between two orientation quaternions or the geodesic distance between two rotation matrices. The final comprehensive distance can be a weighted sum of the positional and orientational distances, with the weights determined according to the importance of position and orientation to state judgment in the actual application scenario.

[0115] Let's take a specific knob position inference as an example to illustrate this comparison process. Assume a knob has three positions. The reference pose data recorded during the teaching phase are as follows: Position 1 corresponds to a 0° rotation angle of the marker around the knob's Z-axis, with the marker's center point coordinates being (15, 25, 5); Position 2 corresponds to a 90° rotation angle, with coordinates (15, 25, 5); Position 3 corresponds to a 180° rotation angle, with coordinates (15, 25, 5). During the operation phase, the system detects a real-time rotation angle of 92° for the non-reflective marker on the knob, with real-time coordinates (16, 24, 6). The system calculates the combined distance between the real-time data and the reference data for each of the three positions. Since the reference coordinates of the three positions are very close, the positional distance contributes little to distinguishing the positions, while the attitude distance becomes the decisive factor: the real-time angle of 92° differs from 90° for position 2 by only 2°, from 0° for position 1 by 92°, and from 180° for position 3 by 88°. Therefore, the system infers that the knob is currently in position 2.

[0116] It should be noted that in the method of inferring the real-time pose state data of a target manipulator by setting non-reflective markers on a rigid structure in the target operating space, the inference of the real-time pose state data of the target manipulator relies on two prerequisites: first, non-reflective markers are set on the rigid structure of the target operating area to construct a real-time local coordinate system; second, it is necessary to solve the rigid transformation matrix from the reference local coordinate system to the real-time local coordinate system, and then map the reference pose state data to the real-time space. The core advantage of using non-reflective markers on the target manipulator for real-time pose state data inference is that the system does not need to perform the construction of a local coordinate system or spatial transformation. It only needs to detect the non-reflective markers carried by the target manipulator itself in the visual image, obtain its real-time pose state data, and then directly compare it with the reference pose state data of the non-reflective markers recorded during the teaching phase to complete the inference of the real-time pose state data of the target manipulator. The entire process does not involve region division, local coordinate system construction, or rigid transformation matrix solution, and is therefore unaffected by changes in camera viewpoint or overall device attitude. This is because the non-reflective marker on the target operating component is rigidly connected to the operating component body, and the two move together. The pose of the non-reflective marker directly reflects the pose of the target operating component, without the need for any intermediate coordinate system transformation.

[0117] This independence brings two significant technical benefits: First, even when non-reflective markers placed on rigid structures within the target operating area are obscured or damaged, the system can still infer the state of target operating components bearing their own markers, thereby improving the robustness and fault tolerance of the entire system. Second, for rotary operating components such as knobs, the accuracy of state inference directly depends on the accuracy of marker attitude calculation. Without the intermediate step of regional coordinate system transformation, the accumulation of transformation matrix solution errors is avoided, thus typically achieving higher gear position determination accuracy.

[0118] It should be understood that for a device panel containing multiple operating components, a scheme using panel markers plus regional coordinate system transformation can be used for buttons and toggle switches, while a scheme using direct comparison of body markers can be used for knobs. Both schemes can be executed in parallel within the same visual image without interference. This hybrid deployment method can fully leverage the advantages of both schemes, ensuring the overall system simplicity while providing higher state inference accuracy for specific operating components.

[0119] The above describes how to infer the real-time pose state data of the target manipulator in the real-time local coordinate system through regional coordinate system transformation and comparison with manipulator body markers. Based on this, it further describes how to directly use these inferred results to guide the robotic arm in performing actual operations, thus forming a complete closed loop from visual perception to state inference to motion execution.

[0120] Specifically, if the real-time pose state data of the target manipulator in the real-time local coordinate system includes the real-time pose state data of the target manipulator under a critical operation, then the critical operation is performed on the target manipulator based on the real-time pose state data of the target manipulator under the critical operation.

[0121] The phrase "real-time pose state data of the target manipulator under a key operation" refers to a task where the target manipulator only needs to perform a single, independent action. For example, the task might simply require pressing a button to change it from a popped-up state to a pressed state, or toggling a switch from the "off" to the "on" position. In this case, the system directly reads the real-time pose state data corresponding to the key operation from the inferred real-time pose state data of the target manipulator. For a button, this is typically a three-dimensional coordinate point representing the target position that the robotic arm's end effector needs to reach; for a knob, in addition to the three-dimensional coordinate point, it may also include an attitude parameter representing the target orientation that the robotic arm's end effector gripper needs to reach. After receiving this data, the robotic arm control unit converts it into motion commands for each joint of the robotic arm, controlling the end effector to move from its current position to the target pose, and then performing actions such as pressing, toggling, or rotating.

[0122] If the real-time pose state data of the target manipulator in the real-time local coordinate system includes the real-time pose state data of the target manipulator under at least two key operations, then according to the preset key operation execution order, based on the real-time pose state data of the target manipulator under each key operation, each key operation is executed sequentially on the target manipulator.

[0123] In actual industrial operations, it is more common for the target actuator to perform multiple consecutive critical operations sequentially. For example, a typical equipment inspection procedure might involve first rotating a knob from its current position to a specified position, then pressing a confirmation button, and finally tossing a switch to the "on" position. In this case, the real-time pose status data of the target actuator inferred by the system contains multiple entries, each corresponding to a critical operation.

[0124] It should be noted that the "preset execution order of key operations" is not determined temporarily during the runtime phase, but rather established and recorded during the teaching phase. During the teaching phase, the operator controls the teaching robotic arm to execute each key operation sequentially according to the correct operating procedure. The system records the reference pose state data corresponding to each key operation in this order and saves the sequence information. During the runtime phase, strictly following this preset order, the system sequentially reads the pose state data of the current key operation from the inferred real-time pose state dataset, controls the robotic arm to adjust the target manipulator to the pose state data of the current key operation, and reads and executes the pose state data of the next key operation only after the current key operation is completed, until all key operations are completed.

[0125] Maintaining consistency between the execution sequence and the teaching sequence is of paramount technical importance. In low-pressure, enclosed equipment, there are often physical interlocks or logical dependencies between operating components. For example, some equipment requires that a knob be rotated to unlock before a button can be pressed; or that a button must be pressed to activate a circuit before a switch can be turned on. If the robotic arm arbitrarily disrupts the operating sequence during operation, it will lead to ineffective operations and may even damage the internal mechanical structure or electrical circuits of the equipment. Therefore, by forcing the execution sequence during the operation phase to be completely consistent with the teaching phase, the logical correctness of the operating procedure and the safety of the equipment are fundamentally guaranteed. This "teaching as procedure" design philosophy allows operators to correctly demonstrate the operating procedure once during the teaching phase, and the system can reproduce this procedure in every subsequent automatic operation without additional procedure programming or logic configuration.

[0126] Specifically, a critical operation queue can be used to manage the execution of multiple critical operations. The preset critical operation execution order recorded during the teaching phase is transformed into an ordered critical operation queue. Each critical operation in the queue contains an operator identifier, a critical operation type, and a reference pose state data index. During the execution phase, after inferring the real-time pose state data of the target operator under each critical operation, the real-time pose state data of all critical operations are added to the critical operation queue according to the preset critical operation execution order. Then, starting from the first critical operation in the critical operation queue, the real-time pose state data of the target operator under that critical operation is read, and the robotic arm is controlled to execute the first critical operation on the target operator according to that real-time pose state data. After execution, the critical operation is removed from the critical operation queue, and the next critical operation in the queue is processed, until the critical operation queue is empty. This queued management method makes the execution logic of multiple critical operations clear and controllable, and facilitates feedback verification based on the reference pose state data index after the execution of critical operations.

[0127] In real-world industrial environments, due to minute positioning errors in robotic arm movements, potential mechanical clearances or wear on the manipulated parts, and sub-pixel-level detection biases in vision systems under low-frequency conditions, a slight deviation may exist between the actual pose of the target manipulated part and the reference pose after the robotic arm executes key operations based on the inferred real-time pose data. If these deviations are not detected and compensated for, in multi-key operations, a small deviation in the previous step may be amplified in subsequent steps, ultimately leading to the failure of the entire operation. Therefore, during the sequential execution of key operations according to a preset execution order, based on the real-time pose data of the target manipulated part under each key operation, a feedback verification and adjustment mechanism can be used to reduce deviations after each key operation and before the next. Specifically, this can be achieved through, but is not limited to, the following methods: First, the actual pose state data of the target manipulator after the current key operation is performed is compared with the reference pose state data.

[0128] The actual pose state data refers to the actual pose state of the target manipulator in the physical world after the robotic arm completes the current key operation. This data is not read from the previously inferred real-time pose state data, but is obtained by re-acquiring a visual image frame at the current moment through the vision system, and then going through the same marker detection, coordinate system transformation and pose inference process as described above to re-acquire the current actual pose state data of the target manipulator.

[0129] Reference pose state data refers to the standard pose state data corresponding to the current critical operation recorded during the teaching phase. It represents the pose state that the current critical operation should ideally achieve, and is the "bullseye" that the robotic arm aims at when performing the operation. Specifically, during the teaching phase, based on the spatial pose transformation relationship between each non-reflective marker and the target manipulator, the reference pose state data of each non-reflective marker in the teaching phase can be obtained by rigidly transforming it.

[0130] The specific comparison method depends on the type of the target actuator and the representation of its pose state data. For button-type actuators, their pose state data is mainly represented by three-dimensional spatial coordinates. During comparison, the Euclidean distance difference between the actual coordinates and the target coordinates is calculated. For example, assuming the center point coordinates of the button in the reference pose state data are (12, 18, 1), while the actual coordinates inferred from the re-acquired image after the operation are (12.3, 18.1, 0.8), then the Euclidean distance difference is... For knob-type actuators, their pose data includes not only position coordinates but also attitude parameters. During comparison, both the position distance difference and the attitude angle difference need to be calculated simultaneously. The attitude angle difference can be obtained by calculating the angle between the actual attitude quaternion and the target attitude quaternion, or by comparing the geodesic distance between the actual rotation matrix and the target rotation matrix. For toggle switch-type actuators, in addition to the position distance difference, the angle between the actual direction vector and the target direction vector can also be calculated as a measure of attitude deviation.

[0131] Secondly, if the comparison result is determined to be satisfactory, the next key operation is executed; if the comparison result is determined to be unsatisfactory, the real-time pose state data of the target manipulator under the current key operation is adjusted based on the visual image until the comparison result between the actual pose state data of the target manipulator and the reference pose state data after the current key operation is executed is satisfactory, and then the next key operation is executed.

[0132] In this context, "meeting the standard" means that the calculated deviation is less than or equal to a preset threshold. The threshold setting needs to be determined comprehensively based on the type of manipulator, the required operational precision, and the positioning accuracy of the robotic arm. For button press operations, the position threshold can be set between 0.5 mm and 1 mm. For example, if the threshold is set to 1 mm, the calculated Euclidean distance difference in the above example is 0.37 mm, which is less than 1 mm, thus the comparison result is considered satisfactory, and the system can safely proceed to the next critical operation. For knob rotation operations, the angle threshold can be set between 1 degree and 2 degrees. For toggle switch operation, the position and angle thresholds can be set separately, for example, a position threshold of 1 mm and an angle threshold of 2 degrees.

[0133] When the comparison result fails to meet the standard, it means that after the robotic arm performs the current critical operation, the target manipulator has not reached the expected target state. There could be several reasons for this, such as: slight slippage between the robotic arm's end effector and the manipulator; unexpected mechanical resistance or jamming in the manipulator itself; sub-pixel-level random errors in the vision system's detection of markers under low-frequency conditions; or slight numerical errors in solving the rigid transformation matrix between the reference local coordinate system and the real-time local coordinate system. These deviations can accumulate and ultimately lead to the failure of subsequent operations.

[0134] At this point, the system enters the adjustment and compensation phase: using the actual pose state data re-inferred from the current frame's visual image, it calculates the deviation vector between the actual pose state data and the reference pose state data. This deviation vector is then used to compensate for the deviation in the robotic arm's motion commands, controlling the robotic arm to perform fine-tuning movements to bring the target manipulator closer to the reference pose state data. After fine-tuning, the system again acquires the poem image through the vision system, re-infers the actual pose state data, and compares it with the reference pose state data. If the comparison result still does not meet the standard, the deviation is recalculated, and fine-tuning compensation is performed again. This process is iterated until the comparison result meets the standard.

[0135] Let's take a specific example of button press deviation compensation to illustrate this iterative process. Assume the coordinates of the center point of the button in the reference pose data when it's pressed are... The actual coordinates inferred from the re-acquired image after the press operation are... Calculate the deviation vector. This deviation vector indicates that the button is 0.8 mm short of being fully pressed in the X-axis direction and 0.5 mm deeper than the target in the Z-axis direction. Based on this, the system generates a fine-tuning command, controlling the robotic arm's end effector to move 0.8 mm along the positive X-axis and simultaneously retract 0.5 mm along the negative Z-axis. After fine-tuning, the system acquires images again and re-infers the actual coordinates, assuming that the current result is... The Euclidean distance difference between the system and the target coordinates is approximately 0.22 mm, which is less than the preset threshold of 1 mm. The comparison result meets the standard, the system stops fine-tuning, and proceeds to the next critical operation.

[0136] It's important to note that to prevent the fine-tuning process from getting stuck in an infinite loop—for example, due to the robotic arm's positioning accuracy limits or physical jamming of the manipulator, preventing it from ever meeting the target—the system typically sets a maximum number of fine-tuning iterations, such as 3 or 5. If the comparison result still doesn't meet the target after reaching the maximum number of iterations, the system will pause the operation and issue an alarm, prompting the operator to intervene and check. This protection mechanism ensures that the system will not endlessly attempt fine-tuning under abnormal circumstances, thus avoiding damage to the equipment or manipulator.

[0137] By introducing a closed-loop mechanism of "execution-comparison-adjustment," the reliability of multi-critical operations under low-frequency vision conditions is significantly improved. Deviations in each step are detected and corrected before proceeding to the next step, preventing deviations from accumulating across operations. Furthermore, it reduces the reliance on the absolute positioning accuracy of the robotic arm. Even with limited repeatability of the robotic arm itself, the closed-loop compensation through visual feedback allows the system to achieve operational accuracy exceeding that of the robotic arm itself. In addition, the system is endowed with adaptive capabilities to changes in the state of the manipulator. If the manipulator experiences slight physical displacement or wear due to long-term use, causing a deviation between its actual position and the position recorded during the teaching phase, the closed-loop compensation mechanism can automatically detect and correct this deviation without requiring re-teaching.

[0138] Based on the above embodiments, this application provides an operating element state inference system, see below. Figure 7 As shown, the operation element state inference system 700 provided in this application embodiment includes at least: Image acquisition unit 701 is used to acquire visual images of the target operating area where the target operating component is located through a vision system; wherein, the visual image contains at least two non-reflective markers set in the target operating area; The real-time construction unit 702 is used to perform marker detection and coordinate system transformation on the visual image to obtain the real-time physical space coordinates of each non-reflective marker, and to construct the real-time local coordinate system of the target operation area based on the real-time physical space coordinates of each non-reflective marker. The spatial transformation unit 703 is used to determine the real-time spatial transformation relationship of the target operating area based on the real-time local coordinate system and the reference local coordinate system of the target operating area; wherein, the reference local coordinate system is a local coordinate system of the target operating area that is pre-constructed based on the reference physical spatial coordinates of each non-reflective marker. The state inference unit 704 is used to infer the real-time pose state data of the target manipulator in the real-time local coordinate system based on the pre-calibrated reference pose state data of the target manipulator in the reference local coordinate system and by adopting the real-time spatial transformation relationship.

[0139] In one possible implementation, the operating element state inference system 700 provided in this application embodiment further includes: The reference construction unit 705 is used to acquire a reference visual image of each operating area within a low-pressure confined space using a vision system. The reference visual image contains at least two non-reflective markers placed on the rigid structure of the target operating area. The reference visual image is used for marker detection and coordinate system transformation to obtain the reference physical space coordinates of each non-reflective marker. Based on the reference physical space coordinates of each non-reflective marker, the origin spatial coordinates are determined. Based on the preset coding order of the two non-reflective markers, the X-axis direction is determined. Based on the normal vector of the plane containing each non-reflective marker, the Z-axis direction is determined. Based on the X-axis and Z-axis directions, the right-hand rule is used to determine the Y-axis direction. Using the origin spatial coordinates as the origin, a reference local coordinate system for the operating area is constructed based on the X-axis, Y-axis, and Z-axis directions.

[0140] In one possible implementation, the operating element state inference system 200 provided in this application embodiment further includes: The teaching calibration unit 706 is used to control the actuator at the end of the teaching robot arm to perform different key operations on each manipulator in different operating areas, and to record the relative pose state data of the manipulator in the reference local coordinate system under different key operations as reference pose state data according to the execution sequence of the key operations; to perform consistency verification on the reference pose state data of each manipulator in different operating areas; and to save the reference pose state data of each manipulator in different operating areas after confirming that the consistency verification is passed.

[0141] In one possible implementation, the real-time construction unit 702 is used to determine the real-time origin spatial coordinates based on the real-time physical spatial coordinates of each non-reflective marker; determine the real-time X-axis direction based on the preset coding order of two non-reflective markers among the non-reflective markers; determine the real-time Z-axis direction based on the normal vector of the plane where each non-reflective marker is located; determine the real-time Y-axis direction based on the real-time X-axis direction and the real-time Z-axis direction using the right-hand rule; and construct a real-time local coordinate system of the operation area with the real-time origin spatial coordinates as the origin, based on the real-time X-axis direction, the real-time Y-axis direction, and the real-time Z-axis direction.

[0142] In one possible implementation, the spatial transformation unit 703 is used to convert the real-time physical spatial coordinates of each non-reflective marker into real-time three-dimensional spatial coordinates in a real-time local coordinate system; based on the real-time three-dimensional spatial coordinates of each non-reflective marker in the real-time local coordinate system and the reference three-dimensional spatial coordinates in the reference local coordinate system, the real-time rigid transformation matrix from the reference local coordinate system to the real-time local coordinate system is solved as the real-time spatial transformation relationship; wherein, the reference three-dimensional spatial coordinates of each non-reflective marker in the reference local coordinate system are obtained in advance by performing coordinate system transformation based on the reference physical spatial coordinates of each non-reflective marker.

[0143] In one possible implementation, the state inference unit 704 is used to map the reference pose state data of the target manipulator in the reference local coordinate system to the real-time local coordinate system based on the three-dimensional rotation matrix and three-dimensional translation vector in the real-time spatial transformation relationship, so as to obtain the real-time pose state data of the target manipulator in the real-time local coordinate system.

[0144] In one possible implementation, if a non-reflective marker is disposed on the target operating component, the visual image includes at least one non-reflective marker disposed on the target operating component. The state inference unit 704 is also used to perform marker detection and coordinate system transformation on the visual image to obtain the real-time pose state data of each non-reflective marker; based on the real-time pose state data of each non-reflective marker and the reference pose state data of each non-reflective marker under different key operations, it infers the real-time pose state data of the target operation under different key operations.

[0145] In one possible implementation, the operating element state inference system 700 provided in this application embodiment further includes: The operation control unit 707 is configured to, if the real-time pose state data of the target operator in the real-time local coordinate system includes the real-time pose state data of the target operator under a key operation, then perform a key operation on the target operator based on the real-time pose state data of the target operator under a key operation; if the real-time pose state data of the target operator in the real-time local coordinate system includes the real-time pose state data of the target operator under at least two key operations, then execute each key operation sequentially on the target operator based on the real-time pose state data of the target operator under each key operation according to a preset key operation execution order.

[0146] In one possible implementation, the operation control unit 707 is used to compare the actual pose state data of the target manipulator after the execution of the current key operation with the reference pose state data; if the comparison result is determined to be satisfactory, the next key operation is executed; if the comparison result is determined to be unsatisfactory, the real-time pose state data of the target manipulator under the current key operation is adjusted based on the visual image until the comparison result of the actual pose state data of the target manipulator and the reference pose state data after the execution of the current key operation is satisfactory, and then the next key operation is executed.

[0147] It should be noted that the principle of the operation device state inference system 700 provided in this application embodiment to solve the technical problem is similar to the operation device state inference method provided in this application embodiment. Therefore, the implementation of the operation device state inference system 700 provided in this application embodiment can refer to the implementation of the operation device state inference method provided in this application embodiment, and the repeated parts will not be described again.

[0148] After introducing the operation state inference method and system provided in the embodiments of this application, the robot provided in the embodiments of this application will be briefly introduced next.

[0149] The robot provided in this application includes at least a robot body, a robot control device disposed inside the robot body, and a robotic arm connected to the robot body; wherein, the end of the robotic arm is equipped with a vision system and an execution structure, both of which are connected to the robot control device. The vision system provides visual images to the robot control device, and the robot control device infers the real-time pose state data of the target manipulator based on the visual images collected by the vision system, and controls the execution mechanism to perform remote or automated operations on the target manipulator based on the real-time pose state data.

[0150] In the embodiments of this application, see the following: Figure 8 As shown, the robot control device 800 includes at least a processor 801, a memory 802, and a computer program stored in the memory 802 and executable on the processor 801. When the processor 801 executes the computer program, it implements the above-described operation state inference method provided in the embodiments of this application.

[0151] In one possible implementation, processor 801 can be a single processing element or a collective term for multiple processing elements. For example, processor 801 can be a central processing unit (CPU), or one or more integrated circuits configured to implement the operation state inference method described in the embodiments of this application. Specifically, processor 801 can be a general-purpose processor, including but not limited to CPUs, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0152] In one possible implementation, memory 802 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 8021 and / or cache memory 8022, and may further include read-only memory (ROM) 8023; memory 802 may also include a program tool 8025 having a set (at least one) of program modules 8024, including but not limited to: operating subsystem, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.

[0153] In one possible implementation, the robot control device 800 may further include a bus 803 connecting different components (including processor 801 and memory 802). The bus 803 represents one or more types of bus structures, including memory bus, peripheral bus, local area bus, etc.

[0154] In one possible implementation, the robot control device 800 can also communicate with one or more devices that enable users to interact with the robot (e.g., remote controls, mobile phones, computers, etc.), and / or with external devices 804 that enable the robot to communicate with one or more other robots (e.g., routers, modems, etc.). This communication can be performed via an input / output (I / O) interface 805. Furthermore, the robot control device 800 can also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 806. Figure 8As shown, network adapter 806 communicates with other modules via bus 803. It should be understood that, although... Figure 8 As not shown, other hardware and / or software modules can be used in conjunction with the robot control device 800, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, Redundant Arrays of Independent Disks (RAID) subsystems, tape drives, and data backup storage subsystems.

[0155] It should be noted that, Figure 8 The robot control device 800 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0156] Furthermore, embodiments of this application also provide a computer-readable storage medium storing computer instructions. When these computer instructions are executed by a processor, they implement the operation state inference method described above in embodiments of this application. Specifically, the computer instructions may be built into or installed in a processor, enabling the processor to implement the operation state inference method described above in embodiments of this application by executing the built-in or installed computer instructions.

[0157] Of course, the above-mentioned operation state inference method provided in the embodiments of this application can also be implemented as a program product, which includes program code. When the program code is executed by the processor, it implements the above-mentioned operation state inference method provided in the embodiments of this application.

[0158] The program product provided in this application embodiment can be any combination of one or more readable media, wherein the readable media can be a readable signal medium or a readable storage medium, and the readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. Specifically, more specific examples of readable storage media (a non-exhaustive list) include: electrical connections with one or more wires, portable disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0159] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0160] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0161] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0162] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of the embodiments of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.

Claims

1. A method for inferring the state of an operating component, characterized in that, include: A visual image of the target operating area where the target operating component is located is acquired through a vision system; wherein the visual image contains at least two non-reflective markers placed within the target operating area; The visual image is subjected to marker detection and coordinate system transformation to obtain the real-time physical space coordinates of each non-reflective marker, and the real-time local coordinate system of the target operation area is constructed based on the real-time physical space coordinates of each non-reflective marker. Based on the real-time local coordinate system and the reference local coordinate system of the target operating area, the real-time spatial transformation relationship of the target operating area is determined; wherein, the reference local coordinate system is a local coordinate system of the target operating area that is pre-constructed based on the reference physical space coordinates of each non-reflective marker. Based on the pre-calibrated reference pose state data of the target manipulator in the reference local coordinate system, the real-time pose state data of the target manipulator in the real-time local coordinate system is inferred using the real-time spatial transformation relationship.

2. The method for inferring the state of an operating component as described in claim 1, characterized in that, Before acquiring a visual image of the target operating area where the target operating component is located through the vision system, the following steps are also included: For each operating area within a low-pressure confined space, a reference visual image of the operating area is acquired using a vision system; wherein, the reference visual image contains at least two non-reflective markers placed within the operating area; The reference physical space coordinates of each non-reflective marker are obtained by performing marker detection and coordinate system transformation on the reference visual image. Based on the reference physical space coordinates of each of the non-reflective markers, determine the origin space coordinates; The X-axis direction is determined based on the preset coding order of two non-reflective markers among the aforementioned non-reflective markers; The Z-axis direction is determined based on the normal vector of the plane containing each non-reflective marker; The Y-axis direction is determined using the right-hand rule based on the X-axis direction and the Z-axis direction. Using the origin spatial coordinates as the origin, a reference local coordinate system for the operating area is constructed based on the X-axis direction, the Y-axis direction, and the Z-axis direction.

3. The method for inferring the state of an operating component as described in claim 1, characterized in that, Before acquiring a visual image of the target operating area where the target operating component is located through the vision system, the following steps are also included: For each manipulator in different operating areas, the actuator at the end of the teaching robot arm is controlled to perform different key operations on the manipulator. According to the execution order of the key operations, the relative pose state data of the manipulator in the reference local coordinate system under different key operations is recorded as the reference pose state data. Consistency verification is performed on the reference pose state data of each operand in different operating areas; After confirming that the consistency check has passed, save the reference pose state data of each operand in different operating areas.

4. The method for inferring the state of an operating component as described in claim 1, characterized in that, Based on the real-time physical space coordinates of each of the non-reflective markers, a real-time local coordinate system for the target operating area is constructed, including: Based on the real-time physical space coordinates of each of the non-reflective markers, determine the real-time origin space coordinates; The real-time X-axis direction is determined based on the preset coding order of two non-reflective markers among the various non-reflective markers; The real-time Z-axis direction is determined based on the normal vector of the plane containing each non-reflective marker; Based on the real-time X-axis direction and the real-time Z-axis direction, the real-time Y-axis direction is determined using the right-hand rule; Using the real-time origin spatial coordinates as the origin, and based on the real-time X-axis direction, the real-time Y-axis direction, and the real-time Z-axis direction, a real-time local coordinate system for the operating area is constructed.

5. The method for inferring the state of an operating component as described in claim 1, characterized in that, Based on the real-time local coordinate system and the reference local coordinate system of the target operating region, the real-time spatial transformation relationship of the target operating region is determined, including: The real-time physical space coordinates of each non-reflective marker are converted into real-time three-dimensional space coordinates in the real-time local coordinate system. Based on the real-time three-dimensional spatial coordinates of each non-reflective marker in the real-time local coordinate system and the reference three-dimensional spatial coordinates in the reference local coordinate system, the real-time rigid transformation matrix from the reference local coordinate system to the real-time local coordinate system is solved as the real-time spatial transformation relationship; wherein, the reference three-dimensional spatial coordinates of each non-reflective marker in the reference local coordinate system are obtained by performing coordinate system transformation in advance based on the reference physical spatial coordinates of each non-reflective marker.

6. The method for inferring the state of an operating component as described in claim 1, characterized in that, Based on the pre-calibrated reference pose state data of the target manipulator in the reference local coordinate system, and using the real-time spatial transformation relationship, the real-time pose state data of the target manipulator in the real-time local coordinate system is inferred, including: Based on the three-dimensional rotation matrix and three-dimensional translation vector in the real-time spatial transformation relationship, the reference pose state data of the target manipulator in the reference local coordinate system is mapped to the real-time local coordinate system to obtain the real-time pose state data of the target manipulator in the real-time local coordinate system.

7. The method for inferring the state of an operating component as described in claim 1, characterized in that, If the non-reflective marker is disposed on the target operating component, then the visual image includes at least one non-reflective marker disposed on the target operating component; After acquiring a visual image of the target operating area where the target operating component is located through the vision system, the following is also included: The visual image is subjected to marker detection and coordinate system transformation to obtain the real-time pose data of each non-reflective marker. Based on the real-time pose state data of each non-reflective marker and the reference pose state data of each non-reflective marker under different pre-calibrated key operations, the real-time pose state data of the target operator under different key operations is inferred.

8. The method for inferring the state of an operating component as described in any one of claims 1-7, characterized in that, Further includes: If the real-time pose state data of the target operator in the real-time local coordinate system includes the real-time pose state data of the target operator under a key operation, then the key operation is performed on the target operator based on the real-time pose state data of the target operator under the key operation. If the real-time pose state data of the target operator in the real-time local coordinate system includes the real-time pose state data of the target operator under at least two key operations, then according to the preset key operation execution order, based on the real-time pose state data of the target operator under each key operation, each key operation is executed sequentially on the target operator.

9. The method for inferring the state of an operating component as described in claim 8, characterized in that, According to the preset key operation execution sequence, when performing each key operation sequentially on the target operator based on the real-time pose state data of the target operator under each key operation, the process further includes: The actual pose state data of the target operation after the current key operation is performed is compared with the reference pose state data; If the comparison result is determined to be satisfactory, the next key operation is executed; if the comparison result is determined to be unsatisfactory, the real-time pose state data of the target manipulator under the current key operation is adjusted based on the visual image until the comparison result between the actual pose state data of the target manipulator and the reference pose state data after the current key operation is executed is satisfactory, and then the next key operation is executed.

10. A system for inferring the state of an operating component, characterized in that, include: An image acquisition unit is used to acquire a visual image of the target operating area where the target operating component is located through a vision system; wherein, the visual image includes at least two non-reflective markers set within the target operating area; The real-time construction unit is used to perform marker detection and coordinate system transformation on the visual image to obtain the real-time physical space coordinates of each non-reflective marker, and to construct the real-time local coordinate system of the target operation area based on the real-time physical space coordinates of each non-reflective marker. A spatial transformation unit is used to determine the real-time spatial transformation relationship of the target operating area based on the real-time local coordinate system and the reference local coordinate system of the target operating area; wherein, the reference local coordinate system is a local coordinate system of the target operating area that is pre-constructed based on the reference physical space coordinates of each non-reflective marker; The state inference unit is used to infer the real-time pose state data of the target manipulator in the real-time local coordinate system based on the pre-calibrated reference pose state data of the target manipulator in the reference local coordinate system and by using the real-time spatial transformation relationship.