Equipment operation real-time guidance teaching method and system based on augmented reality
By establishing spatial semantic relationships and collecting viewpoint and hand posture information in real time in augmented reality technology, and generating viewpoint-adaptive guidance anchor points and operation correction instructions, the problem of insufficient viewpoint change and real-time feedback in existing technologies is solved, and efficient and accurate training for operating complex equipment is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YIKE SHUREN (SUZHOU) EDUCATION TECHNOLOGY CO LTD
- Filing Date
- 2026-01-14
- Publication Date
- 2026-05-12
AI Technical Summary
Existing augmented reality operation guidance technologies fail to fully consider the impact of changes in the operator's actual perspective on visible components and lack real-time feedback mechanisms, resulting in decreased operational efficiency and accuracy, especially in precision operation scenarios involving complex equipment.
By establishing spatial semantic relationships based on device spatial information, a decomposed sequence of operation steps is generated. The operator's perspective and hand posture information are collected in real time, spatial deviation vectors are calculated, perspective-adapted guidance anchors and real-time operation correction instructions are generated, the operation completion status is continuously monitored, and the guidance content is dynamically updated.
It achieves precise matching between operation guidance and the operator's actual perspective and actions, improves the real-time nature, relevance and intuitiveness of teaching, ensures the continuity and smoothness of operation, and significantly improves the efficiency and learning effect of equipment operation training.
Smart Images

Figure CN122018681A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of augmented reality technology, and in particular to a real-time guided teaching method and system for device operation based on augmented reality. Background Technology
[0002] With the rapid development of industrial manufacturing and equipment maintenance, the operation and maintenance of complex equipment places increasingly higher demands on the professional skills of operators. Traditional equipment operation training mainly relies on paper manuals, video tutorials, or apprenticeships, which have many limitations in practical applications. In recent years, the development of augmented reality technology has provided a new solution for equipment operation training. By overlaying virtual guidance information onto real operating scenarios, it can effectively improve operators' learning efficiency and operational accuracy.
[0003] Current augmented reality (AR) teaching systems have already seen some application in the industrial sector, essentially presenting operational instructions as virtual annotations on the device. Operators can wear AR glasses or use mobile devices to view the overlaid virtual information and complete device operation tasks according to preset steps. These systems typically employ image recognition or spatial positioning technology to align virtual information with the real device and display corresponding guidance content to the operator based on predefined operational procedures.
[0004] However, existing augmented reality (AR) operation guidance technologies still have many shortcomings. Current technologies typically employ a fixed-viewpoint guidance method, failing to fully consider the impact of changes in the operator's actual perspective on visible components. When the operator observes the device from different angles, some target operation components may be obscured by other structures and thus invisible. Even in this case, guidance information is still displayed at the obscured location, making it difficult for the operator to understand the instructions and affecting operational efficiency and accuracy. Existing technologies lack a real-time feedback mechanism for the operator's actual actions during operation guidance. They typically only display preset operation steps sequentially, unable to provide dynamic correction guidance based on the actual deviation between the operator's hand position and the target component. Operators are prone to positional deviations or movement errors during operation, especially in precision operation scenarios, where this lack of real-time correction severely impacts operation quality. Existing technologies rely primarily on time delays or manual confirmation to determine the completion status of operation steps, lacking intelligent recognition capabilities based on spatial position relationships. This makes it difficult to accurately determine whether the operator has truly completed the current step, potentially jumping to the next step before completion or remaining at the current step even after completion, affecting the overall continuity and effectiveness of the teaching process. Summary of the Invention
[0005] The embodiments of the present invention provide a method and system for real-time guided teaching of device operation based on augmented reality, which can at least solve some of the problems existing in the prior art.
[0006] A first aspect of the present invention provides a real-time guided teaching method for device operation based on augmented reality, comprising: Based on the spatial information of the equipment, the spatial semantic relationship of the equipment is determined, and a decomposed sequence of operation steps is generated according to the description of the operation task and the spatial semantic relationship. Spatial registration is performed between the real-time collected operator's perspective information and the spatial semantic association to determine the set of visible operation components within the current field of view. Based on the spatial occlusion relationship between the target operation component of the current step in the decomposition sequence and the set of visible operation components, a viewpoint-adapted guide anchor point is generated. Based on the guide anchor point and the operator's hand pose information collected in real time, a spatial deviation vector between the operator's hand and the target operating component corresponding to the guide anchor point is calculated. A real-time operation correction command is generated based on the spatial deviation vector, and the real-time operation correction command is superimposed on the display position of the guide anchor point in an augmented reality visualization form. The system continuously monitors the spatial proximity between the hand pose information and the target operating component corresponding to the guide anchor point. When the operator's hand enters the preset interaction space range of the target operating component, the system identifies the operation completion status. The current execution position of the decomposed sequence is updated according to the operation completion status, and the presentation content of the target operation component and augmented reality guidance information is re-determined based on the updated current execution position.
[0007] Based on the spatial information of the equipment, the spatial semantic association of the equipment is determined. According to the description of the operation task and the spatial semantic association, a decomposed sequence of operation steps is generated, including: Each operable component in the device spatial information is labeled with functional attributes, each operable component is mapped to a predefined operation function category, and the spatial position of each operable component is extracted as spatial coordinate features; Based on the spatial coordinate characteristics, calculate the spatial distance and directional relationships between the operable components, and establish spatial topological connections between the operable components according to the spatial distance and directional relationships; The functional attribute annotations are associated and mapped with the spatial topological connections to form the spatial semantic associations. Parse the target state defined in the operation task description and decompose the target state into multiple sub-target states; Based on the functional attribute annotations in the spatial semantic association, the target operation component corresponding to each sub-target state is matched, and the operation sequence between each target operation component is determined according to the spatial topology connection. Based on the order of operations and the functional attribute labels of each target operation component, a decomposed sequence of operation steps is generated.
[0008] Spatially register the operator's perspective information collected in real time with the spatial semantic association to determine the set of visible operation components within the current field of view. Based on the spatial occlusion relationship between the target operation component of the current step in the decomposition sequence and the set of visible operation components, generate viewpoint-adapted guidance anchor points, including: Extract the spatial viewpoint coordinates and field of view direction vector from the operator's perspective information, and determine the field of view projection area based on the spatial viewpoint coordinates and field of view direction vector; The spatial positions of each operable component recorded in the spatial semantic association are obtained, and the spatial inclusion relationship between the spatial positions of each operable component and the field projection area is determined. Candidate visible components whose spatial positions fall within the field projection area are then selected. For each candidate visible component, the set of visible operational components within the current field of view is determined based on the spatial relationship between the spatial viewpoint coordinates and the spatial position of the candidate visible component. The target operation component for the current step is obtained from the decomposition sequence. When the target operation component does not belong to the set of visible operation components, the spatial position of the associated visible component that is spatially associated with the target operation component and belongs to the set of visible operation components is determined based on the spatial topological connection relationship in the spatial semantic association relationship, so as to generate the viewpoint-adaptive guide anchor point. When the target operating component belongs to the set of visible operating components, the viewpoint-adaptive guide anchor point is generated based on the spatial position of the target operating component.
[0009] For each candidate visible component, based on the spatial relationship between the spatial viewpoint coordinates and the spatial position of the candidate visible component, the set of visible operational components within the current field of view is determined, including: Traverse the candidate visible components, and for any candidate visible component, determine the line-of-sight vector pointing from the spatial viewpoint coordinates to the spatial position of the candidate visible component; Obtain the spatial position and geometric shape information of other geometric components in the three-dimensional geometry of the device, excluding the candidate visible components; The spatial intersection operation is performed between the line-of-sight vector and the spatial position and geometric shape information of the other geometric components to determine whether the line-of-sight vector intersects with other geometric components during its propagation from the spatial viewpoint coordinates to the spatial position of the candidate visible component. When the result of the spatial intersection operation indicates that there is no spatial intersection point, the visibility state of the candidate visible component is determined to be visible; otherwise, the visibility state of the candidate visible component is determined to be invisible. The candidate visible components whose visibility status is visible are added to the temporary set of visible components. After traversing all candidate visible components, the temporary set of visible components is determined as the set of operational components visible within the current field of view.
[0010] Based on the guide anchor point and the operator's hand pose information collected in real time, a spatial deviation vector between the operator's hand and the target operating component corresponding to the guide anchor point is calculated. A real-time operation correction command is generated based on the spatial deviation vector. This real-time operation correction command is overlaid on the display position of the guide anchor point marker in an augmented reality visualization format, including: The spatial position of the target operating component is extracted from the guide anchor point, and the current spatial coordinates and hand posture direction of the operator's hand are extracted from the hand pose information; Calculate the spatial position deviation between the current spatial coordinates of the hand and the spatial position of the target operating component, and calculate the posture angle deviation between the hand posture direction and the predefined operating posture direction required to complete the current step operation. Construct the spatial deviation vector based on the spatial position deviation and the posture angle deviation. Based on the position correction component of the spatial deviation vector, path guidance information indicating the direction and amplitude of hand movement is generated; based on the posture correction component of the spatial deviation vector, posture adjustment information indicating the direction and amplitude of hand rotation is generated; the path guidance information and the posture adjustment information are combined to form the real-time operation correction command. The real-time operation correction command is converted into an augmented reality visualization element, and the augmented reality visualization element is overlaid on the display position of the guide anchor point identifier.
[0011] Continuously monitor the spatial proximity between the hand pose information and the target operating component corresponding to the guide anchor point. When the operator's hand is detected to have entered the preset interaction space range of the target operating component, identify the operation completion status, including: The system continuously acquires the real-time updated hand pose information and extracts the real-time spatial coordinates of the operator's hand from the hand pose information. The spatial position of the target operating component is obtained from the guide anchor point, and the preset interaction space range of the target operating component is determined based on the spatial position of the target operating component and the geometric boundary information of the target operating component; Calculate the spatial distance between the real-time spatial coordinates of the hand and the spatial position of the target operating component, and determine the spatial proximity between the hand pose information and the target operating component based on the spatial inclusion relationship between the spatial distance and the real-time spatial coordinates of the hand relative to the preset interactive space range; When the real-time spatial coordinates of the hand fall within the preset interactive space range based on the spatial proximity, the hand movement features in the hand pose information are extracted, and the hand movement features are matched with the predefined operation action mode of the current step; Based on the matching result between the hand movement features and the predefined operation action pattern of the current step, the operation completion status is identified.
[0012] Calculate the spatial distance between the real-time spatial coordinates of the hand and the spatial position of the target operating component, and based on the spatial inclusion relationship between the spatial distance and the real-time spatial coordinates of the hand relative to the preset interactive space range, determine the spatial proximity between the hand pose information and the target operating component, including: Calculate the Euclidean spatial distance between the real-time spatial coordinates of the hand and the spatial position of the target operating component; Obtain spatial boundary description information of the preset interactive space range, wherein the spatial boundary description information defines the geometric shape and range of the preset interactive space range in three-dimensional space; Based on the real-time spatial coordinates of the hand and the spatial boundary description information, the inclusion relationship between the point and the spatial region is determined to determine whether the real-time spatial coordinates of the hand are within the preset interactive space range. Calculate the shortest distance from the real-time spatial coordinates of the hand to the boundary of the preset interactive space range, determine the distance proximity index based on the Euclidean spatial distance, and determine the spatial position relationship index based on the inclusion relationship determination result of the point and the spatial region and the shortest distance; The distance proximity index and the spatial position relationship index are quantified to obtain the spatial proximity between the hand pose information and the target operating component.
[0013] A second aspect of the present invention provides a real-time guided teaching system for device operation based on augmented reality, comprising: The first unit is used to determine the spatial semantic association of the device based on the device spatial information, and to generate a decomposed sequence of operation steps according to the operation task description and the spatial semantic association. The second unit is used to spatially register the operator's perspective information collected in real time with the spatial semantic association, determine the set of visible operation components within the current field of view, and generate viewpoint-adapted guide anchor points based on the spatial occlusion relationship between the target operation component of the current step in the decomposition sequence and the set of visible operation components. The third unit is used to calculate the spatial deviation vector between the operator's hand and the target operating component corresponding to the guide anchor point based on the guide anchor point and the real-time collected hand posture information of the operator, and to generate a real-time operation correction command based on the spatial deviation vector. The real-time operation correction command is superimposed on the display position of the guide anchor point in an augmented reality visualization form. The fourth unit is used to continuously monitor the spatial proximity between the hand pose information and the target operating component corresponding to the guide anchor point. When the operator's hand is detected to have entered the preset interaction space range of the target operating component, the operation completion status is identified. The fifth unit is used to update the current execution position of the decomposed sequence according to the operation completion status, and to redetermine the presentation content of the target operation component and augmented reality guidance information based on the updated current execution position.
[0014] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0015] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0016] This invention establishes spatial semantic relationships based on equipment spatial information and automatically generates decomposition sequences according to operation tasks, realizing intelligent decomposition and structured expression of complex equipment operation processes. This enables operation guidance to accurately match the actual spatial layout and operation logic of the equipment, improving the accuracy and applicability of teaching guidance and avoiding the tedious process of manually writing detailed operation manuals in traditional methods.
[0017] This invention dynamically identifies the set of visible operating components within the current field of view by collecting operator perspective information in real time and performing spatial registration. It generates viewpoint-adapted guidance anchors based on spatial occlusion relationships and generates real-time operation correction instructions based on the spatial deviation vector between the operator's hand and the target operating component. These instructions are presented intuitively in an augmented reality visualization format, achieving precise matching between operation guidance and the operator's actual perspective and actions. This significantly improves the real-time nature, relevance, and intuitiveness of operation guidance, while reducing the operator's cognitive burden and operational difficulty.
[0018] This invention continuously monitors the spatial proximity of hand posture information to the target operating component, automatically identifies the operation completion status, and dynamically updates the execution position of the decomposed sequence. This achieves automatic switching of operation steps and intelligent updating of guidance information, forming a closed-loop interactive teaching process. It ensures the continuity and smoothness of operation guidance, effectively improving the efficiency and learning effect of equipment operation training, and is particularly suitable for on-site operation training scenarios of complex equipment. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the real-time guided teaching method for device operation based on augmented reality, according to an embodiment of the present invention.
[0020] Figure 2 This is a flowchart illustrating the process of generating real-time operation correction instructions in a visual form according to an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0023] Figure 1 This is a flowchart illustrating the real-time guided teaching method for device operation based on augmented reality, according to an embodiment of the present invention. Figure 1 As shown, the method includes: Based on the spatial information of the equipment, the spatial semantic relationship of the equipment is determined, and a decomposed sequence of operation steps is generated according to the description of the operation task and the spatial semantic relationship. Spatial registration is performed between the real-time collected operator's perspective information and the spatial semantic association to determine the set of visible operation components within the current field of view. Based on the spatial occlusion relationship between the target operation component of the current step in the decomposition sequence and the set of visible operation components, a viewpoint-adapted guide anchor point is generated. Based on the guide anchor point and the operator's hand pose information collected in real time, a spatial deviation vector between the operator's hand and the target operating component corresponding to the guide anchor point is calculated. A real-time operation correction command is generated based on the spatial deviation vector, and the real-time operation correction command is superimposed on the display position of the guide anchor point in an augmented reality visualization form. The system continuously monitors the spatial proximity between the hand pose information and the target operating component corresponding to the guide anchor point. When the operator's hand enters the preset interaction space range of the target operating component, the system identifies the operation completion status. The current execution position of the decomposed sequence is updated according to the operation completion status, and the presentation content of the target operation component and augmented reality guidance information is re-determined based on the updated current execution position.
[0024] When determining spatial semantic relationships based on equipment spatial information, the first step is to perform 3D scanning and modeling of the target equipment to obtain precise 3D structural data. A depth camera is used to collect point cloud data, and a point cloud registration algorithm is used to fuse the multi-angle point cloud data into a complete 3D model of the equipment. For complex industrial equipment, such as CNC machine tools, approximately 50,000 to 100,000 spatial points can be collected with millimeter-level accuracy. Based on this, semantic segmentation is performed, dividing the equipment into multiple functional components. For example, a CNC machine tool is divided into components such as the control panel, spindle, and worktable. A spatial relationship diagram between components is established, recording the relative position, functional attributes, and operational constraints of each component. For example, for industrial control valves, the connection relationship between the valve handle and the valve body, as well as the correspondence between the handle's rotation angle and the valve's open / closed state, are recorded. Furthermore, a mapping is established between textual instructions in the equipment operation manual and specific components in the 3D model, forming a semantic connection from operational descriptions to physical components. The system receives an operation task description, such as "replace printer ink cartridge," and extracts key operation elements using natural language processing technology. It then breaks down the task into ordered basic operation steps, such as "open printer front cover," "remove old ink cartridge," "install new ink cartridge," and "close printer front cover." Based on the dependencies between components and operational constraints, the system optimizes the operation sequence to ensure the rationality and safety of the operation.
[0025] In achieving spatial registration that links viewpoint with spatial semantics, the operator's viewpoint is captured in real-time by a camera on a head-mounted display device, with a sampling rate of no less than 30 frames per second. Feature points are extracted from the video frames and matched with a pre-built 3D model of the device. SLAM technology is used to calculate the precise pose of the camera, with errors controlled within centimeters. The camera coordinate system from the operator's viewpoint is transformed to the world coordinate system of the device model, achieving precise alignment between the virtual augmented content and the real device. Based on the current camera pose, the visibility of each component of the device in the field of view is calculated, considering factors such as component size, distance, and occlusion relationships to construct a set of visible components. For industrial equipment control panels, the visibility status of more than 30 operation buttons and indicator lights can be tracked simultaneously. The target operation component is determined according to the current operation step, such as the valve handle in the "rotate valve handle" operation. The visibility of the target component in the current field of view is analyzed. If it is partially occluded or completely invisible, guidance prompts are generated to guide the operator to adjust the viewpoint. When the target component is visible, visual guidance marks, such as arrows or highlighted outlines, are superimposed on the component as guidance anchor points. Based on the shape and operational characteristics of the components, the position and shape of the guide anchor points are dynamically adjusted to ensure that the guide information is clearly visible.
[0026] When calculating spatial deviations based on guide anchor points and operator hand poses, the operator's hand pose data is acquired in real time using a depth camera or dedicated hand tracking device, with a sampling frequency of no less than 60Hz, capturing finger joint positions with millimeter-level accuracy. Key hand points are identified, including the palm center, fingertip positions, and joints, to construct a hand skeletal model, enabling precise tracking of finger bending, palm orientation, and other states. The three-dimensional spatial distance and orientation differences between the key hand points and the target operating component are calculated, forming a spatial deviation vector. For example, for knob operation, the distance vector from the index fingertip to the knob center and the angle difference between the palm orientation and the knob axis are calculated. Based on the magnitude and direction of the deviation vector, corresponding visual correction instructions are generated, such as displaying a "closer" arrow when the distance is too far and a rotation direction indicator when the angle is incorrect. These correction instructions are overlaid on the guide anchor point position using augmented reality graphics, such as displaying a downward arrow above the target button to indicate the correct operating direction. For fine operations, such as circuit board component installation, millimeter-level position guidance can be provided, intuitively displaying the insertion position and direction.
[0027] During continuous monitoring of hand gestures and interactions with target components, an interaction space model centered on the target component is established, defining different distance threshold regions. For example, for button operations, a 3cm range can be defined as the proximity zone, and a 1cm range as the interaction zone. The positional relationship between key hand points and the interaction space is calculated in real time to determine the current interaction state. The time point when the hand enters the preset interaction space is detected, and gesture changes are analyzed to identify the type of operation, such as pressing, rotating, or dragging. Combining the gesture type and the characteristics of the component, it is determined whether the operation meets expectations. For example, for rotating a valve, it detects whether the hand grip posture and rotation angle reach a predetermined threshold, such as completing a 90-degree clockwise rotation. When the operation is confirmed to be completed correctly, an operation completion status is triggered, and success feedback is provided, such as highlighting the component or displaying a checkmark. Vibration feedback can also be added to enhance the operation perception. For continuous operation tasks, the completion status of each sub-step is recorded to form an operation progress tracking system.
[0028] When updating operation steps and guidance information, the execution position of the decomposed sequence is updated based on the completed operations. The current step is marked as completed, and the system switches to the next operation step, updating the operation progress indicator to display the overall completion percentage of the current task, such as "3 / 7 steps completed." The target operation component is redefined based on the updated current step, such as switching from "opening the front cover of the device" to "removing the filter." The visibility and optimal interaction path of the new target component are recalculated, generating new guidance anchors and visual cues. The presentation format of the augmented reality guidance content is adjusted, including text descriptions, operation diagrams, and interactive prompts. For complex operations, step-by-step animation demonstrations can be provided to intuitively show the correct operation method. This process continues to loop until all steps of the entire operation task are completed, finally displaying a summary of task completion and operation evaluation. This cyclical, progressive real-time guidance method provides operators with continuous and accurate operation guidance, significantly improving the learning efficiency and accuracy of device operation.
[0029] In a practical application case, taking the operation of a home coffee machine as an example, the process begins by acquiring a 3D model of the coffee machine, identifying key components such as the water tank, coffee powder container, and control panel, and establishing a spatial relationship model between these components. For the task of "making a cup of Americano," this is broken down into steps such as "adding water," "adding coffee powder," "placing the coffee cup," "selecting the Americano mode," and "starting the brewing process." When the operator puts on AR glasses and looks at the coffee machine, spatial registration identifies the coffee machine components in the current field of view, displaying a flashing blue arrow above the water tank as a guiding anchor point for the first step. When the operator's hand approaches the water tank, an animation of "lifting upwards" is displayed to indicate the correct action of removing the water tank. If the operator's hand movement is incorrect, a correction arrow is displayed in real time to guide the correct direction. Once the water tank is successfully removed, the step status is updated, and the guiding anchor point is moved to the faucet position to guide the water filling operation. In this way, the operator is guided through all steps, ensuring that each operation is performed correctly.
[0030] In one optional implementation, the spatial semantic association of the devices is determined based on device spatial information, and a decomposed sequence of operation steps is generated according to the operation task description and the spatial semantic association, including: Each operable component in the device spatial information is labeled with functional attributes, each operable component is mapped to a predefined operation function category, and the spatial position of each operable component is extracted as spatial coordinate features; Based on the spatial coordinate characteristics, calculate the spatial distance and directional relationships between the operable components, and establish spatial topological connections between the operable components according to the spatial distance and directional relationships; The functional attribute annotations are associated and mapped with the spatial topological connections to form the spatial semantic associations. Parse the target state defined in the operation task description and decompose the target state into multiple sub-target states; Based on the functional attribute annotations in the spatial semantic association, the target operation component corresponding to each sub-target state is matched, and the operation sequence between each target operation component is determined according to the spatial topology connection. Based on the order of operations and the functional attribute labels of each target operation component, a decomposed sequence of operation steps is generated.
[0031] During the automated processing of equipment operation tasks, spatial information data of the device to be operated is received as input. This spatial information data includes the geometric, positional, and structural information of all interactive elements on the device. For a typical multifunction printer, its spatial information includes multiple operable components such as the paper feed tray, paper output tray, ink cartridge door, power switch, and display panel. This spatial information is parsed to extract the three-dimensional coordinate data of each operable component. The center position coordinates of the paper feed tray are offset from the device origin by 120 mm on the X-axis, 80 mm on the Y-axis, and 50 mm on the Z-axis. The center position coordinates of the paper output tray are offset from the device origin by 120 mm on the X-axis, 280 mm on the Y-axis, and 150 mm on the Z-axis.
[0032] Each operable component undergoes functional attribute labeling. This labeling process matches the components to a pre-established device function ontology library, which defines various operational function categories such as switches, containers, adjustments, and displays. For example, the paper feed tray is identified as having both material receiving and material conveying triggering functions, and is labeled as a container / trigger component with the function tag "Paper Loading Unit." The ink cartridge compartment door is identified as having both blocking and opening / closing functions, and is labeled as a switch component with the function tag "Consumable Replacement Inlet." The power switch is labeled as a switch component with the function tag "Equipment Start Control Unit." The display panel is labeled as a display / interaction component with the function tag "Status Feedback and Parameter Setting Interface." This labeling process assigns clear functional semantic information to each component.
[0033] After obtaining the spatial coordinate characteristics of each component, the spatial distance relationship between any two operable components is calculated. Taking the paper feed tray and paper output tray as examples, the center coordinate points of both are extracted, and the straight-line distance between the two points is calculated to be 206 mm. Simultaneously, the projected distances of the two components on each coordinate axis are calculated: 200 mm on the Y-axis, 100 mm on the Z-axis, and 0 mm on the X-axis. Based on these distance data, the paper feed tray and paper output tray are determined to be adjacent components in the equipment structure, indicating a direct spatial relationship between them. For the ink cartridge compartment door and the paper feed tray, the calculated distance between their center points is 180 mm, which is considered a medium distance relationship.
[0034] Further analysis of the directional relationships between operable components reveals that, for the paper feed tray and paper output tray, the paper output tray is located directly in front of and above the paper feed tray, with a directional vector of 200 mm along the positive Y-axis and 100 mm along the positive Z-axis. For the power switch and display panel, the power switch is located at the lower right of the display panel, with a directional vector of 60 mm along the positive X-axis and 30 mm along the negative Z-axis. These directional relationships reflect the relative spatial arrangement of the components.
[0035] Based on the calculated spatial distance and directional relationships, a spatial topological connection is established between operable components. This topological connection is stored in the form of a graph, with each operable component acting as a node in the graph, and the spatial relationships between components acting as connecting edges. For component pairs with a distance of less than 250 mm and a clear directional relationship, a direct connecting edge is established between them. A connecting edge is established between the paper feed tray and the paper output tray, carrying a distance attribute of 206 mm and a directional attribute of "front and above". A connecting edge is established between the ink cartridge door and the paper feed tray, carrying a distance attribute of 180 mm and a directional attribute of "right". Through this topological connection, a spatial proximity network of equipment components is constructed.
[0036] The labeled functional attributes are mapped to the established spatial topological connections. This mapping process adds semantic annotation information to each edge in the spatial topological connections. Taking the connection edge between the paper inlet tray and the paper outlet tray as an example, combining their functional labels "paper loading unit" and "paper output unit," the semantic relationship of this connection edge is labeled as "material flow path." For the connection between the ink cartridge door and the display panel, the semantic relationship is labeled as "maintenance operation guidance relationship," indicating that the display panel can provide status prompts for the operation of the ink cartridge door. For the connection between the power switch and all other components, the semantic relationship is uniformly labeled as "pre-activation dependency," indicating that the operation of the power switch is a prerequisite for the functioning of other components. Through this mapping, a spatial relationship network containing functional semantics is formed.
[0037] The system receives an operation task description as input, provided in natural language or structured instruction form. For the task description "Complete the initial use preparation of the device," the system parses the defined target state as the device being in a ready state to perform a print job. This target state is then broken down into multiple sub-target states, including device power supply activation, paper feed channel availability, ink cartridge system availability, and user interface responsiveness. Each sub-target state corresponds to a specific functional subsystem of the device reaching a specific operating condition.
[0038] Based on the functional attribute annotations in the spatial semantic association, a corresponding target operation component is matched for each sub-target state. For the sub-target whose power supply state is active, components with the functional tag "start control" are searched, and the power switch is matched as the target operation component. For the sub-target whose paper supply channel is available, the paper tray is matched as the target operation component. For the sub-target whose ink cartridge system is available, the ink cartridge door is matched as the target operation component. For the sub-target whose operating interface is in a responsive state, the display panel is matched as the target operation component.
[0039] Based on the spatial topology connections, the order of operations among the target operating components is determined. The semantic relationships between these components are examined, identifying a "pre-activation dependency" between the power switch and all other components, thus determining that the power switch operation must be prioritized. For the ink cartridge door and paper feed tray, their operational independence is analyzed, determining that there is no mandatory sequential constraint between them. However, considering the continuity of the material flow path, the ink cartridge door operation is prioritized before the paper feed tray operation. For the display panel, it is identified as a status feedback unit and should be confirmed after the operations of other physical components are completed, thus placing it at the end of the operation sequence.
[0040] Based on the determined sequence of operations and the functional attributes of each target operating component, a decomposed sequence of operation steps is generated. This sequence is an ordered list. The first operation is to press the power switch to activate the device's power supply; the expected result is that the device's indicator light illuminates. The second operation is to open the ink cartridge compartment door to expose the ink cartridge installation location; the expected result is that the compartment door is open. The third operation is to install the ink cartridge in the ink cartridge compartment area; the purpose is to load printing consumables; the expected result is that the ink cartridge is inserted into place. The fourth operation is to close the ink cartridge compartment door to seal the consumable compartment; the expected result is that the compartment door is locked and the ink cartridge is recognized. The fifth step is to load paper onto the paper feed tray to provide printing media; the expected result is that the paper is positioned in the tray. The sixth operation is to confirm the parameters on the display panel; the purpose is to complete the settings and enter the ready state; the expected result is that the display panel displays a ready message. This operational step decomposition sequence fully covers the transformation process from the initial state to the target state, with each step associated with a specific operational object, operational action type, and expected state change.
[0041] In one optional implementation, the real-time collected operator's perspective information is spatially registered with the spatial semantic association to determine the set of visible operating components within the current field of view. Based on the spatial occlusion relationship between the target operating component in the current step of the decomposition sequence and the set of visible operating components, perspective-adaptive guidance anchors are generated, including: Extract the spatial viewpoint coordinates and field of view direction vector from the operator's perspective information, and determine the field of view projection area based on the spatial viewpoint coordinates and field of view direction vector; The spatial positions of each operable component recorded in the spatial semantic association are obtained, and the spatial inclusion relationship between the spatial positions of each operable component and the field projection area is determined. Candidate visible components whose spatial positions fall within the field projection area are then selected. For each candidate visible component, the set of visible operational components within the current field of view is determined based on the spatial relationship between the spatial viewpoint coordinates and the spatial position of the candidate visible component. The target operation component for the current step is obtained from the decomposition sequence. When the target operation component does not belong to the set of visible operation components, the spatial position of the associated visible component that is spatially associated with the target operation component and belongs to the set of visible operation components is determined based on the spatial topological connection relationship in the spatial semantic association relationship, so as to generate the viewpoint-adaptive guide anchor point. When the target operating component belongs to the set of visible operating components, the viewpoint-adaptive guide anchor point is generated based on the spatial position of the target operating component.
[0042] During real-time operation guidance, the system continuously collects perspective information provided by the augmented reality device worn by the operator. This perspective information includes two key parameters: spatial viewpoint coordinates and a field of view direction vector. The spatial viewpoint coordinates represent the position of the operator's head in three-dimensional space, expressed as three-dimensional coordinates in the world coordinate system. For example, the operator's current position is represented by coordinates (1.2 meters, 0.8 meters, 1.6 meters). The field of view direction vector describes the direction the operator is looking, expressed as a unit vector. For example, a direction vector of (0.707, 0.0, -0.707) indicates that the operator is looking in a direction that is 45 degrees to the left of the front.
[0043] Based on the acquired spatial viewpoint coordinates and the field of view direction vector, the field of view projection area is calculated. Determining the field of view projection area requires considering the field of view angle parameters of the augmented reality device (ARD). Typically, ARD devices have a horizontal field of view of 60 degrees and a vertical field of view of 45 degrees. Starting from the spatial viewpoint coordinates, extending along the field of view direction vector, a cone-shaped region is constructed in three-dimensional space according to the field of view angle parameters. This cone-shaped region is the field of view projection area. Considering the actual operating distance, the effective depth range of the field of view projection area is set to 0.3 meters to 3.0 meters; spatial areas exceeding this depth range are not included in the field of view projection calculation.
[0044] The spatial position information of each operable component is read from a pre-established spatial semantic association data structure. For example, a mechanical device contains 20 operable components, and the spatial position of each component is stored in three-dimensional coordinates. The position of component A is (1.5 m, 0.6 m, 1.2 m), and the position of component B is (1.8 m, 0.9 m, 1.4 m). The spatial position coordinates of each component are read one by one, and spatial inclusion relationship is determined with the previously determined field of view projection area. The determination process is achieved by calculating the angle between the component position point and the field of view direction vector, as well as the spatial distance from the position point to the viewpoint coordinates. When the angle between a component position point and the field of view direction vector is less than half of the field of view angle, and the spatial distance is within the effective depth range, the component is marked as a candidate visible component.
[0045] For the selected candidate visible components, occlusion determination is further performed. Occlusion determination is based on a ray projection method between the spatial viewpoint coordinates and the candidate visible components. A virtual ray is emitted from the spatial viewpoint coordinates towards the position of each candidate visible component, and the presence of other physical obstructions along the ray path is detected. Information about obstructions is also stored in the spatial semantic association, including the device casing, internal support structure, and other components. For example, if a ray is emitted from viewpoint coordinates (1.2 m, 0.8 m, 1.6 m) towards component A's position (1.5 m, 0.6 m, 1.2 m), a metal panel with a thickness of 0.05 m is detected along the ray path. This panel's spatial position is between (1.3 m, 0.7 m, 1.4 m) and (1.4 m, 0.7 m, 1.3 m), causing the ray to be blocked. In this case, although component A is a candidate visible component, it is not included in the set of visible operational components due to occlusion. For component C, located at (1.3 m, 0.9 m, 1.5 m), the ray path from the viewpoint to this location is unobstructed; therefore, component C is identified as a visible operational component. After occlusion detection, a set of visible operational components within the current field of view is generated. This set includes all components that can indeed be observed by the operator from the current viewpoint.
[0046] Extract the target operation component identifier for the current step from the maintenance task decomposition sequence. For example, if the current step requires the operator to disassemble component D, search the set of visible operation components and determine if component D is included in the set. If component D is not in the set of visible operation components, it means that the component is occluded or outside the field of view from the current perspective, and the operator cannot directly see the target to be operated. Access the spatial topological connection relationship data stored in the spatial semantic association relationship, which records the spatial adjacency and functional connection relationships between components. For example, component D is connected to component E by bolts, and component E shares the same mounting panel with component F; these relationships are clearly recorded in the spatial topological connection relationship. Starting from the target operation component D, search for components that are directly or indirectly related to it along the topological connection relationship, prioritizing directly connected components, and then selecting components that share the same mounting location. Perform an intersection operation on the found related components and the set of visible operation components to filter out related visible components that are both spatially related to the target component and visible from the current perspective.
[0047] Assuming the search results show that component E belongs to the associated visible components, obtain the spatial coordinates of component E (1.4m, 0.85m, 1.45m). Based on this spatial position, generate the spatial coordinates of a guide anchor point. The guide anchor point is not placed directly at the center of component E, but rather offset according to the orientation of component E relative to the target component D. The spatial semantic association records the relative orientation information of component E and component D; for example, component D is located to the right and slightly behind component E. Based on the spatial position of component E, offset 0.1m along the direction pointing to component D to obtain the final position of the guide anchor point (1.5m, 0.85m, 1.4m). The semantic label of this guide anchor point is set to "target component in this direction," prompting the operator to adjust the viewing angle to observe in this direction.
[0048] When the target operable component belongs to the set of visible operable components, it means that the operator can directly see the target to be operated on. The spatial location information of the target operable component is directly extracted; for example, the location of component G is (1.6 meters, 0.75 meters, 1.3 meters). A guide anchor point is generated at this spatial location, and the semantic label of the anchor point is set to "operation location". To enhance the guidance effect, the display parameters of the anchor point are also adjusted according to the geometric bounding box information of the target operable component. For example, if the bounding box dimensions of component G are 0.08 meters long, 0.06 meters wide, and 0.04 meters high, the display radius of the guide anchor point is set to 1.2 times the maximum size of the bounding box to ensure that the virtual identifier can completely cover the target component area and avoid positional deviations for the operator.
[0049] The generated guide anchor point information is then transmitted to the augmented reality rendering module, where the corresponding virtual guide markers are overlaid and displayed in the operator's field of vision, thus providing real-time spatial guidance to the operator.
[0050] In one optional implementation, for each candidate visible component, based on the spatial relationship between the spatial viewpoint coordinates and the spatial position of the candidate visible component, the set of visible operational components within the current field of view is determined, including: Traverse the candidate visible components, and for any candidate visible component, determine the line-of-sight vector pointing from the spatial viewpoint coordinates to the spatial position of the candidate visible component; Obtain the spatial position and geometric shape information of other geometric components in the three-dimensional geometry of the device, excluding the candidate visible components; The spatial intersection operation is performed between the line-of-sight vector and the spatial position and geometric shape information of the other geometric components to determine whether the line-of-sight vector intersects with other geometric components during its propagation from the spatial viewpoint coordinates to the spatial position of the candidate visible component. When the result of the spatial intersection operation indicates that there is no spatial intersection point, the visibility state of the candidate visible component is determined to be visible; otherwise, the visibility state of the candidate visible component is determined to be invisible. The candidate visible components whose visibility status is visible are added to the temporary set of visible components. After traversing all candidate visible components, the temporary set of visible components is determined as the set of operational components visible within the current field of view.
[0051] When determining the set of visible operational components within the current field of view, the visibility of each candidate visible component needs to be assessed individually. This assessment process is based on spatial geometric analysis, using a ray tracing algorithm to verify whether there is line-of-sight occlusion between the observation point and the target component.
[0052] For a candidate visible component, obtain its position coordinates in 3D space. Assume the viewpoint is point P with 3D coordinates (50, 80, 120) in millimeters. The candidate visible component is tagged with point Q with 3D coordinates (150, 200, 180). Calculate the line-of-sight vector from point P to point Q. This vector is obtained by subtracting the starting point coordinates from the ending point coordinates, resulting in vector components (100, 120, 60). To facilitate subsequent calculations, normalize this vector by dividing each component by the vector length to obtain a unit direction vector with a length of approximately 161.2 millimeters. The normalized direction vector components are approximately (0.62, 0.74, 0.37).
[0053] Access the device's 3D geometry database to extract information on all other geometric components besides the currently visible candidate components. These components include the device's outer shell panels, internal support frames, and other functional modules. Each geometric component stores its spatial location parameters and geometric shape description in the database. For example, a rectangular panel is defined by four vertices with 3D coordinates (80, 100, 140), (180, 100, 140), (180, 180, 140), and (80, 180, 140), and is parallel to the XY plane. Another cylindrical component's parameters include the starting coordinates of its axis (120, 150, 100), the ending coordinates of its axis (120, 150, 200), and a radius of 15 mm.
[0054] The spatial intersection operation is performed sequentially on the line of sight vector and each other geometric component. For a rectangular panel component, it determines whether the line containing the line of sight vector intersects the plane containing the panel. The starting point of the line of sight is (50, 80, 120), the direction vector is (0.62, 0.74, 0.37), the normal vector of the plane containing the panel is (0, 0, 1), and the constant term in the plane equation is 140. The intersection point of the line of sight and the plane is calculated by adding the product of the parameter t and the direction vector to the starting point coordinates using the parametric equation of the line of sight, setting the Z-coordinate component to 140, and obtaining the parameter t to be approximately 54.05. Substituting this parameter value into the parametric equation of the line of sight, the coordinates of the intersection point are approximately (83.5, 120, 140). Further verification is made to ensure that the intersection point is within the boundary of the rectangular panel by checking whether the X-coordinate is between 80 and 180 and the Y-coordinate is between 100 and 180. The X coordinate of the intersection point, 83.5, is within the valid range, and the Y coordinate, 120, is also within the valid range. Therefore, it is confirmed that the line of sight has a valid intersection with the panel.
[0055] The distance between the intersection point and the viewpoint is calculated by subtracting the viewpoint coordinates from the intersection point coordinates, resulting in a vector (33.5, 40, 20) with a length of approximately 55.4 mm. Simultaneously, the distance between the candidate visible component and the viewpoint is calculated, which is the original line-of-sight vector with a length of 161.2 mm. Comparing these two distance values reveals that the intersection point is closer to the viewpoint. This means that the line of sight intersects with the rectangular panel before reaching the candidate visible component, and the line of sight is obstructed by the panel.
[0056] For intersection calculations of cylindrical components, different geometric algorithms are used. The axis of the cylinder is determined by its two endpoints, and the intersection point of the line of sight with the side of the cylinder is calculated. The vector from the starting point of the line of sight to the starting point of the cylinder axis is projected onto a plane perpendicular to the axis, and the existence of an intersection point is determined in conjunction with the cylinder radius. Assuming that the minimum distance perpendicular to the cylinder axis in the parametric equation of a certain line of sight vector is 18 mm, and this value is greater than the cylinder radius of 15 mm, then it is determined that the line of sight does not intersect the cylinder.
[0057] After traversing all other geometric components, if at least one valid spatial intersection is found, and the distance of this intersection from the viewpoint is less than the distance of the candidate visible component from the viewpoint, then the candidate visible component is determined to be occluded, and its visibility state is marked as invisible. In the case of the rectangular panel above, since the line of sight is obstructed, the visibility state of the candidate visible component is set to invisible, and it is not added to the temporary set of visible components.
[0058] If no valid occlusion intersections are found after completing the intersection calculations for all other geometric components, the line of sight from the viewpoint to the candidate visible component is considered unobstructed. Assuming another candidate visible component is located at (200, 250, 300) and the line of sight vector is (150, 170, 180), when performing the intersection calculation on this line of sight, if all geometric components either do not intersect with the line of sight or the distance between their intersections is greater than the distance to the target component, then the visibility state of this candidate visible component is marked as visible, and it is added to the temporary set of visible components.
[0059] A dynamic list is maintained as a temporary set of visible components. Whenever a candidate visible component passes visibility verification, its identifier, name, spatial location, and other information are recorded in this set. After processing all candidate visible components, the temporary set of visible components contains all operable components visible from the current viewpoint. This temporary set is output as the final result, marked as the set of operable components visible within the current field of view, for use by subsequent interaction processing modules.
[0060] In one optional implementation, based on the guide anchor point and the operator's hand pose information acquired in real time, a spatial deviation vector between the operator's hand and the target operating component corresponding to the guide anchor point is calculated. A real-time operation correction command is generated based on the spatial deviation vector. The real-time operation correction command is superimposed on the display position of the guide anchor point marker in an augmented reality visualization format, including: The spatial position of the target operating component is extracted from the guide anchor point, and the current spatial coordinates and hand posture direction of the operator's hand are extracted from the hand pose information; Calculate the spatial position deviation between the current spatial coordinates of the hand and the spatial position of the target operating component, and calculate the posture angle deviation between the hand posture direction and the predefined operating posture direction required to complete the current step operation. Construct the spatial deviation vector based on the spatial position deviation and the posture angle deviation. Based on the position correction component of the spatial deviation vector, path guidance information indicating the direction and amplitude of hand movement is generated; based on the posture correction component of the spatial deviation vector, posture adjustment information indicating the direction and amplitude of hand rotation is generated; the path guidance information and the posture adjustment information are combined to form the real-time operation correction command. The real-time operation correction command is converted into an augmented reality visualization element, and the augmented reality visualization element is overlaid on the display position of the guide anchor point.
[0061] Figure 2 This is a flowchart illustrating the process of generating real-time operation correction instructions in a visual form according to an embodiment of the present invention. Figure 2 As shown, during real-time operation guidance, the spatial position information of the target operating component is extracted from the established guidance anchor point data structure. This spatial position information is stored in three-dimensional coordinates, including the X-axis, Y-axis, and Z-axis coordinates of the target operating component in the world coordinate system. For example, the position of an electrical terminal to be operated in the world coordinate system is -120 mm X-axis, 350 mm Y-axis, and 180 mm Z-axis. Simultaneously, the operation area range parameter corresponding to the operating component is extracted from the guidance anchor point data structure. This range parameter defines the three-dimensional spatial boundary where the effective operation occurs, such as a cubic space formed by extending 15 mm in each direction from the center of the terminal.
[0062] While extracting information about the target operating components, the current spatial coordinates of the operator's hand are extracted from the real-time data stream acquired by the hand tracking module. These hand spatial coordinates are derived by analyzing key points of the hand skeleton, specifically using the center point of the hand skeleton or the fingertip position of a particular operating finger as representative coordinates. For example, when the operation requires a touch operation using the right index finger, the three-dimensional coordinates of the right index fingertip are extracted, assuming these coordinates are -85 mm X, 420 mm Y, and 240 mm Z. In addition to coordinate information, the hand's posture direction information is also extracted from the hand pose data. This posture direction is calculated through the vector relationship between key points of the hand skeleton. Specifically, the main direction of the palm is determined by the vector formed by the wrist key point and the middle fingertip key point, the normal direction of the palm is determined by the vector formed by the thumb and little fingertip, and the cross product of these two direction vectors determines the third orthogonal direction of the palm. For example, at a certain moment, the main direction vector of the palm is 0.6 units along the positive X-axis, 0.8 units along the negative Y-axis, and 0 units along the Z-axis, while the normal direction vector of the palm is 0 units along the X-axis, 0 units along the Y-axis, and 1 unit along the positive Z-axis.
[0063] The spatial position deviation between the current spatial coordinates of the hand and the spatial position of the target operating part is calculated. This deviation is calculated by performing component difference operations on the two three-dimensional coordinates. Taking the aforementioned data as an example, the difference between the hand's fingertip coordinate (X: -85 mm) and the target terminal coordinate (X: -120 mm) is +35 mm; the Y-coordinate difference is 420 mm minus 350 mm, resulting in +70 mm; and the Z-coordinate difference is 240 mm minus 180 mm, resulting in +60 mm. This three-dimensional difference constitutes a spatial position deviation vector. The magnitude of this vector is calculated as approximately 99.5 mm using the square root of the sum of the squares of the three components, representing the straight-line distance between the hand and the target position. In subsequent processing, this spatial position deviation vector is decomposed into components along each axis of the world coordinate system, corresponding to correction actions of moving 35 mm to the left, 70 mm downwards, and 60 mm forwards, respectively.
[0064] Simultaneously, the attitude angle deviation between the hand posture direction and the predefined operation posture direction required to complete the current operation step is calculated. This predefined operation posture direction is stored in the operation posture parameters of the guide anchor point, and has different preset values for different types of operation actions. For example, for a vertical downward press operation, the predefined posture requires the palm's main direction vector to be along the negative Z-axis, and the palm's normal direction vector to be in the XY plane. The current hand posture direction vector is compared and analyzed with the predefined posture direction vector. The direction consistency value is obtained through the dot product operation between the two direction vectors. This value ranges from -1 to +1, with the closer the value is to +1, the more consistent the direction. When the dot product result is 0.7, it is converted to an angle value of approximately 45 degrees through inverse cosine operation, indicating a 45-degree deviation angle between the current hand posture and the target posture. Furthermore, the cross product operation of the two vectors determines which rotation axis the posture adjustment needs to be performed around, and the direction of the cross product result vector is the direction of the rotation axis. For example, if the cross product result vector is mainly along the positive Y-axis, it means that a 45-degree rotation adjustment around the Y-axis is required.
[0065] Based on the calculated spatial position and attitude angle deviations, a comprehensive spatial deviation vector data structure is constructed. This data structure includes two parts: a position correction component and an attitude correction component. The position correction component records the displacement deviation values along each axis in three-dimensional space. In the previous example, this includes a deviation of -35 mm on the X-axis, -70 mm on the Y-axis, and -60 mm on the Z-axis. The negative sign indicates that movement is needed in the negative direction of the coordinate axis to approach the target. The attitude correction component records the axial and angular values of rotational adjustments. In the previous example, this includes a parameter of -45 degrees rotation around the Y-axis. The negative sign indicates a clockwise rotation direction. This spatial deviation vector data structure also includes a deviation severity evaluation parameter. By comparing the magnitude of the position deviation with a set threshold, the deviation severity is divided into three levels: severe deviation, moderate deviation, and slight deviation. For example, a position deviation magnitude greater than 80 mm is considered severe deviation, between 50 mm and 80 mm is considered moderate deviation, and less than 50 mm is considered slight deviation.
[0066] Path guidance information is generated based on the position correction component of the spatial deviation vector. This information includes two parts: movement direction indication and movement magnitude indication. The movement direction indication is determined by analyzing the dominant direction of the position deviation vector. When the absolute value of the deviation component in a certain axis is significantly greater than that in other axes, that axis becomes the dominant movement direction. In the previous example, the Y-axis deviation of 70 mm is greater than the X-axis and Z-axis deviations, so the dominant movement direction is downward. When multiple axial deviation values are close, a combined direction indication is generated, such as moving simultaneously to the lower left and forward. The movement magnitude indication is achieved by mapping the deviation values to predefined magnitude levels. For example, a deviation of 0 to 20 mm is mapped to a fine adjustment, 20 to 60 mm to a moderate adjustment, and above 60 mm to a large adjustment. For the aforementioned total deviation of 99.5 mm, a large movement indication is generated, further refined to each axis as a moderate leftward movement, a large downward movement, and a large forward movement.
[0067] Attitude adjustment information is generated based on the attitude correction components of the spatial deviation vector. This information describes the rotational action the hand needs to perform and includes three elements: rotation axis, rotation direction, and rotation angle. The rotation axis is determined by the rotation axis vector obtained from attitude deviation analysis, such as the aforementioned rotation around the Y-axis. The rotation direction is determined by the sign of the rotation angle: a positive value indicates counterclockwise rotation, and a negative value indicates clockwise rotation. The rotation angle value represents the magnitude of adjustment required, which is also mapped to predefined adjustment levels. For example, 0 to 10 degrees is fine-tuning, 10 to 30 degrees is moderate adjustment, and above 30 degrees is significant adjustment. The aforementioned 45-degree attitude deviation is mapped to a significant clockwise rotation adjustment. The attitude adjustment information is also prioritized according to the accuracy requirements of the current operation; for high-precision operations, the correction priority of attitude deviation is higher than that of position deviation.
[0068] The generated path guidance information and attitude adjustment information are combined to form real-time operation correction instructions. These instructions are organized in a structured data format, including an instruction type field, a correction content field, a priority field, and a visualization parameter field. The instruction type field identifies whether the instruction is a position correction, attitude correction, or a combination of both. The correction content field stores specific parameters such as the direction of movement, range of movement, rotation axis, and rotation angle. The priority field sets the execution order of instructions based on the current operation stage and the degree of deviation. For example, position correction instructions are displayed first when the position deviation is large, and attitude correction instructions are displayed first when approaching the target position. The visualization parameter field defines the display attributes of the instruction in the augmented reality environment, including color coding, transparency, and animation effect type. For example, severe deviations are highlighted in red, moderate deviations in orange, and slight deviations in green.
[0069] Real-time operation correction commands are converted into augmented reality visualization elements. This conversion process selects the appropriate visualization format based on the type and content of the correction command. For path guidance information, a 3D arrow model is generated as a visualization element. The arrow's direction indicates the movement direction, its length indicates the movement range, and its color indicates the degree of deviation. For example, for the aforementioned command to move to the lower left and forward, an orange arrow is generated pointing from the current hand position to the target operating part position. The arrow length is scaled proportionally to the actual deviation distance to fit the display range. For posture adjustment information, a rotation indicator symbol is generated as a visualization element. This symbol displays the rotation direction as an arc arrow, and the arc length indicates the rotation angle. For example, for the aforementioned command to rotate 45 degrees clockwise around the Y-axis, a clockwise arc arrow displayed around the hand is generated, with the arc spanning a quarter of a circle to represent a 45-degree angle.
[0070] The generated augmented reality visualizations are overlaid on the display positions of the guide anchor points. This overlay operation is achieved through the augmented reality rendering engine, ensuring that the positions of the visualizations in 3D space remain correlated with the guide anchor points. Specifically, the starting point of the path guide arrow is anchored to the current hand position, and the ending point is anchored to the target operation part position. The arrow updates its starting position and recalculates its direction in real time as the hand moves. The posture adjustment symbol is anchored around the hand and moves with the hand, while updating the displayed content in real time according to changes in hand posture. Depth testing is also performed on the visualizations. When the hand or other objects occlude the visualizations, the occluded parts are displayed semi-transparently or temporarily hidden, ensuring that the operator can always clearly identify the guide information. The update frequency of the visualizations is synchronized with the hand tracking frequency, for example, refreshing the displayed content at a frequency of 30 frames per second, ensuring the smoothness and real-time nature of the guide information.
[0071] In one optional implementation, the spatial proximity of the hand pose information to the target operating component corresponding to the guide anchor point is continuously monitored. When the operator's hand is detected to have entered a preset interaction space range of the target operating component, the operation completion status is identified, including: The system continuously acquires the real-time updated hand pose information and extracts the real-time spatial coordinates of the operator's hand from the hand pose information. The spatial position of the target operating component is obtained from the guide anchor point, and the preset interaction space range of the target operating component is determined based on the spatial position of the target operating component and the geometric boundary information of the target operating component; Calculate the spatial distance between the real-time spatial coordinates of the hand and the spatial position of the target operating component, and determine the spatial proximity between the hand pose information and the target operating component based on the spatial inclusion relationship between the spatial distance and the real-time spatial coordinates of the hand relative to the preset interactive space range; When the real-time spatial coordinates of the hand fall within the preset interactive space range based on the spatial proximity, the hand movement features in the hand pose information are extracted, and the hand movement features are matched with the predefined operation action mode of the current step; Based on the matching result between the hand movement features and the predefined operation action pattern of the current step, the operation completion status is identified.
[0072] During system operation, the continuous monitoring module acquires hand pose information in real time at a preset sampling frequency, set between 30 Hz and 120 Hz, to ensure the capture of subtle changes in the operator's hand movements. A pose data stream containing the 3D coordinates of key hand points is acquired via a depth camera or other 3D sensing device. This pose data stream includes the spatial coordinates of at least 21 key feature points, including the wrist center point, fingertips, and knuckles. The 3D spatial coordinates of the wrist center point are extracted from this data stream as representative real-time spatial coordinates of the hand. These coordinates are represented as specific coordinate values in the world coordinate system; for example, the extracted wrist center point coordinates at a certain moment might be -120 mm on the X-axis, +350 mm on the Y-axis, and +800 mm on the Z-axis.
[0073] The spatial position information of the target operating component is read from a pre-established guide anchor point data structure. This data structure stores the center point coordinates of the target operating component and its associated geometric boundary parameters. Taking a valve knob as an example, the center point coordinates of the target operating component are labeled in the world coordinate system as -100 mm on the X-axis, +300 mm on the Y-axis, and +750 mm on the Z-axis. It also records the cylindrical geometric boundary information of the knob, including dimensions of a radius of 40 mm and a height of 60 mm. Based on this geometric boundary information, a preset interaction space range is determined through a boundary expansion algorithm. This algorithm expands the original geometric boundary along preset distance thresholds in various directions. These distance thresholds are set between 50 mm and 150 mm depending on the type of operating component. For the aforementioned valve knob, its preset interaction space range is determined to be a cylindrical spatial region with a radius extended to 190 mm and a height extended to 210 mm.
[0074] The spatial distance calculation module is activated to perform three-dimensional spatial distance calculations between the extracted real-time spatial coordinates of the hand and the spatial position of the target operating component. This module calculates the coordinate differences between the hand coordinates and the center point of the operating component in the X, Y, and Z axes respectively. The square roots of the sum of the squares of the differences in the X-axis (-20 mm), Y-axis (+50 mm), and Z-axis (+50 mm) directions yield a three-dimensional straight-line distance of approximately 72 mm. Simultaneously, a spatial inclusion relationship determination program is executed. This program substitutes the real-time spatial coordinates of the hand into the boundary conditions of the preset interactive space range for item-by-item checks. For cylindrical interactive spaces, the program calculates the projected distance between the hand coordinates and the center point of the operating component in the XY plane, determines whether this projected distance is less than the expanded radius threshold of 190 mm, and checks whether the Z-axis value of the hand coordinates falls within the expanded height range.
[0075] Based on the calculated spatial distance values and the determination results of spatial inclusion relationships, the spatial proximity level between the hand pose information and the target operating component is determined. Several proximity levels are preset, including a distance greater than 300 mm for the "far away" state, a distance between 150 mm and 300 mm for the "close" state, a distance between 50 mm and 150 mm for the "near" state, and a hand coordinate falling within a preset interaction space for the "entering" state. When the projected distance of the hand is determined to be 54 mm and the Z-axis coordinate difference is 50 mm, both meeting the boundary conditions, the spatial proximity level is determined to be "entering," triggering the subsequent hand motion feature extraction process.
[0076] After confirming that the hand has entered the preset interactive space, the hand motion feature extractor is activated. This extractor separates feature parameters related to the operation from the complete hand pose information. For knob rotation operations, the extractor focuses on analyzing the relative positional relationship between the tips of the thumb and index finger, calculating the spatial distance between the two fingertips to be 35 mm, and the angle between the line connecting the two fingertips and the horizontal plane to be 15 degrees. The extractor also analyzes the normal vector direction of the palm plane, calculating the normal vector through the plane determined by the wrist point, the base of the middle finger, and the base of the little finger, obtaining an angle of 8 degrees between this normal vector and the axis of the operating component. These geometric feature parameters are combined to form the current hand motion feature vector, which includes multi-dimensional features such as fingertip distance, fingertip connection angle, palm normal vector angle, and overall hand movement speed.
[0077] A predefined operation motion pattern is loaded from the operation guidance data of the current step. This pattern defines the hand motion feature constraints required to complete a specific operation. For the knob rotation operation, the predefined pattern requires that the distance between the thumb and index fingertips be within 30 mm to 50 mm, the angle between the line connecting the two fingertips be within -30 degrees to +30 degrees, the angle between the palm normal vector and the knob axis be less than 20 degrees, and the hand exhibit rotational motion around the knob axis with an angular velocity greater than 10 degrees per second. The extracted hand motion feature vector is compared with each constraint of the predefined pattern. The fingertip distance of 35 mm satisfies the first constraint, the line angle of 15 degrees satisfies the second constraint, and the angle between the normal vectors of 8 degrees satisfies the third constraint. Further analysis of the hand pose changes over multiple consecutive frames reveals that the hand has generated a cumulative rotation angle of 45 degrees around the knob axis, and the rotation direction is consistent with the required direction, thus satisfying the rotational motion constraint.
[0078] The matching result evaluation module statistically analyzes the satisfaction of all constraints. When the proportion of satisfied constraints exceeds a preset matching threshold of 80%, the current hand movement feature is considered to have successfully matched the predefined operation action pattern. After successful matching, the continuity of the operation is further verified, requiring the matching state to maintain at least 10 consecutive data acquisition cycles, corresponding to approximately 330 milliseconds, to avoid misjudging operation completion due to accidental hand posture. Once the continuity verification passes, the completion status of the current operation step is identified as complete, and a completion signal is sent to the operation guidance flow controller, triggering the flow to jump to the next operation step or end the current task. Simultaneously, the timestamp of completion, the final hand pose parameters, and the operation time are recorded for subsequent operation quality evaluation and operation behavior analysis.
[0079] In one optional implementation, the spatial distance between the real-time spatial coordinates of the hand and the spatial position of the target operating component is calculated, and based on the spatial inclusion relationship between the spatial distance and the real-time spatial coordinates of the hand relative to the preset interactive space range, the spatial proximity between the hand pose information and the target operating component is determined, including: Calculate the Euclidean spatial distance between the real-time spatial coordinates of the hand and the spatial position of the target operating component; Obtain spatial boundary description information of the preset interactive space range, wherein the spatial boundary description information defines the geometric shape and range of the preset interactive space range in three-dimensional space; Based on the real-time spatial coordinates of the hand and the spatial boundary description information, the inclusion relationship between the point and the spatial region is determined to determine whether the real-time spatial coordinates of the hand are within the preset interactive space range. Calculate the shortest distance from the real-time spatial coordinates of the hand to the boundary of the preset interactive space range, determine the distance proximity index based on the Euclidean spatial distance, and determine the spatial position relationship index based on the inclusion relationship determination result between the point and the spatial region and the shortest distance; The distance proximity index and the spatial position relationship index are quantified to obtain the spatial proximity between the hand pose information and the target operating component.
[0080] When calculating the spatial distance between the real-time spatial coordinates of the hand and the spatial position of the target operating component, a three-dimensional Euclidean distance calculation method is used. Specifically, assuming the real-time spatial coordinates of the hand are point P, with x-coordinate, y-coordinate, and z-coordinate values of -200 mm, 300 mm, and 450 mm respectively, and the spatial position of the target operating component is point Q, with x-coordinate, y-coordinate, and z-coordinate values of 100 mm, 400 mm, and 500 mm respectively, the calculation process involves calculating the coordinate differences between point P and point Q along the x-axis, y-axis, and z-axis, yielding an x-axis difference of 300 mm, a y-axis difference of 100 mm, and a z-axis difference of 50 mm. These three differences are then squared, resulting in a x-axis difference squared to be 90000, a y-axis difference squared to be 10000, and a z-axis difference squared to be 2500. Summing the three square values yields a total of 102500. Taking the square root of this sum gives an Euclidean distance of approximately 320.16 millimeters.
[0081] When obtaining the spatial boundary description information of a preset interactive space, the geometric shape definition parameters of that interactive space are read from the configuration database. This preset interactive space can be defined as a cuboid shape, and its spatial boundary description information includes the center point coordinates, length, width, and height. For example, if the center point coordinates of a preset interactive space are at the origin (x-coordinate 0 mm, y-coordinate 0 mm, z-coordinate 0 mm), the cuboid has a length of 800 mm, a width of 600 mm, and a height of 700 mm. Based on the center point coordinates and dimensions, the range of the cuboid along the x-axis is determined to be from -400 mm to +400 mm, along the y-axis from -300 mm to +300 mm, and along the z-axis from -350 mm to +350 mm. The spatial boundary description information also includes the plane equation parameters of the six faces, each face described by a normal vector and an offset from the origin, thus fully defining the boundary of the three-dimensional spatial region.
[0082] When determining the inclusion relationship between a point and a spatial region, the real-time spatial coordinates of the hand are compared with the boundary surfaces defined in the spatial boundary description information. For the hand coordinates in the above case (-200 mm, 300 mm, 450 mm), the following checks are performed: The x-coordinate value (-200 mm) is checked to see if it falls within the range of -400 mm to +400 mm (yes); the y-coordinate value (300 mm) is checked to see if it falls within the range of -300 mm to +300 mm (yes); the z-coordinate value (450 mm) is checked to see if it falls within the range of -350 mm to +350 mm (no). Since the z-coordinate value exceeds the upper boundary of the preset interaction space range, the real-time spatial coordinates of the hand are determined not to be within the preset interaction space range. This determination result is stored as a Boolean value, where "within the range" is recorded as true and "outside the range" is recorded as false.
[0083] When calculating the shortest distance from the real-time spatial coordinates of the hand to the boundary of the preset interactive space, the perpendicular distances between the hand coordinate point and each of the six faces of the cuboid are calculated. For the two faces along the x-axis, the distance from the point to the plane with x = -400 mm is 200 mm, and the distance to the plane with x = +400 mm is 600 mm. For the two faces along the y-axis, the distance from the point to the plane with y = -300 mm is 600 mm, and the distance to the plane with y = +300 mm is 0 mm. For the two faces along the z-axis, the distance from the point to the plane with z = -350 mm is 800 mm, and the distance to the plane with z = +350 mm is 100 mm. Since this point is located outside the preset interactive space, the smallest non-zero value among all calculated distance values is selected as the shortest distance. In this case, the shortest distance is 100 mm, representing the distance from the hand coordinate point to the upper boundary of the preset interactive space.
[0084] When determining the proximity index, the Euclidean distance is normalized, and a reference distance threshold of 500 mm is set. The calculated Euclidean distance of 320.16 mm is compared with this threshold. The proximity index is calculated using an inverse proportional relationship, meaning the smaller the distance, the greater the proximity index. Specifically, the reference distance threshold is subtracted from the actual Euclidean distance, resulting in a difference of 179.84 mm. This difference is then divided by the reference distance threshold, yielding a normalized proximity index value of approximately 0.36. This index value ranges from 0 to 1; a value closer to 1 indicates a closer spatial distance between the hand and the target operating part.
[0085] When determining the spatial relationship index, both the inclusion relationship determination result and the shortest distance are considered. Since the real-time spatial coordinates of the hand are not within the preset interaction space, the inclusion relationship determination result is false; therefore, a base weight coefficient of 0.5 is set for this case. Simultaneously, the shortest distance of 100 mm is compared with the set spatial reference distance of 200 mm. Subtracting the shortest distance from the reference distance yields 100 mm, which is then divided by the reference distance to obtain a normalized value of 0.5. Multiplying the base weight coefficient by the normalized distance value gives a spatial relationship index of 0.25. If the hand coordinates are within the preset interaction space, the base weight coefficient is set to 1.0, and the shortest distance is 0; in this case, the spatial relationship index is 1.0.
[0086] When quantifying the distance proximity index and the spatial position relationship index, a weighted fusion method is used to calculate the final spatial proximity. A weighting coefficient of 0.6 is assigned to the distance proximity index, and a weighting coefficient of 0.4 is assigned to the spatial position relationship index. Multiplying the distance proximity index value of 0.36 by its weighting coefficient of 0.6 yields 0.216, and multiplying the spatial position relationship index value of 0.25 by its weighting coefficient of 0.4 yields 0.1. Adding these two values together gives the quantified value of spatial proximity: 0.316. This quantified value ranges from 0 to 1. When the value exceeds the preset interaction trigger threshold of 0.7, it is determined that the hand and the target operating component have reached an interactive state, thus triggering the corresponding interactive operation response. In this case, since the spatial proximity value of 0.316 is lower than the trigger threshold, it is determined that the current state does not meet the interaction trigger condition, and the monitoring state continues to wait for the hand to further approach the target operating component.
[0087] The augmented reality-based real-time guided teaching system for device operation, according to an embodiment of the present invention, includes: The first unit is used to determine the spatial semantic association of the device based on the device spatial information, and to generate a decomposed sequence of operation steps according to the operation task description and the spatial semantic association. The second unit is used to spatially register the operator's perspective information collected in real time with the spatial semantic association, determine the set of visible operation components within the current field of view, and generate viewpoint-adapted guide anchor points based on the spatial occlusion relationship between the target operation component of the current step in the decomposition sequence and the set of visible operation components. The third unit is used to calculate the spatial deviation vector between the operator's hand and the target operating component corresponding to the guide anchor point based on the guide anchor point and the real-time collected hand posture information of the operator, and to generate a real-time operation correction command based on the spatial deviation vector. The real-time operation correction command is superimposed on the display position of the guide anchor point in an augmented reality visualization form. The fourth unit is used to continuously monitor the spatial proximity between the hand pose information and the target operating component corresponding to the guide anchor point. When the operator's hand is detected to have entered the preset interaction space range of the target operating component, the operation completion status is identified. The fifth unit is used to update the current execution position of the decomposed sequence according to the operation completion status, and to redetermine the presentation content of the target operation component and augmented reality guidance information based on the updated current execution position.
[0088] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0089] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0090] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A real-time guided teaching method for device operation based on augmented reality, characterized in that: include: Based on the spatial information of the equipment, the spatial semantic relationship of the equipment is determined, and a decomposed sequence of operation steps is generated according to the description of the operation task and the spatial semantic relationship. Spatial registration is performed between the real-time collected operator's perspective information and the spatial semantic association to determine the set of visible operation components within the current field of view. Based on the spatial occlusion relationship between the target operation component of the current step in the decomposition sequence and the set of visible operation components, a viewpoint-adapted guide anchor point is generated. Based on the guide anchor point and the operator's hand pose information collected in real time, a spatial deviation vector between the operator's hand and the target operating component corresponding to the guide anchor point is calculated. A real-time operation correction command is generated based on the spatial deviation vector, and the real-time operation correction command is superimposed on the display position of the guide anchor point in an augmented reality visualization form. The system continuously monitors the spatial proximity between the hand pose information and the target operating component corresponding to the guide anchor point. When the operator's hand enters the preset interaction space range of the target operating component, the system identifies the operation completion status. The current execution position of the decomposed sequence is updated according to the operation completion status, and the presentation content of the target operation component and augmented reality guidance information is re-determined based on the updated current execution position.
2. The method according to claim 1, characterized in that, Based on the spatial information of the equipment, the spatial semantic association of the equipment is determined. According to the description of the operation task and the spatial semantic association, a decomposed sequence of operation steps is generated, including: Each operable component in the device spatial information is labeled with functional attributes, each operable component is mapped to a predefined operation function category, and the spatial position of each operable component is extracted as spatial coordinate features; Based on the spatial coordinate characteristics, calculate the spatial distance and directional relationships between the operable components, and establish spatial topological connections between the operable components according to the spatial distance and directional relationships; The functional attribute annotations are associated and mapped with the spatial topological connections to form the spatial semantic associations. Parse the target state defined in the operation task description and decompose the target state into multiple sub-target states; Based on the functional attribute annotations in the spatial semantic association, the target operation component corresponding to each sub-target state is matched, and the operation sequence between each target operation component is determined according to the spatial topology connection. Based on the order of operations and the functional attribute labels of each target operation component, a decomposed sequence of operation steps is generated.
3. The method according to claim 1, characterized in that, Spatially register the operator's perspective information collected in real time with the spatial semantic association to determine the set of visible operation components within the current field of view. Based on the spatial occlusion relationship between the target operation component of the current step in the decomposition sequence and the set of visible operation components, generate viewpoint-adapted guidance anchor points, including: Extract the spatial viewpoint coordinates and field of view direction vector from the operator's perspective information, and determine the field of view projection area based on the spatial viewpoint coordinates and field of view direction vector; The spatial positions of each operable component recorded in the spatial semantic association are obtained, and the spatial inclusion relationship between the spatial positions of each operable component and the field projection area is determined. Candidate visible components whose spatial positions fall within the field projection area are then selected. For each candidate visible component, the set of visible operational components within the current field of view is determined based on the spatial relationship between the spatial viewpoint coordinates and the spatial position of the candidate visible component. The target operation component for the current step is obtained from the decomposition sequence. When the target operation component does not belong to the set of visible operation components, the spatial position of the associated visible component that is spatially associated with the target operation component and belongs to the set of visible operation components is determined based on the spatial topological connection relationship in the spatial semantic association relationship, so as to generate the viewpoint-adaptive guide anchor point. When the target operating component belongs to the set of visible operating components, the viewpoint-adaptive guide anchor point is generated based on the spatial position of the target operating component.
4. The method according to claim 3, characterized in that, For each candidate visible component, based on the spatial relationship between the spatial viewpoint coordinates and the spatial position of the candidate visible component, the set of visible operational components within the current field of view is determined, including: Traverse the candidate visible components, and for any candidate visible component, determine the line-of-sight vector pointing from the spatial viewpoint coordinates to the spatial position of the candidate visible component; Obtain the spatial position and geometric shape information of other geometric components in the three-dimensional geometry of the device, excluding the candidate visible components; The spatial intersection operation is performed between the line-of-sight vector and the spatial position and geometric shape information of the other geometric components to determine whether the line-of-sight vector intersects with other geometric components during its propagation from the spatial viewpoint coordinates to the spatial position of the candidate visible component. When the result of the spatial intersection operation indicates that there is no spatial intersection point, the visibility state of the candidate visible component is determined to be visible; otherwise, the visibility state of the candidate visible component is determined to be invisible. The candidate visible components whose visibility status is visible are added to the temporary set of visible components. After traversing all candidate visible components, the temporary set of visible components is determined as the set of operational components visible within the current field of view.
5. The method according to claim 1, characterized in that, Based on the guide anchor point and the operator's hand pose information collected in real time, a spatial deviation vector between the operator's hand and the target operating component corresponding to the guide anchor point is calculated. A real-time operation correction command is generated based on the spatial deviation vector. This real-time operation correction command is overlaid on the display position of the guide anchor point marker in an augmented reality visualization format, including: The spatial position of the target operating component is extracted from the guide anchor point, and the current spatial coordinates and hand posture direction of the operator's hand are extracted from the hand pose information; Calculate the spatial position deviation between the current spatial coordinates of the hand and the spatial position of the target operating component, and calculate the posture angle deviation between the hand posture direction and the predefined operating posture direction required to complete the current step operation. Construct the spatial deviation vector based on the spatial position deviation and the posture angle deviation. Based on the position correction component of the spatial deviation vector, path guidance information indicating the direction and amplitude of hand movement is generated; based on the posture correction component of the spatial deviation vector, posture adjustment information indicating the direction and amplitude of hand rotation is generated; the path guidance information and the posture adjustment information are combined to form the real-time operation correction command. The real-time operation correction command is converted into an augmented reality visualization element, and the augmented reality visualization element is overlaid on the display position of the guide anchor point identifier.
6. The method according to claim 1, characterized in that, Continuously monitor the spatial proximity between the hand pose information and the target operating component corresponding to the guide anchor point. When the operator's hand is detected to have entered the preset interaction space range of the target operating component, identify the operation completion status, including: The system continuously acquires the real-time updated hand pose information and extracts the real-time spatial coordinates of the operator's hand from the hand pose information. The spatial position of the target operating component is obtained from the guide anchor point, and the preset interaction space range of the target operating component is determined based on the spatial position of the target operating component and the geometric boundary information of the target operating component; Calculate the spatial distance between the real-time spatial coordinates of the hand and the spatial position of the target operating component, and determine the spatial proximity between the hand pose information and the target operating component based on the spatial inclusion relationship between the spatial distance and the real-time spatial coordinates of the hand relative to the preset interactive space range; When the real-time spatial coordinates of the hand fall within the preset interactive space range based on the spatial proximity, the hand movement features in the hand pose information are extracted, and the hand movement features are matched with the predefined operation action mode of the current step; Based on the matching result between the hand movement features and the predefined operation action pattern of the current step, the operation completion status is identified.
7. The method according to claim 6, characterized in that, Calculate the spatial distance between the real-time spatial coordinates of the hand and the spatial position of the target operating component, and based on the spatial inclusion relationship between the spatial distance and the real-time spatial coordinates of the hand relative to the preset interactive space range, determine the spatial proximity between the hand pose information and the target operating component, including: Calculate the Euclidean spatial distance between the real-time spatial coordinates of the hand and the spatial position of the target operating component; Obtain spatial boundary description information of the preset interactive space range, wherein the spatial boundary description information defines the geometric shape and range of the preset interactive space range in three-dimensional space; Based on the real-time spatial coordinates of the hand and the spatial boundary description information, the inclusion relationship between the point and the spatial region is determined to determine whether the real-time spatial coordinates of the hand are within the preset interactive space range. Calculate the shortest distance from the real-time spatial coordinates of the hand to the boundary of the preset interactive space range, determine the distance proximity index based on the Euclidean spatial distance, and determine the spatial position relationship index based on the inclusion relationship determination result between the point and the spatial region and the shortest distance; The distance proximity index and the spatial position relationship index are quantified to obtain the spatial proximity between the hand pose information and the target operating component.
8. A real-time guided teaching system for device operation based on augmented reality, used to implement the method as described in any one of claims 1-7, characterized in that, include: The first unit is used to determine the spatial semantic association of the device based on the device spatial information, and to generate a decomposed sequence of operation steps according to the operation task description and the spatial semantic association. The second unit is used to spatially register the operator's perspective information collected in real time with the spatial semantic association, determine the set of visible operation components within the current field of view, and generate viewpoint-adapted guide anchor points based on the spatial occlusion relationship between the target operation component of the current step in the decomposition sequence and the set of visible operation components. The third unit is used to calculate the spatial deviation vector between the operator's hand and the target operating component corresponding to the guide anchor point based on the guide anchor point and the real-time collected hand posture information of the operator, and to generate a real-time operation correction command based on the spatial deviation vector. The real-time operation correction command is superimposed on the display position of the guide anchor point in an augmented reality visualization form. The fourth unit is used to continuously monitor the spatial proximity between the hand pose information and the target operating component corresponding to the guide anchor point. When the operator's hand is detected to have entered the preset interaction space range of the target operating component, the operation completion status is identified. The fifth unit is used to update the current execution position of the decomposed sequence according to the operation completion status, and to redetermine the presentation content of the target operation component and augmented reality guidance information based on the updated current execution position.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.