Industrial part installation auxiliary guiding method based on user intention perception

By introducing user intention perception technology and finite state automatons into the augmented reality-assisted installation guidance system, combining three-dimensional tracking and cloud collaborative real-time interaction, the shortcomings of existing systems in the face of complex assembly scenarios are solved, and an efficient and accurate assembly process is achieved.

CN120236038APending Publication Date: 2025-07-01SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510304959.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing augmented reality assisted installation guidance system is difficult to provide timely feedback and flexible response when facing inunique steps, movements and dynamic changes of assembled objects, resulting in inefficient assembly and assembly.

Method used

The industrial parts installation assisted guidance method based on user intention perception is adopted, through building a camera-projector system, building data sets and performing interactive labeling, using a finite state automaton for assembly state perception and state migration, performing hand-part interaction detection, multi-state object three-dimensional tracking, and real-time augmented reality interaction and presentation based on cloud collaboration.

Benefits of technology

It realizes fully automatic augmented reality assembly state perception, pushes relevant auxiliary information, avoids errors, and improves assembly efficiency and operation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236038A_ABST
    Figure CN120236038A_ABST
Patent Text Reader

Abstract

The invention relates to an industrial part installation auxiliary guiding method based on user intention perception, and belongs to the technical field of part installation assistance. The method comprises the following steps: S1, building a camera-projector system; s2, constructing a data set and performing interactive labeling; s3, carrying out assembly state sensing and state transition based on the finite state automaton; s4, performing hand-part interaction detection based on continuous frames; s5, performing three-dimensional tracking on the multi-state object; s6, real-time augmented reality interaction and presentation based on cloud collaboration; and S7, designing and realizing a prototype system. According to the method and the device, full-automatic augmented reality assembly state sensing is realized, and related augmented reality auxiliary information is pushed, so that errors are effectively avoided, and the overall assembly efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an industrial part installation assistance and guidance method based on user intention perception, and belongs to the technical field of part installation assistance. Background Art

[0002] Augmented reality presents virtual information in an online interactive manner in the real world, thereby enhancing the user's perception and experience. This makes it very suitable for guiding the installation of complex equipment and components in the manufacturing industry. Traditional installation guidance methods, such as drawings or videos, require users to think about the order of installation steps and the relative spatial positions between components by themselves. However, an augmented reality assembly guidance system can help users get started quickly. The system can accurately indicate the next component to be installed and provide corresponding assembly navigation, and display the assembly guidance information in an intuitive and user-friendly manner through augmented reality.

[0003] Early installation guidance methods mainly relied on two-dimensional image markers or required users not to move the installed objects to avoid the need for spatial positioning. These solutions all required users to manually align the objects to be installed with the objects in the virtual space. Some methods chose to let users manually switch the guidance information for different steps, but these methods all shifted the burden of identifying, positioning, and detecting the installation progress to the users.

[0004] Some methods relying on dynamic models adopted object tracking algorithms to locate the installation components. However, as the installation progress advanced, the geometric shape and appearance of the target object would change, which might lead to the failure of spatial registration. To solve this problem, some solutions introduced multi-stage pose estimation and tracking technologies, integrating state graphs and perception intensively, which enhanced the visual perception ability of the assembly system. However, these methods all ignored the unpredictability of user behavior during the assembly process. For example, users might accidentally select the wrong components and try to install them. Therefore, an effective assembly guidance system should provide users with sufficient flexibility to make autonomous decisions and then evaluate these decisions and provide corresponding feedback.

[0005] In summary, augmented reality-assisted installation guidance still faces many challenges, such as the non-uniqueness of steps, the movement of assembled objects, and the dynamic changes of assembled components. A common problem with traditional systems is that they can only provide feedback after users complete a certain step, which often cannot prevent incorrect assembly in time and cannot flexibly handle steps with multiple branches. Therefore, the present invention is proposed. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the present invention provides an industrial part installation auxiliary guidance method based on user intention perception, which realizes fully automatic augmented reality assembly state perception and pushes relevant augmented reality auxiliary information, effectively avoiding errors and improving the overall assembly efficiency.

[0007] Term Explanation:

[0008] SAM: Segment Anything Model is a general image segmentation model proposed by Meta AI. SAM is a pre-trained model that can achieve "all segmentation", aiming to solve the generality problem in image segmentation tasks.

[0009] DeAOT: Decoupled Associating Objects Tracker is a segmentation-based multi-object tracking method, an improved tracking algorithm based on AOT (Associating Objects Tracker), mainly used to track multi-object objects in videos, especially to accurately associate objects at the segmentation level.

[0010] The technical solution of the present invention is as follows:

[0011] An industrial part installation auxiliary guidance method based on user intention perception, the steps are as follows:

[0012] S1: Build a camera-projector system;

[0013] S2: Construct a dataset and perform interactive annotation;

[0014] S3: Perform assembly state perception and state transition based on a finite state automaton;

[0015] S4: Perform hand-part interaction detection based on consecutive frames;

[0016] S5: Three-dimensional tracking of multi-state objects;

[0017] S6: Perform real-time augmented reality interaction and presentation based on cloud collaboration;

[0018] S7: Prototype system design and implementation.

[0019] Preferably, in step S1, the internal parameter information of the camera, including the focal length and optical center of the camera, is pre-acquired by an existing camera calibration method, and matched with the internal parameters of the projector to determine their relative posture relationship. First, a checkerboard image with a grid number of at least 20×20 is selected as the source image, and the image is projected onto a workbench using a projector to ensure that enough ORB feature points can be extracted from the image taken by the camera. Then, the complete projection screen is captured by the camera, and the ORB feature point extraction algorithm is used to extract feature points from the source image and the image captured by the camera, respectively. Through the extracted feature points, a correspondence between the images is established to provide basic data for subsequent affine transformation.

[0020] The ORB (Oriented FAST and Rotated BRIEF) algorithm is rotationally invariant and scale invariant, and can effectively extract significant feature points in the image. These feature points are used for matching through the Flann algorithm to ensure that the correspondence between the source image and the camera image is accurate. Through ORB feature point matching, the system can calculate the affine transformation relationship between the source image and the camera-captured image. The affine transformation can map the coordinate system of the camera image to the projector coordinate system, thereby obtaining the exact position of the camera image in the projector coordinate system. This transformation process ensures the exact correspondence between the camera image and the projector screen, avoids geometric distortion, and improves the accuracy of the alignment of virtual information with the actual scene.

[0021] Initialization: First, the video stream of the workspace is captured by the camera and the image data is extracted. The object detection algorithm (such as YOLOv8) is used to detect the parts in the image to obtain the detection frame and classification identification information. The information is passed to the finite state machine (FSM) for subsequent state evaluation and state migration. The initial state of the system is set to the first zero-indegree node in the state diagram to ensure that the assembly process can proceed stably in the predetermined order.

[0022] Preferably, in step S2, the camera-projector system is mounted on a fixed bracket to capture the working area, and the field of view of the camera and the projector completely covers the projection area of ​​the workbench. During operation, the parts to be inspected are randomly placed in the working area, and the operator manually moves or picks up the parts to ensure that the hands and parts are always within the camera capture range. Through repeated operations, image data covering multiple viewing angles and scene changes are collected. This method can quickly generate original image sequences for real annotation and improve the sample coverage of complex parts.

[0023] SAM is used to accurately segment the first frame of the video stream. SAM determines the segmentation region by simple positive and negative sample clicks and generates an initial mask of the object. Subsequently, the segmentation result of the first frame is transmitted to DeAOT for automatically annotating all objects in subsequent frames. DeAOT performs multi-object segmentation on the objects in the video sequence based on the initial mask, generating an annotated image sequence with a mask overlapping with the actual object contour consistency of over 95%. In this way, the workload of manual annotation is significantly reduced, while ensuring the high precision and consistency of the annotation results.

[0024] To improve the quality of the annotation dataset, a problem frame automatic elimination tool is specially developed. This tool reads the image segmentation results and performs morphological preprocessing on the mask images generated for each frame. Specifically, it removes noise points in the mask through opening operation and fills isolated points in the mask through closing operation. The preprocessed mask images will be converted into the COCO dataset format and corresponding annotation files will be generated.

[0025] Using the Unity rendering environment, a diverse synthetic image dataset is generated. First, the positions and rotation angles of the objects are uniformly sampled within the camera's frustum region. At the same time, certain random perturbations are introduced to the spatial coordinates and rotation angles of the objects to simulate the changes in the real scene. Second, by dynamically changing the background, materials, colors, and lighting conditions, the robustness of the data is further enhanced. Finally, by generating object masks and corresponding images with solid color materials, the annotation of the dataset is completed. This method can automatically generate high-quality annotation data covering multiple angles and environments, effectively supplementing the possible deficiencies in the real dataset.

[0026] Based on the introduction of background randomization and multi-dimensional attribute enhancement during synthetic data generation, the diversity of the dataset is further improved. Specifically, images in the COCO dataset are randomly selected as the background, and at the same time, the materials, colors, and lighting of the objects are randomly sampled. This method can effectively simulate the complex situations that may occur in the real scene, thereby improving the robustness and generalization ability of the object detection model in practical applications.

[0027] Preferably, in step S3, through the decomposition of the assembly task and the modeling of the finite state automaton, the state perception and stability judgment of the assembly process are realized, and the state transition is carried out in combination with the real-time object detection data to complete the intelligent management of the assembly task. Specifically: images of the assembly operation area are captured by the camera, and the image data of each frame analyzed by the object detection algorithm (including the number of detected objects, the position of the bounding box, and the class probability, etc.) are transmitted to the finite state automaton. The finite state automaton comprehensively analyzes the multi-frame continuous data based on the image data and the context situation to judge whether the current assembly state is stable. To ensure the accuracy of the judgment, the finite state automaton calculates the mean value P of the object detection results in every 15 frames of image data.avg , the difference P between the maximum value and the minimum value range and the normalized variance V f , the mathematical expressions are as follows:

[0028]

[0029] P range = P max - P min

[0030]

[0031] where P avg is the mean value of the number of parts, P range is the difference between the maximum value and the minimum value of the number of parts, V f is the normalized variance of the mean value, is the number of parts in the f-th frame, N f is the number of frames within the selected time window;

[0032] When the mean value P avg , the difference P between the maximum value and the minimum value range and the normalized variance V f simultaneously meet the following conditions, it is determined that the current state is stable and state transition is allowed:

[0033] |P avg - round(P avg )| < 0.25

[0034] P range < 1.5

[0035] V f < 1

[0036] The finite state automaton sets the current state to S c , and performs a topological sort according to the state diagram to obtain all possible next states S c+1 , the currently detected state is S d , the finite state automaton compares S d with all S c+1 . If the match is successful, the finite state automaton obtains the type and quantity information of each part through the object detection algorithm, and performs data matching according to the pre-defined state diagram (which details the required quantity and type of parts for each state). When the part data detected in real time completely meets the pre-defined conditions of a certain state, that is, the type and quantity of the parts are consistent with the requirements in the state diagram, this is regarded as a successful match. Then the finite state automaton migrates the current state to S c+1 , and calculates S c+2In all cases, the assembly task is decomposed into multiple stages and modeled using a finite state machine. A state is defined for each stage. The finite state machine determines the current state according to the part position and assembly progress. By defining state transition rules, intelligent perception and adjustment of the task process are achieved. When the conditions are met, the finite state machine automatically migrates to the next state.

[0037] Further preferably, in step S3, in order to improve the accuracy of the state machine for state transition in the assembly task, a dynamic context information analysis mechanism is introduced in the state stability judgment. Specifically, the state judgment strategy is dynamically adjusted by comprehensively analyzing the target detection data of the current frame and several frames before and after it, including the number of parts, location distribution, and category confidence, etc. For example, when the target detection result fluctuates greatly, the number of frames for stability judgment is automatically increased to smooth the fluctuation, ensuring that the state transition is triggered only after the data is stable.

[0038] In the process of acquiring target detection results, in order to avoid the influence of detection errors on the judgment of the state machine, a confidence threshold is set for each detected part in combination with the confidence filtering mechanism. When the detection confidence is lower than the threshold, the detection result of the part will be ignored to prevent erroneous target detection results from interfering with the decision of subsequent state migration. At the same time, in order to enhance the continuity of the detection results, a time-weighted fusion strategy is adopted for the inter-frame detection results to improve the stability of the detection results.

[0039] Preferably, in the state migration step S3, the next state set is prioritized using a topological sorting rule, and the highest priority state is gradually matched according to the detected state;

[0040] Use topological sorting rules to prioritize the possible state sets for the next step. This sorting process is based on the dependencies and prerequisites of each state. The priority is determined mainly based on the number and type of parts required to complete each state, ensuring that assembly is carried out along the most efficient path;

[0041] In actual operation, the system obtains the number and confidence of each part in the current image through the YOLO target detection algorithm, and matches it with the part information required for each state preset by the finite state automaton, giving priority to matching the high-priority state closest to the current detection result. For example, if a state requires a specific number and type of parts that are highly consistent with the detection result, this state will be considered the most likely next state;

[0042] If the initial matching fails, backtrack to the current state and re-match to handle possible detection errors and maintain the consistency of the assembly process. In addition, the historical path of each state migration is recorded so that when an abnormality is detected, it can quickly backtrack to the most recent stable state, thereby ensuring the overall accuracy and reliability of the assembly task.

[0043] Preferably, in order to improve the adaptability of the system to complex assembly tasks, in step S3, the state machine adopts a multi-frame cumulative judgment strategy. That is, among the detection data of consecutive multiple frames, the data with large changes in the number and position distribution of parts are screened and filtered, and only the data with high consistency in the temporal change is retained for stability judgment. This strategy effectively reduces the misjudgment of the state caused by part occlusion or light change.

[0044] Preferably, in the state stability judgment of step S3, the state machine combines the phased characteristics of the assembly task and sets dynamic thresholds for the types and quantities of parts. For example, in certain specific assembly stages, the system will automatically adjust the threshold range of the number of parts according to the predefined assembly steps to adapt to the complex changes in the assembly process. This not only improves the accuracy of state judgment but also reduces unnecessary warning prompts.

[0045] Preferably, in order to improve the real-time performance of state transition, in step S3, the detection results of each frame output by the target detection network are sorted in chronological order. By introducing a frame-by-frame calculation mechanism based on a sliding window, a statistical evaluation is performed on the detection results every 15 frames, avoiding the increase of system burden caused by frequent state judgments. This sliding window strategy can make full use of the temporal characteristics of the detection data while maintaining the judgment efficiency, enhancing the response speed of the state machine.

[0046] Preferably, in step S4, the part and hand information of each frame is obtained through the target detection network in the assembly area, and the interaction state is determined. The intersection over union S is calculated through the detection bounding box of the hand and the detection bounding box of the part IoU ;

[0047]

[0048] Among them, B hand represents the detection area of the hand, B p represents the detection area of part p, AoI represents the intersection area of calculating the two bounding boxes, and AoU represents the union area of calculating the two bounding boxes;

[0049] The specific method is as follows: Initialize the interaction score S p of each part p to 0. When the intersection over union S IoU of the hand and the object is not 0, S p accumulates the IoU value (between 0 and 1) of part p in the current frame;

[0050] S p = S p + S IoU

[0051] When IoU is 0, Score uses a fixed decay function R dDecay until it decays to S p = 0;

[0052] S p = R d (S p )

[0053] The default decay function is inverse proportional function decay:

[0054]

[0055] When the interaction score S p exceeds the set threshold T, it is determined that the hand is in an interaction state with the part; when Score is less than the threshold T, it is determined as a non - interaction state. To avoid the problem of multi - target confusion in interaction determination, the interaction score is calculated for each part p, rather than only calculating the interaction score of the hand, because the hand bounding box may interact with multiple parts simultaneously. When the S p of multiple objects are all greater than the threshold T, the one with the highest interaction score is selected as the current interaction object. Through the above method, the system can detect the interaction state between the hand and the assembled part in real - time and provide interaction guidance for the subsequent assembly process.

[0056] Preferably, in step S4 of part selection and state matching, to ensure that the assembled part selected by the operator is consistent with the current assembly state, in step S3, the state machine monitors the hand - part interaction in real - time. When it detects that the hand interacts with the part and the current state does not match the selected part category, the system will prompt the operator to correct the selection operation through the augmented reality interface to avoid the failure of the assembly task caused by incorrect part selection.

[0057] Preferably, during the interaction detection process, to improve the accuracy of IoU calculation, in step S4, a pixel - based precise bounding box matching method is adopted to avoid misjudgment problems caused by bounding box boundary offset. At the same time, a time - weighted mechanism is introduced to smooth the IoU values in consecutive frames and reduce the jitter of the interaction state caused by detection noise.

[0058] Preferably, to further enhance the reliability of the interaction state, when calculating Score in step S4, the hand movement trajectory information is combined. By analyzing the moving direction and speed of the hand in consecutive frames, it is judged whether it is consistent with the spatial position of the target part, thereby assisting in determining the authenticity of the interaction state.

[0059] Preferably, in step S4, in a multi - target interaction scenario, a priority strategy is adopted. When the Scores of multiple parts simultaneously exceed the threshold T, the target part that is most consistent with the hand trajectory direction is preferentially selected as the interaction object. This strategy can effectively reduce the incorrect interaction determination caused by multi - target interference;

[0060] During the interaction detection process, handle abnormal situations. When the IoU between the hand and the part remains 0 and the Score increases abnormally, trigger the misjudgment correction mechanism to re-initialize the Score of the part, avoiding incorrect interaction states caused by detection errors.

[0061] Preferably, in step S5, first capture a real-time RGB image through a mobile device, compress and store the image, and then send it to the cloud server. The cloud server processes the received image data, uses a 3D tracking algorithm to obtain the current pose of the object. At the same time, when the stable state judgment condition is met, adjust the state of the target 3D model. Subsequently, the pose data and state information generated by the cloud are returned to the mobile device to complete the model matching and state update of the multi-state object, achieving accurate real-time 3D tracking.

[0062] To ensure the accuracy and real-time performance of multi-state tracking, construct a 3D model library covering all possible states, use the depth-first search (DFS) algorithm to permute and combine multi-states, automatically generate 3D models in different states, and optimize the models, including reducing the number of patches and removing redundant texture information, etc., so as to improve the efficiency of the models in calculation and storage. In addition, use a finite state automaton (FSM) to manage the object state, and realize the real-time determination and model switching of the object state through state numbers and transition rules.

[0063] More preferably, in step S5, before the RGB image captured by the mobile device is transmitted to the cloud, use the JPEG encoding method for compression, and at the same time appropriately reduce the image resolution (for example, reduce from 2048×1024 to 512×256) to reduce the bandwidth consumption during transmission and improve the data transmission efficiency, thereby reducing the impact of communication latency on the real-time tracking effect. When generating the 3D model, the system optimizes the 3D model in each state, including limiting the number of model patches to less than 75000 and removing redundant texture maps at the same time, so as to reduce the storage space requirements and reduce the time overhead during model rendering and calculation, and improve the real-time tracking ability of multi-state objects;

[0064] Use the depth-first search (DFS) algorithm to automatically model and sort the multi-states of the object. The system comprehensively explores the possible states of the object under different assembly conditions to ensure that all states are completely covered. Combining with the Meshlab library, the system further processes the generated 3D models, including automatically annotating the part structure numbers and names for quick identification in subsequent model matching.

[0065] In addition, the system manages the multi-state model using a finite state automaton (FSM). Each state is assigned a unique number, and the transition rules between states are defined based on the assembly process of the object. The FSM quickly determines the matching degree between the current state and the target state according to the received pose data and triggers the corresponding model switching operation, thus ensuring the coherence and accuracy of multi-state tracking. To improve the robustness of multi-state tracking, time series analysis technology is introduced in pose estimation to smooth the pose data of consecutive frames, reduce the jitter phenomenon caused by detection errors or noise, improve the stability of state switching, and enable the system to maintain high-precision tracking ability in complex environments.

[0066] Preferably, in step S6, after the cloud server finishes processing, it sends the latest pose and model state information back to the mobile device. At this time, the mobile device application uses augmented reality technology to fuse the updated 3D model with the user's actual environment and display a detailed assembly animation. Based on the tracked model, the animation adds dynamic assembly guidance for the parts being assembled in the current state to dynamically guide the user to perform the assembly work according to the correct steps.

[0067] After receiving the data, the mobile device superimposes a virtual model on the real scene based on the augmented reality platform to achieve precise tracking and dynamic presentation of multi-state objects. Through frame queue management and synchronous coding technology, the system solves the problems of transmission delay and data packet loss. Combining 3D coordinate transformation and occlusion processing, it ensures the seamless fusion and precise alignment of the virtual model and the real scene.

[0068] Preferably, in step S6, during the augmented reality presentation process, the conversion of 3D coordinates and the rendering display of the virtual model are realized based on the Unity platform and OpenGL technology. To improve the alignment accuracy between the model and the real scene, the X, Y, and Z axis data of the 3D coordinates are filtered, and the pose change of the model in consecutive frames is optimized through an interpolation algorithm to ensure the smoothness and stability of the display process.

[0069] Preferably, in step S6, to solve the occlusion problem of the fusion between the virtual model and the real scene, the system adopts a dynamic occlusion processing technology. The specific method is to generate an occlusion area based on the depth information of the virtual model and achieve different degrees of occlusion effects by adjusting the transparency of the occlusion area. In the unoccluded state, the virtual model is displayed transparently, only presenting the model edges; in the partially occluded state, the transparency is reduced proportionally to ensure the natural superposition of the virtual and real objects; in the fully occluded state, the virtual model presents a low transparency and is fully integrated with the real object.

[0070] Preferably, in step S6, in combination with the frame synchronization coding technology, the state switching of the virtual model is corrected in real time with the pose change of the real object. During the model switching process, each frame of data is marked with a frame number and matched with the pose information to ensure that the virtual model accurately corresponds to the actual state of the current object when the state is switched, thereby avoiding the problem of inconsistency between the virtual and the real caused by the model switching delay;

[0071] To implement the frame synchronization coding technology, a frame number is marked in each frame of data, and the pose information calculated in the cloud also contains the corresponding frame number. The specific matching process is as follows:

[0072] Binding of frame number and pose information: When the mobile device encodes and sends an image, each image frame is assigned a unique frame number, and the frame number is encoded into the image data packet and sent to the cloud server. After the cloud server calculates the pose corresponding to the frame of the image, it will bind the calculated pose with the frame number of the image frame and return them to the mobile device together;

[0073] Receiving and matching: After receiving the data packet containing the pose and the frame number, the mobile device will search for the image frame that matches the returned frame number in the buffer queue. By comparing the frame number stored in the buffer queue with the frame number returned by the cloud, the mobile device accurately finds the image frame corresponding to the current pose information.

[0074] Through this frame synchronization coding and pose matching method, the system can ensure that even in the case of network delay or unstable data transmission, the accurate and real-time update of the virtual model can be achieved, thereby greatly improving the user experience and interaction quality of the augmented reality application.

[0075] Preferably, in step S6, for the tracking of multi-state objects in complex scenarios, by fusing the real-time transmitted object image and the historical pose trajectory information, the rendering position and pose of the virtual model are dynamically adjusted. In a scenario with a large environmental change, the stability and robustness of the virtual model in a dynamic environment are enhanced by increasing the length of the frame queue to smooth the pose estimation result.

[0076] Preferably, in step S7, an augmented reality assembly guidance prototype system is constructed to implement the assembly task guidance function of multi-module collaboration. Based on the multi-state model management and real-time state perception technology, this system completes the full-process coverage from part detection, state judgment to dynamic guidance during the assembly process. During the assembly task, the user obtains real-time guidance information through a mobile device, including the superimposed display of virtual models and operation prompts, dynamically adjusts the guidance content according to the current part state, and combines the data collected by the vision sensor to achieve precise planning and real-time feedback of the assembly steps. In addition, the system introduces the virtual-real fusion technology to enable seamless docking between the virtual model and the actual scene, providing intuitive and clear assembly guidance for users, significantly improving the operation efficiency and accuracy. Through this system, assembly workers can quickly understand and complete the required operations in complex tasks, effectively improving the assembly quality and the reliability of the overall process.

[0077] Further preferably, in step S7, in the assembly application example, a scheme combining multi-state part models and augmented reality guidance is adopted. The system generates corresponding virtual models according to the geometric features and assembly requirements of each part and dynamically displays them in the actual scene. In the part picking link, the position of the target part is automatically identified and marked through a vision algorithm, providing clear pick-and-place guidance; in the assembly link, the connection path and operation prompts of the virtual model are dynamically displayed to ensure the accuracy of the assembly action.

[0078] The mobile augmented reality assembly guidance module further enhances the interaction experience between the user and the system. It real-time displays the operation steps and assembly status through color identification and animation effects. For example, the correct parts are highlighted in green, the wrong operation prompts are red warnings, and the uncompleted steps are shown in gray, ensuring that users can quickly understand and execute the assembly operations, thus significantly improving the assembly efficiency and quality.

[0079] The beneficial effects of the present invention are as follows:

[0080] 1. Compared with the existing assembly guidance method based on static rules, the present invention uses a finite state automaton to dynamically manage the assembly process and solves the problem that the same type of method cannot adapt to the changes in complex assembly scenarios by real-time evaluating the state stability, significantly improving the assembly efficiency and operation accuracy.

[0081] 2. The present invention uses augmented reality technology to accurately superimpose the virtual model on the actual scene. Compared with the existing augmented reality method based on static guidance, through three-dimensional coordinate transformation and occlusion processing technology, it realizes the seamless fusion of the virtual model and the real scene, solving the problem of large alignment errors between the model and the actual scene in the same type of method.

[0082] 3. In the multi-state object tracking of the present invention, by combining the mobile terminal and cloud collaborative processing technology, the problems of inaccurate real-time tracking caused by data transmission delay and loss in similar methods are solved through frame synchronization encoding and model optimization strategies, ensuring high-precision three-dimensional pose estimation in a dynamic environment.

[0083] 4. The automated data annotation and synthetic data generation platform designed by the present invention, compared with similar methods that traditionally rely on manual annotation, significantly reduces the annotation cost and improves the data quality by introducing the SAM and DeAOT algorithms, solving the problems of low annotation efficiency and insufficient sample diversity in existing methods.

[0084] 5. The augmented reality assembly guidance prototype system based on multi-module collaboration of the present invention solves the problems of single guidance information and insufficient flexibility for assembly tasks in similar methods through full-process coverage from part detection, status judgment to dynamic assembly guidance, providing more efficient and reliable auxiliary support for complex assembly scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] Figure 1 is the usage framework of the present invention. The camera, projector and mobile device serve as visual sensors and display devices. The assembly state machine is the core that controls the entire assembly guidance, including user interaction detection, part status maintenance and assembly status maintenance. It is deployed on a standard desktop computer together with multi-state object tracking and communicates with the mobile device through a local area network.

[0086] Figure 2 is the schematic diagram of the guidance process of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0087] The present invention will be further described below by way of examples in conjunction with the drawings, but is not limited thereto.

[0088] Example 1:

[0089] This example provides an industrial part installation auxiliary guidance method based on user intention perception, and the steps are as follows:

[0090] S1: Build a camera-projector system;

[0091] The internal parameter information of the camera, including the focal length and optical center of the camera, is obtained in advance through the existing camera calibration method, and matched with the internal parameters of the projector to determine their relative posture relationship. First, a checkerboard image with a grid number of at least 20×20 is selected as the source image, and the image is projected onto the workbench using a projector to ensure that enough ORB feature points can be extracted from the image taken by the camera. Then, the complete projection screen is captured by the camera, and the feature points of the source image and the image captured by the camera are extracted using the ORB feature point extraction algorithm. Through the extracted feature points, the correspondence between the images is established to provide basic data for subsequent affine transformation.

[0092] The ORB (Oriented FAST and Rotated BRIEF) algorithm is rotationally invariant and scale invariant, and can effectively extract significant feature points in the image. These feature points are used for matching through the Flann algorithm to ensure that the correspondence between the source image and the camera image is accurate. Through ORB feature point matching, the system can calculate the affine transformation relationship between the source image and the camera-captured image. The affine transformation can map the coordinate system of the camera image to the projector coordinate system, thereby obtaining the exact position of the camera image in the projector coordinate system. This transformation process ensures the exact correspondence between the camera image and the projector screen, avoids geometric distortion, and improves the accuracy of the alignment of virtual information with the actual scene.

[0093] Initialization: First, the video stream of the workspace is captured by the camera and the image data is extracted. The object detection algorithm (such as YOLOv8) is used to detect the parts in the image to obtain the detection frame and classification identification information. The information is passed to the finite state machine (FSM) for subsequent state evaluation and state migration. The initial state of the system is set to the first zero-indegree node in the state diagram to ensure that the assembly process can proceed stably in the predetermined order.

[0094] S2: Build the dataset and perform interactive annotation;

[0095] The camera-projector system is installed on a fixed bracket to capture the working area. The field of view of the camera and projector completely covers the projection area of ​​the workbench. During operation, the parts to be inspected are randomly placed in the working area, and the operator manually moves or picks up the parts to ensure that the hands and parts are always within the camera capture range. Through repeated operations, image data covering multiple perspectives and scene changes are collected. This method can quickly generate original image sequences for real annotation and improve the sample coverage of complex parts.

[0096] SAM is used to accurately segment the first frame of the video stream. SAM determines the segmentation region by simple positive and negative sample clicks and generates an initial mask of the object. Subsequently, the segmentation result of the first frame is transmitted to DeAOT for automatically annotating all objects in the subsequent frames. DeAOT performs multi-object segmentation on the objects in the video sequence based on the initial mask and generates an annotation image sequence with a mask overlapping consistency of more than 95% with the actual object contour. In this way, the workload of manual annotation is significantly reduced, while ensuring the high accuracy and consistency of the annotation results.

[0097] To improve the quality of the annotation dataset, a problem frame automatic elimination tool is specially developed. This tool reads the image segmentation result and performs morphological preprocessing on the mask image generated for each frame. Specifically, the noise in the mask is removed through opening operation, and the isolated points in the mask are filled through closing operation. The preprocessed mask image will be converted into the COCO dataset format and the corresponding annotation file will be generated.

[0098] Using the Unity rendering environment, a diverse synthetic image dataset is generated. First, the positions and rotation angles of the objects are uniformly sampled within the camera's frustum region. At the same time, certain random perturbations are introduced to the spatial coordinates and rotation angles of the objects to simulate the changes in the real scene. Secondly, by dynamically changing the background, materials, colors, and lighting conditions, the robustness of the data is further enhanced. Finally, by generating the object masks and corresponding images with solid-color materials, the annotation of the dataset is completed. This method can automatically generate high-quality annotation data covering multiple angles and environments, effectively supplementing the possible deficiencies in the real dataset;

[0099] Based on the introduction of background randomization and multi-dimensional attribute enhancement during the generation of synthetic data, the diversity of the dataset is further improved. Specifically, images in the COCO dataset are randomly selected as the background, and the materials, colors, and lighting of the objects are randomly sampled. This method can effectively simulate the complex situations that may occur in the real scene, thereby improving the robustness and generalization ability of the object detection model in practical applications.

[0100] S3: Assembly state perception and state transition are carried out based on the finite state automaton;

[0101] Through the decomposition of assembly tasks and the modeling of finite state automata, the state perception and stability judgment of the assembly process are realized, and state migration is carried out in combination with real-time object detection data to complete the intelligent management of assembly tasks. Specifically: capture the images of the assembly operation area through a camera, and transmit the image data of each frame analyzed by the object detection algorithm (including the number of detected objects, the position of the bounding box, and the category probability, etc.) to the finite state automaton. The finite state automaton comprehensively analyzes the multi-frame continuous data based on the image data and the context situation to judge whether the current assembly state is stable. To ensure the accuracy of the judgment, the finite state automaton calculates the mean value P of the object detection results in every 15-frame image data avg , the difference between the maximum value and the minimum value P range and the normalized variance V f , and the mathematical expressions are as follows:

[0102]

[0103] P range = P max - P min

[0104]

[0105] Among them, P avg is the mean value of the number of parts, P range is the difference between the maximum value and the minimum value of the number of parts, V f is the normalized variance of the mean value, is the number of parts in the f-th frame, N f is the number of frames within the selected time window;

[0106] When the mean value P avg , the difference between the maximum value and the minimum value P range and the normalized variance V f simultaneously meet the following conditions, it is judged that the current state is stable and state switching is allowed:

[0107] |P avg - round(P avg )| < 0.25

[0108] P range < 1.5

[0109] V f < 1

[0110] The finite state automaton sets the current state to S c , and performs topological sorting according to the state diagram to obtain all possible next states S c+1 , the currently detected state is S d , the finite state automaton will Sd Match with all S c+1 If the match is successful, the finite state automaton obtains the type and quantity information of each part through the object detection algorithm. According to the pre-defined state diagram (which details the required quantity and type of parts for each state), data matching is performed. When the part data detected in real-time fully meets the predefined conditions of a certain state, that is, the type and quantity of the parts are consistent with the requirements in the state diagram, this is regarded as a successful match. Then the finite state automaton migrates the current state to S c+1 , and calculates all cases of S c+2 Decompose the assembly task into multiple stages and model it using a finite state machine. Define a state for each stage. The finite state automaton determines the current state based on the part position and assembly progress. By defining state transition rules, intelligent perception and adjustment of the task process are achieved. When the conditions are met, the finite state automaton automatically migrates to the next state.

[0111] To improve the accuracy of state transitions in the assembly task by the state machine, a dynamic context information analysis mechanism is introduced during state stability judgment. Specifically, by comprehensively analyzing the object detection data of the current frame and several frames before and after it, including part quantity, position distribution, and category confidence, etc., the state judgment strategy is dynamically adjusted. For example, when the object detection results fluctuate greatly, the number of frames for stability judgment is automatically increased to smooth the fluctuations, ensuring that state transitions are only triggered after the data stabilizes.

[0112] During the process of obtaining object detection results, to avoid the influence of detection errors on the state machine judgment, a confidence filtering mechanism is combined. A confidence threshold is set for each detected part. When the detection confidence is lower than the threshold, the detection result of this part will be ignored, preventing incorrect object detection results from interfering with the decision-making of subsequent state transitions. At the same time, to enhance the continuity of detection results, a time-weighted fusion strategy is adopted for inter-frame detection results to improve the stability of detection results.

[0113] In the state transition step, use the topological sorting rule to prioritize the set of next states and gradually match the highest-priority state according to the detected state;

[0114] Use the topological sorting rule to prioritize the set of possible next states. This sorting process is based on the dependency relationships and preconditions of each state. The determination of priority mainly depends on the quantity and type of parts required for each state, ensuring that the assembly proceeds along the most efficient path;

[0115] In actual operation, the system obtains the quantity and confidence of each part in the current image through the YOLO object detection algorithm, and matches it with the part information required for each state preset by the finite state automaton. It preferentially matches the high-priority state that is closest to the current detection result. For example, if a state requires a specific quantity and type of parts that highly match the detection result, this state will be considered the most likely next state;

[0116] If the initial match fails, it backtracks to the current state and rematches to handle possible detection errors and maintain the coherence of the assembly process. In addition, the historical path of each state transition is recorded so that in case of an abnormal situation, it can quickly backtrack to the nearest stable state, thus ensuring the overall accuracy and reliability of the assembly task.

[0117] S4: Perform hand-part interaction detection based on consecutive frames;

[0118] Obtain the part and hand information of each frame through the object detection network in the assembly area, and determine the interaction state. Calculate the intersection over union S between the detection bounding box of the hand and the detection bounding box of the part IoU ;

[0119]

[0120] Among them, B hand represents the detection area of the hand, B p represents the detection area of part p, AoI represents the intersection area of calculating the two bounding boxes, and AoU represents the union area of calculating the two bounding boxes;

[0121] The specific method is: Initialize the interaction score S p of each part p to 0. When the intersection over union S IoU of the hand and the object is not 0, S p accumulates the IoU value (between 0 and 1) of part p in the current frame;

[0122] S p = S p + S IoU

[0123] When IoU is 0, Score uses a fixed decay function R d to decay until it decays to S p = 0;

[0124] S p = R d (S p )

[0125] The default decay function is inverse proportional function decay:

[0126]

[0127] When the interaction score S p exceeds the set threshold T, it is determined that the hand is in an interaction state with the part; when Score is less than the threshold T, it is determined as a non-interaction state. To avoid the problem of multi-target confusion in interaction determination, the interaction score is calculated for each part p, rather than only calculating the interaction score of the hand, because the hand bounding box may interact with multiple parts simultaneously. When the S p of multiple objects are all greater than the threshold T, the one with the highest interaction score is selected as the current interaction object. Through the above method, the system can detect the interaction state between the hand and the assembled parts in real time and provide interaction guidance for the subsequent assembly process.

[0128] In a multi-target interaction scenario, a priority strategy is adopted. When the Scores of multiple parts simultaneously exceed the threshold T, the target part that is most consistent with the hand trajectory direction is preferentially selected as the interaction object. This strategy can effectively reduce the misjudgment of interaction caused by multi-target interference;

[0129] During the interaction detection process, abnormal situations are handled. When the IoU between the hand and the part remains 0 and the Score increases abnormally, the misjudgment correction mechanism is triggered to re-initialize the Score of the part to avoid incorrect interaction states caused by detection errors.

[0130] S5: Three-dimensional tracking of multi-state objects;

[0131] First, a real-time RGB image is captured by a mobile device, and after compressing and storing the image, it is sent to a cloud server. The cloud server processes the received image data and uses a three-dimensional tracking algorithm to obtain the current pose of the object. At the same time, when the stable state judgment condition is met, the state of the target three-dimensional model is adjusted. Subsequently, the pose data and state information generated by the cloud are returned to the mobile device to complete the model matching and state update of the multi-state object, achieving accurate real-time three-dimensional tracking.

[0132] Before the RGB image captured by the mobile device is transmitted to the cloud, the JPEG encoding method is used for compression, and at the same time, the image resolution is appropriately reduced (for example, reduced from 2048×1024 to 512×256) to reduce the bandwidth consumption during transmission and improve the data transmission efficiency, thereby reducing the impact of communication latency on the real-time tracking effect. When generating the three-dimensional model, the system optimizes the three-dimensional model in each state, including limiting the number of model patches to less than 75000 and removing redundant texture maps, thereby reducing the storage space requirement and reducing the time overhead during model rendering and calculation to improve the real-time tracking ability of multi-state objects;

[0133] Use the Depth-First Search (DFS) algorithm to automatically model and sort the multi-states of an object. The system comprehensively explores the possible states of the object under different assembly conditions to ensure that all states are completely covered. Combining with the Meshlab library, the system further processes the generated 3D model, including automatically annotating the part structure numbers and names for quick identification in subsequent model matching.

[0134] In addition, the system uses a Finite State Automaton (FSM) to manage the multi-state model. Each state is assigned a unique number, and the transition rules between states are defined based on the assembly process of the object. The FSM quickly determines the matching degree between the current state and the target state according to the received pose data and triggers the corresponding model switching operation, thus ensuring the coherence and accuracy of multi-state tracking. To improve the robustness of multi-state tracking, time series analysis technology is introduced in pose estimation to smooth the pose data of consecutive frames, reduce the jitter phenomenon caused by detection errors or noise, improve the stability of state switching, and enable the system to maintain high-precision tracking ability in complex environments.

[0135] S6: Cloud collaborative real-time augmented reality interaction and presentation;

[0136] After the cloud server finishes processing, it sends the latest pose and model state information back to the mobile device. At this time, the mobile device application fuses the updated 3D model with the user's actual environment through augmented reality technology and displays a detailed assembly animation. Based on the tracked model, the animation adds dynamic assembly guidance for the parts being assembled in the current state to dynamically guide the user to perform the assembly work according to the correct steps.

[0137] After receiving the data, the mobile device superimposes the virtual model on the real scene based on the augmented reality platform to achieve precise tracking and dynamic presentation of the multi-state object. Through frame queue management and synchronous coding technology, the system solves the problems of transmission delay and data packet loss. Combining 3D coordinate transformation and occlusion processing, it ensures the seamless fusion and precise alignment of the virtual model and the real scene.

[0138] During the augmented reality presentation process, the 3D coordinate transformation and the rendering and display of the virtual model are realized based on the Unity platform and OpenGL technology. To improve the alignment accuracy between the model and the real scene, the X, Y, and Z axis data of the 3D coordinates are filtered, and the pose change of the model in consecutive frames is optimized through interpolation algorithms to ensure the smoothness and stability of the display process.

[0139] Combined with frame synchronization coding technology, the state switching of the virtual model is corrected in real time with the pose change of the real object. During the model switching process, each frame of data is marked with a frame number and matched with the pose information to ensure that the virtual model accurately corresponds to the actual state of the current object when the state is switched, thus avoiding the problem of inconsistency between virtual and real caused by model switching delay;

[0140] To implement frame synchronization coding technology, a frame number mark is added to each frame of data, and the pose information calculated in the cloud also contains the corresponding frame number. The specific matching process is as follows:

[0141] Binding of frame number mark and pose information: When the mobile device encodes and sends an image, each image frame is assigned a unique frame number, and this frame number is encoded into the image data packet and sent to the cloud server. After the cloud server calculates the pose corresponding to this frame of image, it will bind the calculated pose with the frame number of this image frame and return them to the mobile device together;

[0142] Receiving and matching: After the mobile device receives the data packet containing the pose and frame number, it will search for the image frame that matches the returned frame number in the buffer queue. By comparing the frame number stored in the buffer queue with the frame number returned by the cloud, the mobile device accurately finds the image frame corresponding to the current pose information.

[0143] S7: Prototype system design and implementation.

[0144] Build an augmented reality assembly guidance prototype system to realize the assembly task guidance function of multi-module collaboration. Based on multi-state model management and real-time state perception technology, this system completes the full process coverage from part detection, state judgment to dynamic guidance during the assembly process. In the assembly task, the user obtains real-time guidance information through the mobile device, including the superimposed display and operation prompts of the virtual model, dynamically adjusts the guidance content according to the current part state, and combines the data collected by the visual sensor to achieve accurate planning and real-time feedback of the assembly steps. In addition, the system introduces virtual-real fusion technology to enable seamless docking of the virtual model and the actual scene, providing intuitive and clear assembly guidance for users, greatly improving the operation efficiency and accuracy. Through this system, assembly workers can quickly understand and complete the required operations in complex tasks, effectively improving the assembly quality and the reliability of the overall process.

[0145] In the assembly application example, a scheme combining multi-state part models and augmented reality guidance is adopted. The system generates corresponding virtual models according to the geometric features and assembly requirements of each part and dynamically displays them in the actual scene. In the part picking link, the position of the target part is automatically recognized and marked through a visual algorithm, providing clear pick-and-place guidance; in the assembly link, the connection path and operation prompts of the virtual model are dynamically displayed to ensure the accuracy of the assembly action.

[0146] The mobile augmented reality assembly guidance module further enhances the user-system interaction experience. It uses color identification and animation effects to display operation steps and assembly status in real time. For example, correct parts are highlighted in green, incorrect operation prompts are shown as red warnings, and unfinished steps are represented in gray, ensuring that users can quickly understand and execute assembly operations, thus significantly improving assembly efficiency and quality.

[0147] To verify the effectiveness of the industrial part installation assistance and guidance method based on user intention perception proposed in this embodiment, a set of comparative experiments were designed. The existing industrial augmented reality guidance system (such as Vuforia) was selected as the comparison object, and the two methods were evaluated in multiple typical industrial assembly task scenarios. The experimental indicators include assembly efficiency, assembly accuracy, and adaptability to complex assembly scenarios. In this embodiment, a control system based on the Vuforia framework was built on a mobile device to conduct three-dimensional registration performance tests. The method proposed in this embodiment and Vuforia both run on the same mobile platform. Among them, Vuforia uses the Model Target Generator to generate three-dimensional model targets. Since the two instances used in the experiments of this embodiment are objects with insufficient texture, it is set to the LOW_FEATURE mode in Vuforia. The image captured by the mobile device camera is streamed to OBS and converted into a virtual camera. While testing the effect of Vuforia, the method of this embodiment is used to perform tracking processing on the same video stream. The above experiments are repeated three times for each scenario. The specific results are as follows. The system of this embodiment directly uses the given three-dimensional model and ensures that other configuration parameters of the two systems are the same. The experiment follows the following specific steps: Set the same initial pose, and under the same image sequence, compare and analyze the tracking effects of the two methods on the same object in different states and perspectives.

[0148]

[0149] It can be seen from the experimental data that the method of the present invention is significantly superior to the existing Vuforia system in terms of tracking success rate. Especially in complex lighting and multi-part interference scenarios, it still maintains high stability. Vuforia shows obvious drift phenomena in the above scenarios, and the success rate of tracking target parts is low, unable to meet the actual industrial needs.

[0150] To verify the actual effect of the augmented reality assembly guidance system proposed by this method, a user study was designed and implemented. The specific design is as follows: 16 college students were recruited as experimental participants. All participants had no professional background in augmented reality technology to ensure the universality of the experimental results. The task objective of the experimental design was to compare the augmented reality assembly system with the traditional assembly guidance system (two-dimensional drawings) to evaluate the working efficiency and user experience of the augmented reality system.

[0151] The measurement indicators are divided into two categories:

[0152] Efficiency indicators: Record the assembly time and compare the completion times of the experimental group (augmented reality system) and the control group (two-dimensional drawings).

[0153] Subjective indicators: A detailed questionnaire survey was designed, including key dimensions such as the ease of use of the system, user satisfaction, and error rate.

[0154] The questionnaire used a 5-level Likert scale to quantify the participants' evaluations of the system usage. The questionnaire content covered feedback on the system's interface design, function implementation, and interaction methods. The collected data showed that compared with the traditional two-dimensional drawings, the augmented reality assembly system proposed in this embodiment reduced the assembly time by an average of 34% and the user error rate by 47%. The statistical results of the questionnaire survey showed that the participants generally believed that the system interface was clear and the operation was simple, and the "intuition" and "flexibility" of the system received the highest evaluations.

[0155] Table 6-1 Comparative results of objective indicators Tab.6-1 Comparative results of objective indicators

[0156]

[0157] Table 6-2 Comparative results of subjective indicators Tab.6-2 Comparative results of subjective indicators

[0158]

[0159]

[0160] Example 2:

[0161] This embodiment provides an industrial part installation assistance and guidance method based on user intention perception. As described in Embodiment 1, the difference is that an auxiliary depth camera is added during the initialization process to capture the three-dimensional depth information of the assembly area, and the RGB data and depth data are aligned through multi-sensor fusion technology to ensure the robustness of target detection and state perception. During the calibration process, a dynamic calibration method is used in combination with a high-precision calibration board to dynamically adjust the relative poses of the camera and the projector, so as to adapt to the lighting and field of view changes in different assembly scenarios. The system newly adds an ambient light intensity sensor to monitor the light conditions in the working area in real time and automatically adjust the exposure parameters of the camera and the brightness of the projector, so that the system can still maintain good detection and guidance capabilities even under complex lighting conditions.

[0162] In terms of dataset annotation, this embodiment optimizes the method for generating synthetic data. In the Unity environment, in addition to conventional augmentation operations such as randomizing the background, materials, and lighting, typical occlusion situations during the part assembly process are simulated, such as the hand occluding the part, the tool covering part of the part, etc., so as to improve the adaptability of the target detection model to complex assembly scenarios. In addition, multi-view samples are added in the real data collection, including the position changes of the parts under different assembly sequences and the state records at the completion stage of the assembly. In data annotation, the SAM model and the DeAOT algorithm are combined to optimize the segmentation strategy for transparent parts or parts with reflective surfaces, generating high-quality mask annotations and significantly improving the quality of the dataset.

[0163] In the state machine-based assembly process management, this embodiment newly adds a state backtracking mechanism. When an abnormal operation (such as incorrect part position or incorrect assembly order) is detected, the system will backtrack to the nearest stable state and re-prompt the user for the correct operation of the current step. At the same time, the state machine combines a dynamic weight adjustment mechanism to dynamically adjust the trigger conditions for state transition according to the confidence of part detection and the reliability of user operations, thereby improving the accuracy and real-time performance of assembly guidance.

[0164] In terms of hand-part interaction detection and state perception, this embodiment adds a hand movement trajectory analysis function, and combines time series analysis technology to track the interaction state between the hand and the part in real time. When the system detects that the hand trajectory is inconsistent with the movement direction of the target part, an error prompt will be issued through the augmented reality interface to guide the user to adjust the operation path to ensure the accuracy and coherence of the assembly operation.

[0165] Embodiment 3:

[0166] This embodiment provides an industrial part installation assistance and guidance method based on user intention perception. As described in Embodiment 2, the difference is that this embodiment improves the three-dimensional tracking and management of multi-state objects, adds a depth estimation module based on binocular stereo vision, and calculates the three-dimensional position and pose information of objects in the assembly area in real time through binocular cameras. Combining with the pose records in the state machine, precise tracking of multi-state objects is achieved. Compared with monocular RGB images, binocular stereo vision can significantly improve the depth measurement accuracy of target objects, especially in scenarios where multiple parts overlap or are highly complex.

[0167] On the basis of the finite state automaton, a task priority scheduling function is newly added. During the planning process of the assembly task, the priority of the task is dynamically adjusted in combination with the importance of the parts, the user's assembly progress, and possible abnormal situations. For example, when an assembly error of a key part is detected, the system will first prompt the user to correct the error instead of continuing the current assembly step, thereby ensuring the correctness of the overall assembly process. In addition, the state machine introduces a time-weighted smoothing strategy in the state stability judgment, performs weighted processing on multi-frame detection data, and effectively reduces the state judgment fluctuations caused by changes in light or angle.

[0168] In the augmented reality assembly guidance, this embodiment newly adds a gesture recognition module. The user can control the system prompt content through simple gesture operations, such as zooming in on the virtual model, rotating the viewing angle, or pausing the current operation prompt. The display content of the augmented reality interface is further optimized, and the virtual model and the actual scene are more naturally integrated by combining the transparency gradient technology, making the alignment effect between the virtual model and the physical part more accurate.

[0169] Embodiment 4:

[0170] This embodiment provides an industrial part installation assistance and guidance method based on user intention perception. As described in Embodiment 3, the difference is that this embodiment further optimizes the state transition rules of the state machine. In the state judgment process, multi-modal perception technology is combined, and the image features of part detection, the mechanical characteristics of user operations, and the real-time data of assembly tools are comprehensively analyzed to improve the accuracy of state transition. For example, when the system detects that the user is operating with a specific tool (such as a screwdriver or welding equipment), it will automatically adjust the state transition conditions of the current step to ensure that the assembly task conforms to the actual operation process.

[0171] In terms of three-dimensional tracking and management, this embodiment newly adds a real-time model update mechanism. By capturing the surface change data of the parts in real time (such as slight deformation or pose offset caused by installation), the geometric shape and state information of the three-dimensional model are dynamically adjusted, so that the display content of the virtual model always remains consistent with the actual parts. This mechanism effectively solves the adaptability problem of traditional static models in dynamic assembly scenarios.

[0172] In augmented reality guidance, a new voice prompt function is added. Users can obtain detailed operation instructions for the current step or jump to the specified assembly stage through voice interaction. In addition, a multi-user collaboration mode is added. In the same assembly task, multiple users are supported to view the assembly guidance information in real time through their respective mobile devices and complete complex assembly tasks collaboratively by sharing virtual models.

[0173] Embodiment 5:

[0174] This embodiment provides an industrial part installation assistance and guidance method based on user intention perception. As described in Embodiment 4, the difference is that this embodiment further improves the state perception and interaction guidance capabilities in complex scenarios. During the planning process of the assembly task, a path prediction algorithm is introduced to predict possible error paths based on the user's historical operation data and give a warning prompt before the error occurs. For example, when the system detects that the user's operation trajectory deviates from the target part, it will dynamically display the correction path through the augmented reality interface to reduce the occurrence of assembly errors.

[0175] In terms of hand-part interaction detection, this embodiment combines a machine learning model to perform pattern recognition on the user's operation behavior and automatically distinguish normal assembly actions from abnormal interference actions. For example, when it detects that the contact force between the user's hand and the part is abnormal, the system will immediately pause the state transition of the current step and prompt the user to check whether the part is placed correctly. This function is especially suitable for assembly scenarios with high requirements for operation accuracy, such as precision machinery assembly or electronic component installation.

[0176] In the augmented reality interface display, this embodiment adds a dynamic occlusion optimization function. By analyzing the occlusion relationship between the virtual model and the real scene in real time, the system dynamically adjusts the transparency, brightness, and texture features of the virtual model, enabling users to clearly distinguish the current part from the virtual prompt content. In addition, the system automatically adjusts the projection angle of the virtual model in combination with the scene data of the depth camera to perfectly match the user's perspective, thus significantly improving the intuitiveness of the assembly guidance and the user experience.

[0177] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An industrial parts installation auxiliary guidance method based on user intention perception, characterized in that: Here are the steps: S1: Build the camera-projector system; S2: Build the dataset and perform interactive annotation; S3: Assembly state perception and state migration based on finite state automata; S4: Hand-part interaction detection based on continuous frames; S5: 3D tracking of multi-state objects; S6: Real-time augmented reality interaction and presentation based on cloud collaboration; S7: Prototype system design and implementation.

2. The industrial parts installation auxiliary guidance method based on user intention perception according to claim 1 is characterized in that: In step S1, the internal parameter information of the camera, including the focal length and optical center of the camera, is obtained in advance, and matched with the internal parameters of the projector to determine their relative posture relationship. First, a checkerboard image with a grid number of at least 20×20 is selected as the source image, and the image is projected onto the workbench by a projector. Then, the complete projection screen is captured by the camera, and the feature points of the source image and the image captured by the camera are extracted using the ORB feature point extraction algorithm. The corresponding relationship between the images is established through the extracted feature points. Initialization, first use the camera to capture the video stream of the workspace and extract the image data, and use the target detection algorithm to detect the parts in the image to obtain the detection frame and classification identification information. The information is passed to the finite state automaton for subsequent state evaluation and state migration. The initial state of the system is set to the first zero-indegree node in the state diagram.

3. The industrial parts installation auxiliary guidance method based on user intention perception according to claim 2 is characterized in that: In step S2, the camera-projector system is installed on a fixed bracket to capture the working area. The field of view of the camera and the projector completely covers the projection area of ​​the workbench. During operation, the parts to be inspected are randomly placed in the working area, and the operator manually moves or picks up the parts to ensure that the hands and parts are always within the camera capture range. Through repeated operations, image data covering multiple perspectives and scene changes are collected; SAM is used to segment the first frame of the video stream. SAM determines the segmentation area by clicking on positive and negative samples and generates an initial mask of the object. Subsequently, the segmentation result of the first frame is transmitted to DeAOT for annotating all objects in subsequent frames. DeAOT performs multi-target segmentation on objects in the video sequence based on the initial mask and generates an annotated image sequence with a mask overlap consistency of more than 95% with the actual object contour. Using the Unity rendering environment, a diverse set of synthetic image datasets is generated. First, the position and rotation angle of the object are uniformly sampled within the camera's frustum. At the same time, a certain amount of random perturbation is introduced to the object's spatial coordinates and rotation angle to simulate changes in the real scene. Secondly, the robustness of the data is enhanced by dynamically changing the background, material, color, and lighting conditions. Finally, the dataset is labeled by generating object masks and corresponding images of pure color materials.

4. The industrial parts installation auxiliary guidance method based on user intention perception according to claim 3 is characterized in that: In step S3, the state perception and stability judgment of the assembly process are realized through assembly task decomposition and finite state automaton modeling, and state migration is performed in combination with real-time target detection data to complete the intelligent management of assembly tasks. Specifically, the image of the assembly operation area is captured by the camera, and each frame of image data analyzed by the target detection algorithm is transmitted to the finite state automaton. The finite state automaton performs a comprehensive analysis of multiple frames of continuous data based on the image data and context to determine whether the current assembly state is stable. To ensure the accuracy of the judgment, the finite state automaton calculates the mean value P of the target detection result in every 15 frames of image data. avg , the maximum and minimum difference P range and the normalized variance V f , the mathematical expression is as follows: P range =P max -P min Among them, P avg is the mean number of parts, P range is the difference between the maximum and minimum number of parts, V f is the normalized variance of the mean, is the number of parts in the fth frame, N f is the number of frames in the selected time window; When the mean P avg , the maximum and minimum difference P range and the normalized variance V f When the following conditions are met at the same time, the current state is considered stable and state switching is allowed: |P avg -round(P avg )|<0.25 P range <1.5 V f <1 The finite state automaton sets the current state to S c , and perform topological sorting based on the state diagram to obtain all possible states S for the next step c+1 , the current detection state is S d , the finite state automaton will S d With all S c+1 If the match is successful, the finite state automaton obtains the type and quantity information of each part through the target detection algorithm, and performs data matching according to the pre-defined state diagram. When the part data detected in real time fully meets the pre-defined conditions of a certain state, that is, the type and quantity of the part are consistent with the requirements in the state diagram, this is considered a successful match, and then the finite state automaton migrates the current state to S c+1 , and calculate S c+2 In all cases, the assembly task is decomposed into multiple stages and modeled using a finite state machine. A state is defined for each stage. The finite state machine determines the current state based on the part position and assembly progress. By defining state transition rules, intelligent perception and adjustment of the task process are achieved. When the conditions are met, the finite state machine will automatically migrate to the next state. Preferably, a dynamic context information analysis mechanism is introduced when judging state stability, and the state judgment strategy is dynamically adjusted by comprehensively analyzing the target detection data of the current frame and several frames before and after it, including the number of parts, location distribution and category confidence; In the process of acquiring target detection results, a confidence threshold is set for each detected part in combination with the confidence filtering mechanism. When the detection confidence is lower than the threshold, the detection result of the part will be ignored. At the same time, in order to enhance the continuity of the detection results, a time-weighted fusion strategy is adopted for the inter-frame detection results.

5. The industrial parts installation auxiliary guidance method based on user intention perception according to claim 4 is characterized in that: In the state migration step S3, the topological sorting rule is used to prioritize the next state set, and the highest priority state is gradually matched according to the detected state; The system uses the YOLO target detection algorithm to obtain the number and confidence of each part in the current image, and matches it with the part information required for each state preset by the finite state automaton, giving priority to matching the high-priority state closest to the current detection result; If the initial matching fails, go back to the current state and match again. In addition, record the historical path of each state transition.

6. The industrial parts installation auxiliary guidance method based on user intention perception according to claim 5 is characterized in that: In step S4, the object detection network in the assembly area is used to obtain the part and hand information of each frame, and the interaction state is determined. The intersection-and-union ratio S is calculated by the detection bounding box of the hand and the detection bounding box of the part. IoU ; Among them, B hand Represents the hand detection area, B p Represents the detection area of ​​part p, AoI represents the intersection area of ​​two bounding boxes, and AoU represents the union area of ​​two bounding boxes; The specific method is: initialize the interaction score S of each part p p is 0, when the intersection of the hand and the object is greater than S IoU When it is not 0, S p Accumulate the IoU value of part p in the current frame; S p =S p +S IoU When IoU is 0, Score uses a fixed decay function R d Attenuate to S p =0; S p =R d (S p ) The default decay function is an inverse proportional decay function: When the interaction score S p When the score exceeds the set threshold T, the hand is judged to be in an interactive state with the part; when the score is less than the threshold T, it is judged to be in a non-interactive state. To avoid the problem of multi-target confusion in interaction judgment, the interaction score is calculated for each part p instead of only the hand interaction score. When there are multiple objects with S p When both are greater than the threshold T, the object with the highest interaction score is selected as the current interaction object; Preferably, in step S4, in a multi-target interaction scenario, a priority strategy is adopted, and when the scores of multiple parts simultaneously exceed the threshold T, the target part with the most consistent direction with the hand trajectory is preferentially selected as the interaction object; During the interactive detection process, abnormal situations are handled. When the IoU between the hand and the part continues to be 0 and the Score increases abnormally, the misjudgment correction mechanism is triggered and the Score of the part is reinitialized.

7. The industrial parts installation auxiliary guidance method based on user intention perception according to claim 6 is characterized in that: In step S5, firstly, a real-time RGB image is captured by a mobile device, and the image is compressed and stored before being sent to a cloud server. The cloud server processes the received image data and obtains the current position and posture of the object using a three-dimensional tracking algorithm. At the same time, when the state stability judgment condition is met, the target three-dimensional model state is adjusted. Subsequently, the position and posture data and state information generated in the cloud are returned to the mobile device, completing the model matching and state update of the multi-state object, and realizing real-time three-dimensional tracking. Preferably, in step S5, before the RGB image captured by the mobile device is transmitted to the cloud, it is compressed using JPEG encoding and the image resolution is reduced.

8. The industrial parts installation auxiliary guidance method based on user intention perception according to claim 7 is characterized in that: In step S6, after the cloud server completes the processing, it sends the latest posture and model status information back to the mobile device. At this time, the mobile device application integrates the updated 3D model with the user's actual environment through augmented reality technology, and displays a detailed assembly animation. The animation is based on the tracked model and adds dynamic assembly instructions for the parts being assembled in the current state, dynamically guiding the user to perform the assembly work according to the correct steps; Preferably, in the augmented reality presentation process, the conversion of three-dimensional coordinates and the rendering and display of virtual models are realized based on the Unity platform and OpenGL technology, the X, Y and Z axis data of the three-dimensional coordinates are filtered, and the posture changes of the model in continuous frames are optimized through the interpolation algorithm to ensure the smoothness and stability of the display process; Combined with frame synchronization coding technology, the state switching of the virtual model and the posture change of the real object are corrected in real time. During the model switching process, each frame of data is marked with a frame number and matched with the posture information to ensure that the virtual model accurately corresponds to the actual state of the current object when the state is switched; In order to implement frame synchronization encoding technology, a frame number mark is added to each frame of data. At the same time, the pose information calculated in the cloud also contains the corresponding frame number. The specific matching process is as follows: Binding frame number tag with pose information: When the mobile device encodes and sends an image, each image frame is assigned a unique frame number, and the frame number is encoded into the image data packet and sent to the cloud server. After the cloud server calculates the pose corresponding to the frame image, it binds the calculated pose with the frame number of the image frame and returns them to the mobile device together; Receiving and matching: After receiving a data packet containing the posture and frame number, the mobile device will search for an image frame that matches the returned frame number in the buffer queue. By comparing the frame number stored in the buffer queue with the frame number returned by the cloud, the mobile device can accurately find the image frame corresponding to the current posture information.

9. The industrial parts installation auxiliary guidance method based on user intention perception according to claim 8, characterized in that: In step S7, an augmented reality assembly guidance prototype system is constructed to realize the assembly task guidance function of multi-module collaboration. The system is based on multi-state model management and real-time state perception technology, and completes the full process coverage from part detection, state judgment to dynamic guidance of the assembly process. In the assembly task, the user obtains real-time guidance information through the mobile device, including the superimposed display and operation prompts of the virtual model, dynamically adjusts the guidance content according to the current part status, and combines the data collected by the visual sensor to realize accurate planning and real-time feedback of the assembly steps. In addition, the system introduces virtual-reality fusion technology to enable seamless connection between the virtual model and the actual scene, providing users with intuitive and clear assembly guidance.

10. The industrial parts installation auxiliary guidance method based on user intention perception according to claim 9, characterized in that: In step S7, during the part picking process, the position of the target part is automatically identified and marked through a visual algorithm, providing clear guidance on picking and placing. During the assembly process, the connection path and operation prompts of the virtual model are dynamically displayed to ensure the accuracy of the assembly action. The operation steps and assembly status are displayed in real time through color identification and animation effects.