Work behavior intelligent identification and guidance method and system

By integrating multimodal perception information and adaptive learning, accurate identification and personalized guidance of operator behavior are achieved, solving the problems of accuracy and adaptability in operation behavior identification and guidance in existing technologies, and improving human-machine collaboration efficiency and production process stability.

CN122196936APending Publication Date: 2026-06-12SHANDONG HAIDE INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG HAIDE INTELLIGENT TECH CO LTD
Filing Date
2026-05-14
Publication Date
2026-06-12

Smart Images

  • Figure CN122196936A_ABST
    Figure CN122196936A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of industrial automation and artificial intelligence, and particularly relates to a work behavior intelligent identification and guidance method and system. The method comprises: collecting multi-modal perception information of a work environment; identifying a current work behavior of an operator based on the multi-modal perception information; performing matching analysis on the current work behavior and a predefined target work behavior sequence to generate behavior difference information containing time sequence and quality deviation; and generating and outputting guidance instructions for guiding the operator based on the behavior difference information, the instructions being capable of controlling action parameters of an auxiliary execution device or output rhythm of interface prompts. Through multi-modal fusion perception, fine behavior difference analysis, dynamic personalized guidance and adaptive learning, the present application realizes precise, intelligent and flexible real-time guidance and collaborative control of the work process, effectively improving work quality, efficiency and human-machine collaboration level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial automation and artificial intelligence, specifically to a method and system for intelligent recognition and guidance of work behavior. Background Technology

[0002] In manufacturing, logistics assembly, and precision operations, human-machine collaboration has become a key model for improving production flexibility and efficiency. However, existing work behavior recognition and guidance technologies still have significant limitations. Traditional methods mainly rely on a single vision sensor to monitor operator actions or deploy sensors in fixed locations to detect workpiece status. This single-modal perception approach is easily affected by occlusion, lighting changes, and environmental noise in complex and dynamic industrial scenarios, resulting in low accuracy in behavior recognition and difficulty in fully understanding the work context. Most existing guidance systems are based on pre-programmed fixed instructions or simple conditional triggering rules, such as playing static electronic work instructions or triggering audible and visual alarms at specific times. This static guidance method lacks adaptability to individual differences in operators (such as skill level and operating rhythm) and cannot be dynamically adjusted according to the real-time status of the work, often causing unnecessary interference to skilled workers or insufficient guidance to novices. When deviations occur in the work, the system can usually only issue an alarm for the result, unable to trace the root cause of the deviation and provide targeted corrective guidance, let alone have the ability to autonomously fine-tune to compensate for errors in the early stages of deviation. Furthermore, the system's rigid decision-making logic prevents it from learning and optimizing from historical interaction data, making it difficult for guidance strategies to evolve with changes in production conditions or the improvement of operator skills. Therefore, there is an urgent need for a comprehensive solution that can deeply integrate multi-dimensional information, accurately identify work intentions and effects in real time, and provide personalized, dynamic, and intelligent guidance accordingly, ultimately achieving efficient human-machine collaboration and autonomous assurance of work quality.

[0003] Therefore, the existing technology still needs further development. Summary of the Invention

[0004] The purpose of this invention is to overcome the above-mentioned technical deficiencies and provide a method and system for intelligent identification and guidance of work behavior to solve the problems existing in the prior art.

[0005] To achieve the above-mentioned technical objectives, according to a first aspect of the present invention, the present invention provides a method for intelligent identification and guidance of work behavior, comprising: S1. Collect multimodal sensing information of the working environment; S2. Identify the operator's current work behavior based on the multimodal perception information; S3. Match and analyze the current operation behavior with the predefined target operation behavior to generate behavior difference information; S4. Based on the behavioral difference information, generate and output guidance instructions to guide the operator to perform the target operation behavior; wherein, the guidance instructions are used to control the output rhythm of the action parameters or interface prompts of at least one auxiliary execution device.

[0006] Specifically, the step of identifying the operator's current work behavior based on the multimodal perception information includes: The visual information in the multimodal perception information is analyzed in a temporal sequence to extract the continuous motion trajectory of the operator's limb key points; based on the continuous motion trajectory, the operator's action units under the preset work sequence constraints are identified; the continuous action units are combined and parsed to form the current work behavior with work semantics.

[0007] Specifically, the target operation behavior includes a multi-target behavior sequence arranged in the order of the operation steps; the matching analysis between the current operation behavior and the predefined target operation behavior includes: determining the target operation node corresponding to the current operation behavior in the multi-target behavior sequence; determining whether the current operation behavior is completed within the expected time window of the target operation node; if it is not completed within the expected time window, the behavior difference information includes timing deviation information.

[0008] Specifically, after identifying the operator's current work behavior based on the multimodal perception information, the process further includes: The real-time status of the current work object is identified based on the multimodal perception information; the current work behavior is associated with the real-time status to determine whether the immediate effect of the current work behavior on the current work object meets the expected process effect; if not, the behavior difference information includes effect deviation information.

[0009] Specifically, the step of generating and outputting guidance instructions based on the behavioral difference information to guide the operator to perform the target task behavior includes: When the behavioral difference information is the effect deviation information, the potential behavioral root causes that lead to the effect deviation are inferred based on the historical data model; and the guidance instruction containing the operation points for correcting the potential behavioral root causes is generated based on the inferred potential behavioral root causes.

[0010] Specifically, the method further includes an adaptive learning step: Collect the multimodal perception information, the corresponding current work behavior, the behavior difference information, and the execution result data of the guidance instructions generated by the operator during multiple operations; based on the execution result data, perform incremental learning optimization on the decision model used to generate the guidance instructions, so as to reduce the behavior difference information generated in subsequent operations for the same operator.

[0011] Specifically, the incremental learning optimization of the decision model used to generate the guidance instructions based on the execution result data includes: Based on the operator's historical skill proficiency tags, personalized guidance strategy sub-models are established and updated for different skill proficiency groups. When a specific operator is identified, the personalized guidance strategy sub-model corresponding to the operator's skill proficiency group is invoked to generate guidance instructions adapted to the operator's habits.

[0012] Specifically, the method also includes an exception handling step: When the current operation behavior is identified as belonging to a predefined abnormal behavior category, an anomaly analysis process is triggered. The anomaly analysis process includes: predicting the possible operational defects caused by the abnormal behavior based on a knowledge base, and autonomously generating a preliminary process parameter adjustment plan; outputting the process parameter adjustment plan to the relevant execution equipment to attempt to compensate for the impact caused by the abnormal behavior.

[0013] Specifically, the anomaly analysis process also includes: After outputting the process parameter adjustment scheme, the multimodal perception information is continuously monitored to evaluate the improvement effect of the process parameter adjustment scheme on the operation results; the abnormal behavior, the process parameter adjustment scheme and the improvement effect are associated and stored to form closed-loop knowledge data for optimizing subsequent anomaly analysis and decision-making.

[0014] According to a second aspect of the present invention, a work behavior intelligent recognition and guidance system is provided, comprising: The sensing module is used to collect multimodal sensing information about the working environment; The behavior recognition module, connected to the perception module, is used to identify the operator's current work behavior based on the multimodal perception information; The analysis and decision-making module, connected to the behavior recognition module, is used to match and analyze the current operation behavior with the predefined target operation behavior to generate behavior difference information, and generate guidance instructions based on the behavior difference information; The guidance output module, connected to the analysis and decision module, is used to output the guidance instructions to control the output rhythm of the action parameters or interface prompts of at least one auxiliary execution device, thereby guiding the operator to perform the target operation.

[0015] Beneficial effects: Compared with existing technologies, the intelligent identification and guidance method and system for work behavior provided by this invention can bring the following significant benefits: First, by integrating and deeply fusing multimodal sensory information such as vision, force, and hearing, the system constructs a holographic and robust digital twin of the work site. This overcomes the limitations and fragility of single-sensor perception, making the identification of complex operator behaviors, the detection of subtle workpiece conditions, and the understanding of environmental context more accurate and reliable, laying a solid data foundation for subsequent intelligent decision-making.

[0016] Secondly, this invention creatively decomposes work behavior into a multi-layered recognition architecture of "motion trajectory - action unit - work semantics," and performs precise matching and quantitative analysis with predefined target behavior sequences in both temporal and effect dimensions. This not only enables the judgment of "whether the right thing is being done," but also quantitatively assesses "whether it is done on time" and "whether the effect meets the standard," thereby generating refined behavioral difference information that includes temporal and quality deviations, making guidance decisions more informed.

[0017] Third, based on refined behavioral difference information, the system can generate and output highly dynamic and personalized guidance instructions. These instructions go beyond just interface prompts; they can directly control the motion parameters (such as speed and trajectory) of collaborative robots and other assistive devices, achieving physical-level collaborative adaptation. The system can dynamically adjust the pace and intensity of prompts based on the operator's proficiency, providing detailed guidance for beginners and seamless collaboration for experienced users, thus achieving true "personalized instruction" and flexible human-machine collaboration.

[0018] Fourth, the system's built-in adaptive learning mechanism and closed-loop knowledge base enable it to continuously evolve. By collecting human-computer interaction data and incrementally optimizing the guidance strategy model, the system can continuously narrow the behavioral differences for specific production lines and specific operators, making the guidance strategy increasingly precise. Simultaneously, the closed-loop system for anomaly handling and effect evaluation allows the system to autonomously accumulate "craftsman experience," improving its ability to respond to unknown anomalies and achieving a leap from rigid automation to autonomous intelligence.

[0019] Finally, the system of this invention possesses preliminary autonomous anomaly handling and root cause diagnosis capabilities. When abnormal behavior or quality deviations are detected, the system not only issues an alarm but also infers potential root causes based on historical data models and attempts to autonomously adjust process parameters for compensation, or provides targeted operational guidance. This integrated "detection-diagnosis-handling" capability shifts quality control from pre-inspection to real-time process intervention, significantly reducing scrap rates and rework costs, and enhancing the resilience and stability of the production process. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating the intelligent identification and guidance method for work behavior provided in a specific embodiment of the present invention. Detailed Implementation

[0021] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Based on the embodiments in this application, other similar embodiments obtained by those skilled in the art without creative effort should all fall within the scope of protection of this application. Furthermore, directional terms mentioned in the following embodiments, such as "up," "down," "left," and "right," are only for reference to the directions in the accompanying drawings; therefore, the directional terms used are for illustrative purposes and not for limiting the invention.

[0022] The present invention will be further described below with reference to the accompanying drawings and preferred embodiments.

[0023] Please see Figure 1 This invention provides a method for intelligent recognition and guidance of work behavior, comprising: S1. Collect multimodal sensing information of the working environment.

[0024] It should be further explained that, in the specific implementation of this method, the step of "collecting multimodal perception information of the working environment" is achieved through an integrated sensor network. This network includes at least: multiple global and local industrial cameras (such as Baslerac A2440-75um, resolution 2448×2048, frame rate 75fps) deployed above and to the side of the workstation to capture visual information of the operator's limbs, tools, and workpieces from different perspectives; a six-dimensional force / torque sensor (such as ATI Mini40) and a nine-axis inertial measurement unit (IMU, such as Bosch BMI160) deployed on the operator's wrist, tool handle, or specific fixture to collect force feedback, torque, and attitude changes during the operation; and a high-precision microphone array embedded in the worktable to collect ambient sound, voice commands, and characteristic audio of tool-workpiece contact. All sensor data are time-synchronized through a precision timing module (such as the PTP protocol), sampled at a frequency of not less than 100Hz, packaged into a unified timestamped data stream, and transmitted to the edge computing unit.

[0025] S2. Identify the operator's current work behavior based on the multimodal perception information.

[0026] It should be further explained that the identification of the operator's current work behavior based on the aforementioned multimodal perception information specifically includes: The visual information in the multimodal perception information is analyzed in a temporal sequence to extract the continuous motion trajectory of the operator's limb key points; based on the continuous motion trajectory, the operator's action units under the preset work sequence constraints are identified; the continuous action units are combined and parsed to form the current work behavior with work semantics.

[0027] It should be further explained that the specific engineering implementation of this step includes the following detailed sub-steps: (1) Time series analysis and trajectory extraction: First, the synchronously acquired video stream is decoded. Using a pre-trained open-source human pose estimation algorithm library, Open Pose (using the COCO human keypoint model, outputting 18 keypoints) or the more accurate HR Net (outputting 17 keypoints), each frame is processed to obtain the two-dimensional pixel coordinates of the operator's body keypoints (such as left wrist, right wrist, left elbow, right elbow, nose, etc.). Considering viewpoint and occlusion, a multi-view camera is preferred, and the two-dimensional coordinates are converted to three-dimensional coordinates (unit: millimeters) relative to the world coordinate system using triangulation.

[0028] For frame t, obtain the set of coordinates of all keypoints. Where K is the total number of key points. This represents the 3D coordinates of the k-th keypoint at time t. Connecting the coordinates of the same keypoint in consecutive frames (e.g., 1 second, or 75 frames) forms the motion trajectory of that keypoint within a short time window. At the same time, IMU data (such as the angular velocity and acceleration of the tool handle) are integrated to correct errors in visual estimation during rapid movements and to help determine the activation status of the tool (such as whether the electric screwdriver is turned on).

[0029] (2) Motion unit recognition: The system predefines a library containing 15 basic motion units (MPs), such as: MP1 - "Reach" (hand moves towards the target), MP2 - "Grasp" (finger closes to grasp an object), MP3 - "Transport Empty" (moves without hands), MP4 - "Transport Loaded" (moves while holding an object), MP5 - "Position" (fine positioning and alignment), MP6 - "Assemble" (assembly and insertion), MP7 - "Use Tool" (use a tool), MP8 - "Release Load" (release an object), etc. Each MP has its corresponding trajectory feature template. These templates are established by collecting demonstration actions of skilled workers and extracting the statistical features of their key point trajectories (such as trajectory curvature, velocity profile, and acceleration extrema).

[0030] Furthermore, during recognition, a sliding time window (e.g., window length 0.5 seconds, step size 0.1 seconds) is used for continuous trajectories. Segmentation is performed. For each trajectory segment within a window, its feature vector is extracted. Then, a hybrid approach combining Dynamic Time Warping (DTW) distance calculation and Long Short-Term Memory (LSTM) neural network classifiers is used for identification.

[0031] First, calculate the DTW distance between the segment and all MP templates. The DTW distance measures the similarity between two time series that may have different lengths; the smaller the distance, the more similar they are. Select the top 3 MPs with the smallest DTW distance as candidates.

[0032] At the same time, the feature sequence of this segment Input a pre-trained two-layer LSTM classification network (each hidden layer has 128 units), and the network outputs the probability distribution of the segment belonging to each MP.

[0033] Finally, the DTW distance results (converted to probabilities) are weighted and fused with the output probabilities of the LSTM network. The weighting coefficients are typically set to 0.4 (DTW) and 0.6 (LSTM) because LSTM is better at capturing contextual temporal dependencies. The MP with the highest fused probability is selected as the action unit identified at the center of the window. The "preset operation sequence constraint" is represented here as a state machine or probabilistic graphical model. For example, in the "tightening screws" operation, the "Use Tool" action unit is likely to be preceded by the "Position" unit. The system uses this constraint to post-process and smooth the classification results. For example, using the Viterbi algorithm, a globally optimal MP sequence that conforms to the operation logic is found among the recognition results of adjacent time windows, thereby correcting the misidentification of isolated frames.

[0034] (3) Semantic parsing of job behavior: The identified action unit sequence (e.g., ["Reach", "Grasp", "Transport Loaded", "Position", "Use Tool"]) needs to be parsed into high-level behaviors with clear job objects and intentions. This is achieved through a rule-based and knowledge graph-based parsing engine. The system maintains a job knowledge graph, where nodes represent tools (e.g., "Electric Screwdriver E3"), parts (e.g., "M4x10 Screw"), and procedures (e.g., "Tightening base plate screws"). The parsing engine receives the MP sequence and combines it with objects identified in real time from visual information (using the YOLOv5 object detection model) and their spatial relationships (e.g., "the screwdriver head is within 5mm directly above the screw") to map the MP sequence to the operation nodes in the knowledge graph. For example, if the detected object is "M4x10 Screw", and the MP sequence is ["Position", "Use Tool"], and the tool is "Electric Screwdriver E3", then the parsed current job behavior is "Use Electric Screwdriver E3 to tighten screw M4x10". For more complex behaviors, a sequence-to-sequence (Seq2Seq) model can be used to encode the MP sequence and the object detection sequence together and decode them into behavior labels described in natural language.

[0035] Understandably, this three-tiered recognition architecture—"lower-level trajectory - middle-level unit - higher-level semantics"—enables the system to penetrate complex and ever-changing raw data and stably and accurately understand the micro and macro actions being performed by the operator. It not only recognizes "hand movement," but also understands "performing the specific step of tightening an M4x10 screw," providing precise semantic input for subsequent accurate behavioral compliance comparisons and discrepancy analysis. This is the cornerstone of the entire system's effective operation.

[0036] S3. Match and analyze the current operation behavior with the predefined target operation behavior to generate behavior difference information.

[0037] It should be further explained that step S3 specifically includes two core components: behavior matching and difference quantification. First, the system matches the identified current behavior (CB) with a predefined multi-target behavior sequence (Target Behavior Sequence) following a strict process order. Each target behavior (TB) in this sequence has a unique sequence identifier i and is associated with a standardized semantic description, expected tool, and expected workpiece. The matching process is accomplished by calculating the semantic similarity between the semantic description of CB and the current expected node and its neighboring nodes (e.g., the previous TB_i-1, the next TB_i+1) in the sequence. Semantic similarity calculation can be performed using text embedding vectors obtained based on a Sentence-BERT pre-trained model, and cosine similarity can be calculated. If the similarity score between CB and the current expected node TB_i is the highest and exceeds a preset matching threshold (e.g., 0.85), the match is considered successful, and CB corresponds to TB_i. If the match fails, it may indicate a skipped step, omission, or misoperation in the process, and the system will generate difference information indicating a "misaligned behavior sequence."

[0038] Furthermore, after a successful match, the system generates specific behavioral difference information from two dimensions: (1) Temporal difference: This is one of the core quantitative indicators of matching analysis. The system records the actual time when CB is determined to be completed. And the expected end time of the matched target behavior TB_i. Compare. Timing deviation. The calculation formula is: in, It is a timestamp indicating when the system determines the completion of the current task based on the behavior recognition model. It is the planned completion time for target behavior i, and also a timestamp. If This generates time-series difference information containing positive deviation values, indicating operational delays. Threshold The preferred settings are related to the production cycle time, and are usually... ,in This refers to the cycle time of the production line. For example, when the cycle time is 60 seconds, Set to 6 seconds. The reason for choosing this value is that it provides the operator with approximately 10% cycle time to accommodate normal fluctuations such as minor material differences or brief shifts in attention, avoiding overly sensitive alarms; at the same time, any delay exceeding this percentage usually indicates a real obstacle (such as defective parts or tool malfunctions), requiring timely system intervention to prevent it from becoming a production bottleneck.

[0039] (2) Effectiveness / Quality Difference: This dimension assesses whether the results of the action execution meet the process standards. After the CB is identified, the system immediately obtains the real-time status parameters of the work object through dedicated sensors (such as torque sensors and vision inspection units). (e.g., actual torque, adhesive path width). Compare this with the expected process performance standard range corresponding to this TB. Perform a comparison. If... This generates effect difference information, including deviation parameters, actual values, standard ranges, and deviation amounts. ( (This is the standard median value). For example, the standard torque for fastening screws is... N·m, if the actual measurement is 10.5 N·m, then N·m.

[0040] Understandably, through the precise matching and dual-dimensional (time and quality) difference quantification described above, step S3 transforms abstract operational behaviors into concrete and measurable performance indicators. This not only enables the real-time detection of operations deviating from the standard, but more importantly, it provides the structured difference data foundation required for subsequent intelligent decision-making, making guidance no longer based on vague experience-based judgments, but on precise quantitative analysis.

[0041] S4. Based on the behavioral difference information, generate and output guidance instructions to guide the operator to perform the target operation behavior; wherein, the guidance instructions are used to control the output rhythm of the action parameters or interface prompts of at least one auxiliary execution device.

[0042] It should be further explained that step S4 serves as a bridge between decision-making and execution. Based on the structured difference information generated in S3, it dynamically generates and executes targeted guidance strategies. The logic for generating guidance instructions is a function of the difference type and degree.

[0043] For time-series difference information ( The system is designed to help operators catch up on schedules or optimize their pace. Specific design features include: ① Boot command generation: The system generates boot instructions based on... Size rating (e.g., mild delay) Severe delay Different strategies may be adopted. For minor delays, an "acceleration prompt" instruction may be generated: speeding up the automatic page turning speed of the electronic work instruction (e-SOP) interface, or changing the highlighted arrow indicating the next operation from a steady flashing to a rapid flashing in the augmented reality (AR) glasses interface, visually urging acceleration. For severe delays, a "collaborative assistance" instruction may be generated: sending instructions to the collaborative robot to adjust its movement speed parameters in the next or current collaborative task, such as increasing the robot's movement speed from the standard value of 100% to 120%, enabling it to deliver parts to the operator more quickly, or remove completed workpieces more quickly, thereby saving time for the operator.

[0044] ② Command Output and Control: The "Acceleration Hint" command is sent to the AR glasses' client application via Wi-Fi or LAN in JSON format, controlling its rendering engine to change the animation rhythm. The "Collaborative Assistance" command directly modifies the robot controller's speed planning parameters through the robot control interface (such as UR Cap, Ether CAT).

[0045] For information on differences in effectiveness / quality ( (Out of tolerance), the system is designed to correct errors and prevent defects from occurring, and its specific design includes: ① Guiding instruction generation: The system first calls the root cause reasoning model (such as Gradient Boosting Decision Tree (GBDT) trained based on historical data) to analyze the causes... The possible causes are identified (e.g., "incorrect tool angle" or "incorrect parameter settings"). Then, a "corrective guidance" instruction is generated based on the inferred cause. For example, if the inferred cause is "insufficient torque due to screwdriver tilt," the instruction would be: In the AR glasses, overlay a virtual laser beam perpendicular to the workpiece surface directly above the screw hole as a visual reference, and display the text "Please align with the green guide line."

[0046] ② Command Output and Control: This command is also sent to the AR glasses for visualization rendering. Simultaneously, if the discrepancy is strongly correlated with equipment parameters (such as insufficient dispensing volume), the system may generate a "parameter adjustment command" in parallel, directly modifying the pressure or time parameters in the dispensing controller (PLC) via an industrial communication protocol (such as OPCUA), for example, adjusting the dispensing time from 0.5 seconds to 0.6 seconds for real-time compensation.

[0047] Understandably, step S4 achieves a closed loop from "problem discovery" to "problem resolution." Based on accurate diagnosis of discrepancies, it outputs multi-level, adaptive guidance instructions, ranging from information prompts to adjustments in physical parameters. This not only directly assists operators in correcting deviations in real time but also creates conditions in the physical environment that make it easier for operators to perform tasks correctly by adjusting the motion parameters of auxiliary equipment. This significantly improves work quality, efficiency, and first-pass yield, truly embodying the core value of intelligent and collaborative work guidance.

[0048] Furthermore, the "auxiliary execution device" includes a collaborative robot (such as UR5e), a programmable tightening shaft, an automatic dispensing valve, augmented reality (AR) glasses (such as Microsoft HoloLens 2), and a workstation touchscreen display. The "motion parameters" specifically refer to the motion speed (unit: mm / s), acceleration (unit: mm / s²), and trajectory path point sequence of the collaborative robot's end effector, or the final target torque (unit: N·m) and rotational speed (unit: rpm) of the tightening shaft. The "output rhythm of interface prompts" specifically refers to the timing, duration, and disappearance of virtual indicator arrows, highlighted borders, and text prompts within the AR glasses' field of view, or the triggering conditions and delay time for automatic page turning in the electronic work instruction (e-SOP) interface.

[0049] Understandably, by constructing a high-precision, highly synchronized multimodal perception network, the system can acquire comprehensive and dynamic data on "people, machines, materials, methods, and environment" at the work site, laying a reliable data foundation for subsequent refined behavioral understanding. By directly binding guidance commands with underlying equipment control parameters and the rhythm of human-machine interaction, guidance can extend from the information level to the physical execution level, achieving "flexible" control of the work process. This allows for proactive adaptation to human rhythm through adjusting the robot's coordinated movements, and also for dynamically adjusting the rhythm of information prompts to match human cognitive load, thereby significantly improving the naturalness of human-machine collaboration and overall work efficiency.

[0050] Specifically, the target operation behavior includes a multi-target behavior sequence arranged in the order of the operation steps; the matching analysis between the current operation behavior and the predefined target operation behavior includes: determining the target operation node corresponding to the current operation behavior in the multi-target behavior sequence; determining whether the current operation behavior is completed within the expected time window of the target operation node; if it is not completed within the expected time window, the behavior difference information includes timing deviation information.

[0051] It should be further explained that the implementation of matching analysis involves precise time management and calculation: (1) Constructing a multi-objective behavior sequence: The multi-objective behavior sequence is derived from the product's process flow diagram (P-Flow) and standard operating procedure (SOP). Each target behavior (TB) includes a semantic description of the behavior (e.g., "take M4x10 screws"), the associated tools and parts, and its sequence number in the entire sequence. Each TB is associated with an "expected time window". The settings for this window are crucial. A preferred approach is: ① Based on the "Model Time Tracing Method" (MODAPTS) or "Time Measurement System" (MTM), assign a standard time value to each basic action unit (MP), then decompose a TB into multiple MPs, sum their standard times, and obtain the theoretical duration of the TB. .

[0052] ②Then, by analyzing the historical data of multiple skilled workers (e.g., 5 workers) in actual operation, the statistical distribution of the actual duration of each TB is calculated, and the mean is taken. and standard deviation .

[0053] ③Finally, the first The expected time window for each TB is set as follows: , .in, It is a reasonable interval between actions, such as 1 second. The settings cover approximately 95% of the normal completion time for skilled workers, providing a reasonable tolerance for the target—neither too stringent nor too restrictive—while still capturing significant delays. (Select) Based on the mean ± The rule of thumb that covers approximately 95.4% of the data points is a commonly used boundary in quality control that strikes a good balance between controlling false positives (false alarms) and false negatives (missed alarms).

[0054] (2) Behavior Matching and Node Determination: The system maintains a pointer in real time, pointing to the target behavior node that is currently "expected" to occur in the multi-target behavior sequence. When a current behavior (CB) is identified, the system calculates the similarity between its semantic description and the semantic descriptions of TBs near the pointer position (such as the previous, current, and next). The similarity calculation can use cosine similarity based on word vectors or sentence embedding similarity based on pre-trained language models (such as BERT). If the similarity with the TB pointed to by the current pointer is the highest and exceeds the threshold (such as 0.85), the match is successful, and the CB corresponds to this TB node. If the similarity is insufficient, it may match with the next TB, and check whether there are any skipped steps or omissions.

[0055] (3) Timing Deviation Calculation: Let the actual time when the CB is determined to be completed by the system (e.g., when the system identifies "Release Load" or the tool leaves the workpiece, etc.) be 1. The expected end time of the i-th matched TB is... Timing deviation Defined as: The unit is seconds (s). Here, It is a timestamp indicating when the system determines the completion of the current task based on the behavior recognition model. This is the planned completion time for target behavior i, which is also a timestamp. Then, it's determined whether a timeout has occurred: if... If the timeout occurs, timing deviation information is generated. This is the trigger threshold for timing deviations. Taking into account normal fluctuations in actual operations, It should not be set to 0. A preferred setting is: ,in This refers to the cycle time of the production line. For example, for a workstation with a cycle time of 60 seconds, Set to 6 seconds. The rationale for this setting is that it allows the operator a 10% cycle time buffer to handle normal situations such as minor material differences or brief distractions, avoiding unnecessary stress on the operator; at the same time, any delay exceeding 10% may indicate a real problem (such as defective parts or operational difficulties), requiring system intervention to prevent the workstation from becoming a bottleneck and affecting the overall production line output.

[0056] Understandably, by breaking down work processes into sequences and managing their precise timing, the system achieves dynamic and quantitative monitoring of the workflow. Timing deviation information is not merely a "timeout" alarm, but also crucial data input for production cycle management and line balancing optimization. It enables the guidance system to provide predictive prompts, such as providing more prominent hints for the next action before an operator might exceed their time limit, or adjusting the assisting rhythm of collaborative robots, thereby proactively maintaining a smooth production flow. This is a significant manifestation of lean manufacturing at the digital level.

[0057] Specifically, after identifying the operator's current work behavior based on the multimodal perception information, the method further includes: identifying the real-time state of the current work object based on the multimodal perception information; associating the current work behavior with the real-time state to determine whether the immediate effect of the current work behavior on the current work object meets the expected process effect; if not, the behavior difference information includes effect deviation information.

[0058] It should be further explained that the identification of the status of the work object and the judgment of its effect are key links in achieving process quality control: (1) Real-time status identification of the work object: Specialized sensing and identification schemes are adopted for different processes. For the tightening process, the torque curve and final rotation angle of the tightening process are directly read and recorded by a high-precision torque sensor and angle encoder installed on the electric screwdriver. Final torque value (Unit: N·m) and angle (Unit: °) represents the real-time status of the work object (bolted connection pair), and the specific design includes: ① For dispensing / coating processes, a high-resolution linear scan camera (e.g., 2048 pixels) is mounted vertically downwards above the dispensing path to scan the adhesive before it cures. Image processing algorithms (e.g., edge detection, grayscale thresholding) are used to extract the contour of the adhesive path and calculate its average width. (Unit: mm), glue path height (assisted by laser triangulation rangefinder, unit: mm), and whether there are defects such as glue breakage or air bubbles.

[0059] ② For the welding process, use an infrared thermal imager to monitor the temperature field distribution of the weld joint and its surrounding heat-affected zone. The structured light 3D scanner was used to obtain the three-dimensional morphology of the weld joint after welding, and the height of the weld joint was calculated. and diameter .

[0060] ③ For assembly clearance inspection, a laser displacement sensor is used to measure the clearance between two assembled parts. .

[0061] (2) Performance Conformity Judgment: Each target operation (TB) is associated with a set of "expected process performance" criteria, usually given in the form of parameter ranges. For example, for the operation of "fastening screws M4x10", the expected performance criterion is: final torque N·m (target value 12.0 N·m), the rotation angle should continue to rotate after the torque reaches the target value. The judgment logic is: check whether the actual state value falls within the expected range. For example, for torque, the judgment condition is: For image quality, such as the uniformity of dispensing width, this can be assessed by calculating the standard deviation of the actual adhesive path width sequence. To determine, the requirements mm. For weld joint morphology, the normalized cross-correlation coefficient (NCC) or structural similarity index (SSIM) between the actual weld joint profile and the standard profile template can be calculated. The formula for calculating SSIM is: Where x and y are the actual solder joint image block and the standard solder joint image block, respectively. It is its mean. It is its variance. It is its covariance. This is a constant used for stability calculations. If If the value is below 0.85, it is considered unsatisfactory. The SSIM threshold of 0.85 is chosen based on extensive image quality assessment practices. This value can effectively distinguish between similarity acceptable to the human eye and obvious quality differences. A value below 0.85 usually means that there are visible defects that may affect functionality or reliability.

[0062] (3) Association and Deviation Generation: The system associates the identified "current work behavior" with the "real-time status of the work object" collected immediately afterward on the timeline. For example, the torque value read within 100 milliseconds after the tightening behavior ends (the electric screwdriver stops rotating) is identified as the effect of the behavior. If it is determined to be non-compliant, deviation information of the effect is generated. This information includes at least: deviation parameters (such as "torque"), actual value (such as 10.5 N·m), expected range ([11.5,12.5] N·m), deviation amount (such as -1.5 N·m), and severity level (for example, if the deviation amount accounts for more than 50% of the tolerance zone, it is marked as "severe").

[0063] Understandably, this step shifts quality control from "outcome inspection" to "process monitoring." Traditionally, issues such as torque, adhesive properties, and weld quality might only be discovered at the final inspection station or even during customer use. This method performs online, real-time assessment of the effect the moment the operation is completed. Once a deviation is detected, it can immediately pinpoint which operation and which workpiece caused the problem, enabling immediate traceability and interception of quality issues. This significantly reduces the flow of defective products and rework costs, making it a core technological means to achieve "zero-defect" production and the digitalization and transparency of the manufacturing process.

[0064] Specifically, when the behavioral difference information is the effect deviation information, based on the behavioral difference information, a guidance instruction is generated and output to guide the operator to perform the target operation behavior, including: Based on historical data models, the potential behavioral root causes that lead to deviations in the effect are inferred; based on the inferred potential behavioral root causes, the guiding instructions containing the key points of operation to correct the potential behavioral root causes are generated.

[0065] It should be further explained that root cause reasoning and precise guidance are the core manifestations of system intelligence: (1) Historical Data Model Construction and Inference: The historical data model preferably adopts the Gradient Boosting Decision Tree (GBDT) model or the Random Forest model, because it has good nonlinear fitting ability and feature importance ranking function. The training data of the model comes from historical homework data samples collected over a long period of time. Each sample is a feature vector, containing: a) Multimodal perception features within a time window before the effect occurs (e.g., 5 seconds before tightening), including dozens of features such as the average speed of the operator's hand key points, acceleration variance, tool posture angle, screwdriver feed force, initial alignment offset of screw and screw hole (calculated visually), and rotation speed of electric screwdriver when started.

[0066] b) The sample's label, representing the root cause category leading to the performance deviation, such as "Root_Cause_1: Screw tilted during engagement," "Root_Cause_2: Excessive rotation speed," "Root_Cause_3: Incorrect screw type," "Root_Cause_4: Foreign object in the screw hole," and "Root_Cause_5: Uneven workpiece surface." These root cause labels were initially manually annotated by process engineers based on video playback and data analysis. Once a certain number (e.g., 5000 labeled anomalous samples) were accumulated, they were used for supervised learning. During inference, when a new performance deviation (e.g., insufficient torque) is detected, the system extracts perceptual features from the same time window preceding the current operation and inputs them into the trained GBDT model. The model outputs probability predictions for each potential root cause category, for example: P("Screw tilted during engagement") = 0.52, P("Excessive rotation speed") = 0.25, P("Others") = 0.23. The system selects the category with the highest probability (if its probability exceeds the threshold of 0.4) as the most likely potential root cause.

[0067] (2) Targeted Guidance Instruction Generation: The system maintains a knowledge base mapping "root cause - operation points - guidance instructions". This knowledge base is stored in a structured manner. For example, for the root cause "screw tilting during insertion", the corresponding operation point is "keep the screwdriver axis perpendicular to the workpiece surface and apply stable downward pressure". The guidance instruction is specified as follows: control the AR glasses to project a continuously flashing virtual laser column perpendicular to the workpiece surface (as a vertical reference) directly above the screw hole in the operator's field of vision, while displaying a text prompt on the side of the field of vision: "Please keep the tool aligned with the green laser column before pressing down". For the root cause "excessive speed", the operation point is "reduce the initial speed to ensure the screw is smoothly inserted". The guidance instruction is specified as follows: send an instruction to the electric screwdriver driver through the PLC to temporarily adjust the speed parameter of its next action from the default 800rpm to 500rpm, and pop up a prompt box on the touch screen: "The speed has been reduced for you, please try again". This guidance directly acts on the micro-operational level that causes the problem, providing clear corrective measures.

[0068] (3) Threshold selection: The probability threshold of 0.4 is chosen based on a trade-off between precision and recall. During the model development phase, by plotting PR curves on the validation set, the probability threshold corresponding to the point where the recall is relatively high while maintaining a high precision (e.g., >85%) is selected. 0.4 is an empirical value, which means that targeted guidance is triggered when the model has a moderate to high degree of confidence in a certain root cause, neither too conservative (to avoid missing the real cause) nor too aggressive (to avoid giving frequent incorrect guidance).

[0069] Understandably, this step transforms the system's role from a passive "defect detector" to a proactive "process diagnostic assistant" and "intelligent coach." Instead of simply reporting "torque non-compliant," it analyzes "it might be because the screw is crooked" and guides the operator to "align it vertically." This deep causal inference and precise intervention significantly reduces the knowledge threshold and time cost required for operators to troubleshoot problems, enabling rapid correction of poor operating habits, improving first-pass yield, and providing significant value for new employee training.

[0070] Specifically, the method further includes an adaptive learning step: collecting the multimodal perception information generated by the operator during multiple operations, the corresponding current operation behavior, the behavior difference information, and the execution result data of the guidance instructions; based on the execution result data, performing incremental learning optimization on the decision model used to generate the guidance instructions, so as to reduce the behavior difference information generated in subsequent operations for the same operator.

[0071] It should be further explained that the adaptive learning step aims to enable the system to continuously evolve, and its implementation includes a complete data closure loop and model update mechanism: (1) Data Collection and Storage: The system creates a data package for each operation instance (i.e., completing a complete product or process unit). This data package contains: a time-series snapshot of the full-cycle multimodal perception information from the start to the end of the operation (to save storage, only key features can be saved instead of raw stream data); all "current operation behaviors" sequences identified by the system and their timestamps; all calculated "behavioral difference information" (including timing deviations and effect deviations); any "guidance instructions" generated by the system and their issuance timestamps; and most importantly—"execution result data". Execution result data needs to be quantitatively defined. For example, for a guidance instruction to correct "screw tilt", its execution result can be comprehensively judged by comparing whether the standard deviation of the operator's hand verticality decreased before and after the guidance was issued, and whether the subsequent tightening torque was qualified.

[0072] Furthermore, a quantifiable result reward It can be designed as follows: ,in, It is the degree of improvement of relevant behavioral indicators (such as verticality) after guidance (later value - previous value, normalized to [-1,1]). It represents the pass / fail rating of the post-guidance work effect (1 for pass, 0 for fail). and It is weight, for example They place greater emphasis on the final quality results. It is the instant reward value for a single guidance session, and it is a scalar. It represents the degree of improvement of relevant behavioral indicators after guidance. It is calculated by comparing the average value of the indicator within a time window before and after guidance, and then undergoing maximum and minimum normalization to make it range between -1 and 1, with positive values ​​indicating improvement. It is the binary pass / fail rating of the post-guidance operation effect, where 1 indicates pass and 0 indicates fail. and These are weighting coefficients, assigning relative importance to behavioral improvement and outcome compliance, respectively. Here, they are set to 0.3 and 0.7, reflecting an outcome-oriented approach, meaning that final quality compliance is more important than intermediate behavioral improvement. This data is stored in a time-series database.

[0073] (2) Decision model and incremental learning: The decision model used to generate guidance instructions is essentially a policy function. It receives the current state. (Including perceived features, identified behaviors, and differential information, etc.), outputting the guiding actions to be taken. (Such as prompt type, content, device parameter adjustment values). In the initial stage, It can be based on rules or a simple model. The goal of adaptive learning is to optimize this policy. This can be modeled as a reinforcement learning (RL) problem. State It is a set of the above features. Action This is the boot instruction space. (Reward) That is, the definition in the previous step The system uses the Proximal Policy Optimization (PPO) algorithm to optimize the policy network. Incremental learning is performed. Every so often (e.g., after collecting N=2000 new data samples), the system initiates an offline training cycle. Using these samples (states, actions, rewards, new states), the advantage function is calculated, and then the PPO update step is executed to update the parameters of the policy network. The PPO algorithm aims to maximize the expected cumulative reward. Through importance sampling and pruning mechanisms, it achieves stable and efficient policy updates. The learning objective is to make the updated policy... When faced with similar situations, they can choose the guiding action that yields a higher reward (i.e., is more effective at narrowing behavioral differences).

[0074] (3) Specification of learning objectives: "Reducing behavioral differences" is already reflected in the reward function. For example, if, after guidance, the temporal deviation in subsequent task cycles... The average value decreases, or the effect deviates. A decrease in the average value of each state leads to an increase in the expected reward value for subsequent states, thereby driving the policy model to learn guidance methods that produce these results. The system periodically (e.g., weekly) evaluates the performance of new policies on the validation set to ensure performance improvements.

[0075] Understandably, adaptive learning mechanisms enable the system to continuously acquire feedback from actual human-computer interactions and optimize its guidance strategies. This makes the system no longer static and one-size-fits-all, but dynamic and personalized. It can gradually learn the optimal guidance timing and methods for specific production lines, specific products, and even specific operators, realizing the transformation from "general logic" to "specific optimization," thereby continuously improving the overall efficiency and adaptability of human-machine collaboration.

[0076] Specifically, the incremental learning optimization of the decision model used to generate the guidance instructions based on the execution result data includes: establishing and updating personalized guidance strategy sub-models for different skill levels of operators based on the operator's historical job proficiency tags; when a specific operator is identified, calling the personalized guidance strategy sub-model corresponding to the skill level group to which the operator belongs, so as to generate the guidance instructions adapted to the operator's habits.

[0077] It should be further explained that this is a further refinement and personalized implementation of the adaptive learning steps: (1) Dynamic assessment of proficiency tags: The system maintains a dynamically updated proficiency score for each operator. This score is calculated based on historical performance data from multiple dimensions, using the following formula: .in, It is a comprehensive score of the operator's proficiency, usually normalized to 0-100. It is an efficiency score, calculated based on the ratio of the operator's average cycle time to the standard time, such as... The higher the value, the higher the efficiency. It is a quality score, calculated based on the operator's historical First Pass Yield (FPY), such as... . It is the guidance dependence coefficient, which is calculated as the frequency with which the operator receives guidance (number of guidances / total number of operation steps). The lower the value, the stronger the independence. These are weighting coefficients, for example, they can be set to... This indicates that efficiency and quality are equally important, and the degree of reliance on guidance also takes into account. These are weighting coefficients, satisfying... Here, the values ​​are set to 0.4, 0.4, and 0.2, reflecting that efficiency (E) and quality (Q) are considered core and equally important indicators when assessing proficiency, while dependence on guidance (C) is considered a secondary but not negligible indicator. The system recalculates the proficiency of each operator weekly or after completing a certain number of tasks. Value. Then, according to Cluster the values ​​(e.g., K-means clustering, K=3) to divide all operators into three groups: "novice" (S<60), "skilled worker" (60≤S<85), and "expert" (S≥85).

[0078] (2) Establishment and updating of personalized sub-models: The system maintains three independent guidance strategy sub-models for the above three groups, denoted as... Each sub-model has the same structure (e.g., all are neural networks), but their parameters are independent. During the PPO incremental learning described above, the data samples generated by the operator are only used to update the sub-model corresponding to their skill level group. For example, the (state, action, reward) data generated by a "novice" operator is only used to update... The parameters. In this way, each sub-model specifically learns how to interact optimally with an operator of a corresponding skill level. For example, You will learn to provide more basic, detailed, and proactive guidance when faced with common hesitations and mistakes made by beginners; It will then learn to only provide very restrained, non-intrusive prompts when it detects minute, high-risk biases that even experts might easily overlook.

[0079] (3) Identification and Model Invocation: Operators start their workstations by swiping their cards, using facial recognition, or logging in with their personal employee ID. After identifying their identity, the system immediately queries the database for their latest proficiency group tags and loads the corresponding personalized guidance strategy sub-model into memory. In subsequent operations, all behavior recognition and differential analysis results will be input into this dedicated sub-model, which will determine whether to provide guidance, when to provide guidance, and how to provide guidance. For newly hired operators without historical data, the system defaults to assigning a "newbie" sub-model and begins accumulating their data.

[0080] Understandably, by establishing personalized guidance models for operators of varying skill levels, the system achieves true "personalized instruction." This avoids interfering with skilled workers by providing detailed instructions for novices, and also avoids providing overly simplistic guidance even to experts. This tiered strategy significantly improves the effectiveness of guidance and the user experience. Furthermore, as operators' skills improve (their skill level score S increases, potentially changing the group they belong to), the system automatically switches to the model corresponding to the new group, achieving synchronous evolution of the guidance strategy and the operator's skill development. This provides a powerful tool for the digital management and adaptive training of personnel skills in manufacturing enterprises.

[0081] Specifically, the method also includes an exception handling step: When the current operation behavior is identified as belonging to a predefined abnormal behavior category, an anomaly analysis process is triggered. The anomaly analysis process includes: predicting the possible operational defects caused by the abnormal behavior based on a knowledge base, and autonomously generating a preliminary process parameter adjustment plan; outputting the process parameter adjustment plan to the relevant execution equipment to attempt to compensate for the impact caused by the abnormal behavior.

[0082] It should be further explained that the abnormality handling procedure endows the system with preliminary autonomous decision-making and fault-tolerant control capabilities, and its implementation mechanism is as follows: (1) Predefined Abnormal Behavior Recognition: Abnormal behavior categories are predefined by process engineers during system initialization, and typically include: A1 - Tool drop, A2 - Significant workpiece placement deviation (e.g., visual detection of deviation > 3mm), A3 - Screw stripping (judged by sound spectrum analysis or a sharp drop in torque curve), A4 - Dispensing failure (line scan camera detects adhesive path width as 0 for 5 consecutive frames), A5 - Abnormal increase in welding spatter (high-speed camera captures spatter particles exceeding the threshold of 50 / second), A6 - Abnormal operator gestures (e.g., suspected hammering of the workpiece). The recognition of these abnormal behaviors usually relies on specific sensors and algorithms. For example, tool drop can be judged by combining the weightlessness state of the tool handle IMU and a sudden drop in force sensor readings; screw stripping can be detected by analyzing the sound signal during tightening and detecting a sudden increase in energy in a specific frequency band (e.g., 2kHz-4kHz).

[0083] (2) Knowledge Base and Autonomous Solution Generation: The system maintains a graph database of "abnormal behavior - defect prediction - compensation solution". Each abnormal behavior node is associated with multiple possible "defect" nodes, as well as the "occurrence probability" edge from the abnormality to each defect. Each defect node is associated with one or more "compensation solution" nodes. The compensation solution node contains specific process parameter adjustment instructions. For example: ① Abnormal behavior: "A4 - Dispensing failure".

[0084] ②Possible defects: "D1 - Insufficient sealing" (probability 0.8), "D2 - Poor electrical connection" (probability 0.2).

[0085] ③ Compensation plan for “D1 - Insufficient sealing”: “C1: After the glue break point, increase the dispensing pressure of the glue valve by 15% and reduce the moving speed by 20% to perform glue replenishment with a length of 10mm.”

[0086] ④ Compensation scheme for "D2 - Poor electrical connection": "C2: At the point of adhesive breakage, activate the auxiliary conductive adhesive dispensing head for patching." When the system detects the "A4 - Adhesive dispensing failure" anomaly, it traverses the knowledge base to find all associated defects and their probabilities. Then, based on the severity of the defect (with adjustable weights) and its probability of occurrence, it calculates the expected utility of each compensation scheme. The scheme with the highest utility, such as C1, is selected, and its parameters (pressure increased by 15%, speed decreased by 20%) are converted into specific executable equipment instructions: G-code commands or PLC write commands.

[0087] (3) Autonomous Compensation: The generated process parameter adjustment plan is sent to the controller of the relevant execution equipment in real time. For example, the adjusted dispensing parameters are sent to the motion control card of the dispensing machine via the EtherCAT bus. After sending the command, the system waits for confirmation feedback from the equipment. At the same time, a prompt message is displayed to the operator on the interface (such as HMI), such as "Dispensing interruption detected, automatic dispensing replenishment program has been executed, please pay attention." This ensures that the operator is aware of the system's automatic intervention.

[0088] (4) Reasons for Threshold and Parameter Selection: The anomaly identification thresholds, such as "offset > 3mm" and "splash particles > 50 / second", are set based on process tolerances and a large amount of experimental data. An offset of 3mm usually exceeds the tolerance of locating pins or guide holes in most assembly processes, which will inevitably lead to assembly difficulties or interference. The splash threshold of 50 / second is based on the experience of welding process experts. Exceeding this number usually means insufficient shielding gas, improper voltage and current parameters, or material surface contamination, resulting in a high risk of weld quality. The specific parameters in the compensation scheme (such as a 15% increase in pressure) are derived from the process test database, which records the most effective combination of compensation parameters for such anomalies.

[0089] Understandably, this step gives the system a preliminary "self-healing" capability. It can react quickly to some common, well-defined abnormal operating conditions, attempting online compensation before defects solidify or affect downstream processes. This not only has the potential to save individual products and prevent scrap, but more importantly, it maintains the continuous flow of the production line, reduces production interruption time caused by abnormal downtime and handling of abnormalities, and improves overall equipment efficiency (OEE).

[0090] Specifically, the anomaly analysis process also includes: After outputting the process parameter adjustment scheme, the multimodal perception information is continuously monitored to evaluate the improvement effect of the process parameter adjustment scheme on the operation results; the abnormal behavior, the process parameter adjustment scheme and the improvement effect are associated and stored to form closed-loop knowledge data for optimizing subsequent anomaly analysis and decision-making.

[0091] It should be further explained that this is the closed-loop optimization part of the abnormal autonomous handling steps, ensuring that the system can learn from each intervention: (1) Evaluation of Improvement Effect: After the system executes the compensation plan, a monitoring window is immediately activated. For example, in the case of "re-applying glue after glue dispensing interruption", the monitoring window is the visual inspection of the re-applied area and the subsequent normal dispensing section (e.g., the subsequent 20mm glue path) after the glue re-applying is completed. The evaluation indicators are the width continuity and height consistency of the glue path. The average width of the glue path in the re-applied section can be calculated. and standard deviation , with the normal segment and Compare and define the improvement effect index. for: .in, It is the improvement effect index, which ranges from (-∞, 1]. The closer it is to 1, the better the improvement effect. Equal to 1 means complete recovery to normal. It is the average width of the adhesive path after compensation, in mm. This is the average width of the normal section of the adhesive path, in mm. It is the absolute difference between the two, reflecting the degree of mean recovery. It is the standard deviation of the width of the adhesive path after compensation, in mm, reflecting consistency. This is a weighting coefficient, usually set to 0.5, representing the penalty weight for consistency (standard deviation). The formula considers both whether the compensated adhesive width returns to normal levels and its uniformity. If If so, the compensation plan is considered very effective; if They consider it moderately effective; if If the result is poor or ineffective, it is considered to be unsatisfactory.

[0092] (2) Closed-loop knowledge storage: The complete record of this abnormal event, including: the type of abnormal behavior, the timestamp of occurrence, the multimodal data fragment at the moment of the abnormality, the triggered compensation plan (including all adjustment parameters), and the calculated improvement effect index. This is stored as a new "case" in the graph database mentioned above. This case will be associated with similar cases already existing in the knowledge base.

[0093] (3) Knowledge base optimization: The optimization of the knowledge base is reflected in two aspects. First, the utility of the solution is updated: For multiple different compensation solutions that cause the same defect, the system will continuously track their history. Value. Each solution node will maintain an average improvement effect. and number of executions When a new case is added, the corresponding solution... and Updated. If the same anomaly occurs again in the future, the system will prioritize recommending [recommendations / features]. High-value solutions. Secondly, knowledge discovery: After accumulating a large number of cases, new and undefined "anomaly-defect-compensation" relationships can be discovered through data mining techniques (such as association rule learning). After confirmation by process engineers, these relationships can be added to the knowledge base as new rules, thereby expanding the system's anomaly handling capabilities.

[0094] (4) Threshold reason: Improvement effect index The judgment thresholds of 0.5 and 0.8 are set according to actual process requirements. This means that the quality indicators after compensation are very close to those under normal conditions, and within the process tolerance range, they can be considered completely effective. This means that the compensation failed to bring the key quality indicators back to the qualified level, the plan was not effective and needs to be downgraded or re-evaluated.

[0095] Understandably, this closed-loop learning process is crucial for the system's intelligence to move from "rule-based automation" to "experience-based autonomous optimization." It allows the system's anomaly handling capabilities to move beyond fixed rules initially programmed by engineers, enabling it to continuously learn, experiment, and evolve from real production data. Effective compensation strategies are reinforced, while ineffective ones are weakened. Over time, the accumulated "experience" becomes a unique and highly valuable process knowledge asset for the production line, capable of handling increasingly complex and personalized anomalies, significantly improving the resilience and intelligence level of the production system.

[0096] This invention provides another embodiment, which offers a work behavior intelligent recognition and guidance system, the work behavior intelligent recognition and guidance system comprising: (1) Sensing module, used to collect multimodal sensing information of the working environment.

[0097] ① It should be further explained that this system is a concrete implementation of the aforementioned method at the software and hardware level. The detailed composition and cooperation relationship of its various modules are as follows: The sensing module is a heterogeneous sensor array and its data acquisition and fusion unit, specifically including: ② Two global RGB industrial cameras (such as Hikvision MV-CH250-10UM) for panoramic monitoring and one zoom lens camera (such as Basler ace 2) for capturing local details are connected to a gigabit Ethernet switch via GigE Vision interface. ③ A set of data gloves (such as Manus Prime II Haptic) or multiple nine-axis IMU sensors (sampling rate 100Hz) worn on the wrist for collecting fine hand movements of the operator, which communicate with the host via a Bluetooth 5.0 hub; ④ A six-dimensional force sensor (such as the Yuli Instruments IIT-Mini40) for acquiring force / torque data at the tool end, outputting via a built-in analog-to-Ethernet module; and a ring microphone array (such as the ReSpeaker 6-Mic Array) for acquiring ambient and tool sounds, connected via a USB interface.

[0098] Furthermore, all sensors are synchronized at the nanosecond level by a single synchronization controller (such as SMC Networks' IEEE 1588 PTP master clock). Data acquisition is performed by multi-threaded acquisition software on an industrial edge computer (such as Advantech UNO-2484G with a built-in Intel i7 processor). This software timestamps the data streams from each sensor and encapsulates them into standard ROS (Robot Operating System) topic messages, which are then published to the system's internal communication bus.

[0099] (2) Behavior recognition module, connected to the perception module, used to identify the operator's current work behavior based on the multimodal perception information.

[0100] It should be further noted that the behavior recognition module is deployed on an industrial server equipped with a high-performance GPU (such as an NVIDIA RTX A5000), running deep learning model containers based on the PyTorch or TensorFlow framework. This module subscribes to ROS topics published by the perception module, first running an Open Pose or Media Pipe model for real-time 2D pose estimation (latency <50ms), then running a custom 3D pose LSTM network for temporal action unit classification, and finally calling a behavior semantic parsing service based on BERT fine-tuning to output a structured current job behavior object (including behavior type, tool, artifact, and time).

[0101] (3) Analysis and decision module, connected to the behavior recognition module, used to match and analyze the current operation behavior with the predefined target operation behavior to generate behavior difference information, and generate guidance instructions based on the behavior difference information.

[0102] It should be further explained that the analysis and decision-making module is the core logical processing unit of the system and can be deployed on the same server or in a cloud virtual machine. It consists of a real-time database (such as Redis), a rule engine (such as Drools), and multiple microservices (such as time series analysis service, difference calculation service, root cause reasoning service, adaptive learning model service, and anomaly handling service). This module receives the results from the behavior recognition module, queries the database for target behavior sequences and process standards, and executes a series of complex logics such as matching, deviation calculation, guided decision-making, and anomaly handling. Its internal decision model (such as the PPO policy network) and knowledge base (such as the Neo4j graph database) are periodically updated by adaptive learning steps and closed-loop knowledge storage steps.

[0103] (4) A guidance output module, connected to the analysis and decision module, is used to output the guidance instructions to control the output rhythm of the action parameters or interface prompts of at least one auxiliary execution device, thereby guiding the operator to perform the target operation.

[0104] It should be further explained that the guidance output module is a multi-protocol interface adaptation and instruction distribution unit. It receives standardized guidance instruction objects from the analysis and decision-making module and then converts them into specific device protocol instructions: for collaborative robots (such as UR5e), it sends UR Script programs via the UR Cap interface or Socket communication to control their movement; for AR glasses (such as HoloLens 2), it calls their REST API via Wi-Fi to send JSON instructions containing graphics, text, and location, driving them to render prompts; for PLCs / industrial control computers, it writes to control registers via OPC UA or Modbus TCP protocol to control indicator lights, buzzers, or adjust tightening axis parameters. This module ensures low latency (end-to-end <200ms) and reliable transmission of instructions.

[0105] Understandably, this system integrates complex multimodal sensing, intelligent recognition, real-time decision-making, and precise control technologies into a cohesive whole through a modular hardware architecture and software design. The sensing module, acting as the system's "sensors," provides comprehensive, accurate, and synchronous field data; the behavior recognition module, acting as the system's "visual and understanding cortex," extracts advanced semantic information from massive amounts of data; the analysis and decision-making module, acting as the system's "brain," incorporates process knowledge, artificial intelligence models, and adaptive learning capabilities to make intelligent decisions; and the guidance and output module, acting as the system's "nerve endings and effectors," accurately translates decisions into changes in the physical world or the operator's perception. This architecture gives the system excellent scalability (flexible addition or removal of sensors or actuators), maintainability, and high real-time performance, meeting the stringent requirements of real-time, precise, and intelligent guidance for human-machine collaborative operations in modern intelligent manufacturing.

[0106] In a preferred embodiment, this application also provides an electronic device, the electronic device comprising: The computer device includes a memory and a processor, wherein the memory stores computer-readable instructions that, when executed by the processor, implement the intelligent recognition and guidance method for job behavior. The computer device can be broadly categorized as a server, terminal, or any other electronic device with the necessary computing and / or processing capabilities. In one embodiment, the computer device may include a processor, memory, network interface, communication interface, etc., connected via a system bus. The processor of the computer device can be used to provide the necessary computing, processing, and / or control capabilities. The memory of the computer device may include a non-volatile storage medium and internal memory. The non-volatile storage medium may store an operating system, computer programs, etc. The internal memory can provide an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface and communication interface of the computer device can be used to connect and communicate with external devices via a network. When the computer program is executed by the processor, it performs the steps of the method of the present invention.

[0107] This invention can be implemented as a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the steps of the methods of embodiments of the invention to be performed. In one embodiment, the computer program is distributed across multiple network-coupled computer devices or processors, such that the computer program is stored, accessed, and executed in a distributed manner by one or more computer devices or processors. A single method step / operation, or two or more method steps / operations, may be executed by a single computer device or processor or by two or more computer devices or processors. One or more method steps / operations may be executed by one or more computer devices or processors, and one or more other method steps / operations may be executed by one or more other computer devices or processors. One or more computer devices or processors may execute a single method step / operation, or execute two or more method steps / operations.

[0108] Those skilled in the art will understand that the method steps of this invention can be performed by a computer program instructing related hardware, such as a computer device or processor, to perform the steps of this invention when executed. Depending on the context, any references herein to memory, storage, databases, or other media may include non-volatile and / or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.

[0109] The technical features described above can be combined arbitrarily. Although not all possible combinations of these technical features are described, any combination of these technical features should be considered to be covered by this specification, provided that such combination does not contain contradictions.

[0110] The specific embodiments of the present invention described above do not constitute a limitation on the scope of protection of the present invention. Any other corresponding changes and modifications made in accordance with the technical concept of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A method for intelligent recognition and guidance of work behavior, characterized in that, Includes the following steps: S1. Collect multimodal sensing information of the working environment; S2. Identify the operator's current work behavior based on the multimodal perception information; S3. Match and analyze the current operation behavior with the predefined target operation behavior to generate behavior difference information; S4. Based on the behavioral difference information, generate and output guidance instructions to guide the operator to perform the target operation behavior; wherein, the guidance instructions are used to control the output rhythm of the action parameters or interface prompts of at least one auxiliary execution device; After identifying the operator's current work behavior based on the multimodal perception information, the method further includes: The real-time status of the current work object is identified based on the multimodal perception information; the current work behavior is associated with the real-time status to determine whether the immediate effect of the current work behavior on the current work object meets the expected process effect; if not, the behavior difference information includes effect deviation information. The step of generating and outputting guidance instructions based on the behavioral difference information to guide the operator to perform the target operation includes: When the behavioral difference information is the effect deviation information, the potential behavioral root causes that lead to the effect deviation are inferred based on the historical data model; and the guidance instruction containing the operation points for correcting the potential behavioral root causes is generated based on the inferred potential behavioral root causes.

2. The intelligent recognition and guidance method for work behavior according to claim 1, characterized in that, The process of identifying the operator's current work behavior based on the multimodal perception information specifically includes: The visual information in the multimodal perception information is analyzed in a temporal sequence to extract the continuous motion trajectory of the operator's limb key points; based on the continuous motion trajectory, the operator's action units under the preset work sequence constraints are identified; the continuous action units are combined and parsed to form the current work behavior with work semantics.

3. The intelligent recognition and guidance method for work behavior according to claim 2, characterized in that, The target operation behavior includes a multi-target behavior sequence arranged in the order of the operation steps; the matching analysis between the current operation behavior and the predefined target operation behavior includes: determining the target operation node corresponding to the current operation behavior in the multi-target behavior sequence; determining whether the current operation behavior is completed within the expected time window of the target operation node; if it is not completed within the expected time window, the behavior difference information includes timing deviation information.

4. The intelligent identification and guidance method for work behavior according to claim 1, characterized in that, The method also includes an adaptive learning step: Collect the multimodal perception information generated by the operator during multiple operations, the corresponding current operation behavior, the behavior difference information, and the execution result data of the guidance instructions; Based on the execution result data, the decision model used to generate the guidance instructions is incrementally optimized to reduce the behavioral differences generated in subsequent operations for the same operator.

5. The intelligent recognition and guidance method for work behavior according to claim 4, characterized in that, The incremental learning optimization of the decision model used to generate the guidance instructions based on the execution result data includes: Based on the operator's historical skill proficiency tags, personalized guidance strategy sub-models are established and updated for different skill proficiency groups. When a specific operator is identified, the personalized guidance strategy sub-model corresponding to the operator's skill proficiency group is invoked to generate guidance instructions adapted to the operator's habits.

6. The intelligent identification and guidance method for work behavior according to claim 1, characterized in that, The method also includes an autonomous exception handling step: When the current operation behavior is identified as belonging to a predefined abnormal behavior category, the abnormal analysis process is triggered; The anomaly analysis process includes: predicting potential operational defects caused by the abnormal behavior based on a knowledge base, and automatically generating preliminary process parameter adjustment plans. The process parameter adjustment scheme is output to the relevant execution equipment in an attempt to compensate for the impact of the abnormal behavior.

7. The intelligent recognition and guidance method for work behavior according to claim 6, characterized in that, The anomaly analysis process also includes: After outputting the process parameter adjustment scheme, the multimodal perception information is continuously monitored to evaluate the improvement effect of the process parameter adjustment scheme on the operation results; the abnormal behavior, the process parameter adjustment scheme and the improvement effect are associated and stored to form closed-loop knowledge data for optimizing subsequent anomaly analysis and decision-making.

8. A work behavior intelligent recognition and guidance system, characterized in that, The method for intelligent identification and guidance of work behavior according to any one of claims 1-7 includes: The sensing module is used to collect multimodal sensing information about the working environment; The behavior recognition module, connected to the perception module, is used to identify the operator's current work behavior based on the multimodal perception information; The analysis and decision-making module, connected to the behavior recognition module, is used to match and analyze the current operation behavior with the predefined target operation behavior to generate behavior difference information, and generate guidance instructions based on the behavior difference information; The guidance output module, connected to the analysis and decision module, is used to output the guidance instructions to control the output rhythm of the action parameters or interface prompts of at least one auxiliary execution device, thereby guiding the operator to perform the target operation.