Intelligent assembly guiding and error correcting system based on man-machine cooperation
By recognizing the operator's cognitive intent in real time and generating an appropriate human-machine collaboration strategy, combined with augmented reality technology and adaptive grasping methods, the problem of rigid human-machine collaboration strategies in existing technologies is solved, and efficient complex assembly collaboration is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAIAN SENIOR VOCATIONAL & TECH SCHOOL
- Filing Date
- 2026-03-17
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, the lack of real-time perception and response to the operator's cognitive intentions leads to rigid human-machine collaboration strategies and insufficient adaptability. In particular, in non-standard or complex assembly scenarios, the guidance information does not match the intention of human operation, which affects the efficiency of collaboration.
By collecting and analyzing the operator's multimodal physiological and behavioral signals in real time, the robot dynamically identifies cognitive intentions, generates appropriate human-machine collaboration strategies, and uses augmented reality technology to provide guidance information. The robot's end effector adaptively grasps the workpiece and collaboratively completes precision alignment and assembly actions. At the same time, it monitors key process parameters in real time and provides early warnings of anomalies.
This upgrade from one-way programmed guidance to two-way adaptive collaboration enhances the intelligence and efficiency of the assembly process, ensuring that operators and robots work together efficiently to complete complex assembly tasks.
Smart Images

Figure CN122008221A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-machine collaborative intelligent assembly, and in particular to an intelligent assembly guidance and error correction system based on human-machine collaboration. Background Technology
[0002] In the field of human-machine collaborative intelligent assembly, existing technologies can provide visual guidance through augmented reality and utilize collaborative robots to complete basic collaborative tasks such as assisted handling. However, existing systems generally lack the ability to perceive the operator's real-time cognitive state, and their collaborative strategies are often fixed or pre-programmed, unable to dynamically adjust according to the operator's actual understanding and operational intentions. This results in less natural and smooth human-machine interaction, and when dealing with non-standard or complex assembly scenarios, mismatches or even mutual interference may occur between guidance information and human operational intentions, limiting further improvements in collaborative efficiency and adaptability.
[0003] This invention aims to address the core problem in existing technologies: the lack of real-time perception and response to operator cognitive intentions leads to rigid human-machine collaboration strategies and insufficient adaptive capabilities. By collecting and analyzing the operator's multimodal physiological and behavioral signals in real time, dynamically identifying cognitive intentions, and generating optimal human-machine collaboration strategies and guidance methods accordingly, this system achieves a paradigm shift from one-way programmed guidance to two-way adaptive collaboration. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides an intelligent assembly guidance and error correction system based on human-machine collaboration to solve the problem of insufficient human-machine collaboration adaptive capability caused by the inability to perceive and respond to the operator's cognitive intentions in real time.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides an intelligent assembly guidance and error correction system based on human-machine collaboration, which includes a task instantiation module that generates assembly task instances aligned with the physical environment based on the assembly order and environmental perception information. The cognitive intent state module collects and processes multimodal physiological and behavioral signals of operators in real time to identify cognitive intent states related to the current assembly task. The human-machine collaboration strategy module dynamically generates an appropriate human-machine collaboration strategy based on the recognized cognitive intent state and the process requirements of the current assembly task instance. The operation mode pre-configuration module generates operator-oriented guidance information based on the human-machine collaboration strategy and pre-configures the robot's operation mode accordingly. The assembly module controls the robot's end effector to grip the workpiece in an adaptive manner and physically buffers the initial pose error during the gripping process. Under the guidance of the human-machine collaboration strategy, the operator and the robot work together to complete the precise alignment and assembly of the workpiece. The experience fragment module monitors key process parameters in real time and provides early warnings of anomalies during collaborative assembly, performs multi-dimensional quality compliance checks on assembled components, and encapsulates the entire collaborative assembly process data into experience fragments.
[0007] As a preferred embodiment of the intelligent assembly guidance and error correction system based on human-machine collaboration described in this invention, the following steps are included: Based on the assembly order and combined with environmental perception information, generating an assembly task instance aligned with the physical environment. The industrial edge computing server calls the product digital twin model according to the assembly order and drives the high-resolution 3D vision sensor to acquire color point cloud data of the workbench environment; Industrial edge computing servers register and label product digital twin models with color point cloud data, and combine standard assembly sequences and process parameters to generate assembly task instances aligned with the physical environment.
[0008] As a preferred embodiment of the intelligent assembly guidance and error correction system based on human-machine collaboration described in this invention, the system includes the following steps: real-time acquisition and processing of the operator's multimodal physiological and behavioral signals, and identification of the cognitive intention state related to the current assembly task. Wearable biosignal acquisition headbands, augmented reality glasses, and directional microphone arrays collect multimodal physiological and behavioral signals from operators in real time, which are then processed and fused by industrial edge computing servers; The deep learning model identifies the cognitive intent state related to the current assembly task.
[0009] As a preferred embodiment of the intelligent assembly guidance and error correction system based on human-machine collaboration described in this invention, the following steps are included: dynamically generating an appropriate human-machine collaboration strategy based on the recognized cognitive intent state and the process requirements of the current assembly task instance: Industrial edge computing servers extract current process requirements from assembly task instances aligned with the physical environment; The cognitive intent state is input into the dynamic decision engine for evaluation and matching with the process requirements of the current sub-step. Based on the evaluation and matching results, the dynamic decision engine outputs an appropriate human-machine collaboration strategy.
[0010] As a preferred embodiment of the intelligent assembly guidance and error correction system based on human-machine collaboration described in this invention, the following steps are included: generating operator-oriented guidance information according to the human-machine collaboration strategy and pre-configuring the robot's corresponding operating mode: The industrial edge computing server extracts relevant guidance content from assembly task instances aligned with the physical environment based on human-machine collaboration strategies, drives the augmented reality rendering engine to convert the guidance content into virtual-real fusion visual cues, and the augmented reality glasses project the virtual-real fusion visual cues into the operator's field of vision. Based on the human-machine collaboration strategy, robot control instructions are generated, and the robot controller adjusts the collaborative robot to the specified operating mode according to the robot control instructions.
[0011] As a preferred embodiment of the intelligent assembly guidance and error correction system based on human-machine collaboration described in this invention, the following steps are included: controlling the robot's end effector to adaptively grasp the workpiece and physically buffering the initial pose error during the grasping process: The robot controller drives the robot body to move to the workpiece gripping point. The intelligent end effector receives the gripping command and contacts the workpiece surface in a low-stiffness mode. The flexible contact surface array of the intelligent end effector generates adaptive deformation according to the contact pressure distribution to envelop the workpiece. During the enveloping process of the workpiece, the intelligent end effector buffers the initial pose error through adaptive deformation physical properties. After the gripping is stable, the intelligent end effector switches the local area of the contact surface to a high stiffness mode according to the process requirements.
[0012] As a preferred embodiment of the intelligent assembly guidance and error correction system based on human-machine collaboration described in this invention, wherein: under the guidance of the human-machine collaboration strategy, the operator and the robot collaboratively complete the precise alignment and assembly of the workpiece, including the following steps: When the human-machine collaboration strategy is in the human-led - machine compliant assistance mode, the operator manually guides the collaborative robot to move the workpiece to the rough assembly position. The collaborative robot's hand-eye vision system and six-dimensional force / torque sensor collect alignment deviation data and contact force data in real time. Based on the alignment deviation data and contact force data, the robot end effector is calculated to provide fine-tuning compensation instructions. The robot controller then executes these instructions to assist in completing precise alignment and assembly operations.
[0013] As a preferred embodiment of the intelligent assembly guidance and error correction system based on human-machine collaboration described in this invention, the system includes the following steps: Real-time monitoring and anomaly warning of key process parameters during collaborative assembly execution. The collaborative robot's six-dimensional force / torque sensor collects contact force and torque data in real time during the assembly process; the collaborative robot's hand-eye vision system collects image and position data in real time during the assembly process; and the directional microphone array collects sound data during the assembly process. The industrial edge computing server integrates and analyzes contact force and torque data, image and location data, and compares them with standard process parameters. When any parameter in the contact force and torque data, image and location data, or sound data exceeds a preset threshold, it triggers an anomaly warning.
[0014] As a preferred embodiment of the intelligent assembly guidance and error correction system based on human-machine collaboration described in this invention, the system includes: performing multi-dimensional quality conformity inspection on the assembled components, comprising the following steps: The assembled components are 3D scanned to obtain point cloud data. The obtained point cloud data is then precisely registered with the standard 3D model in the product's digital twin model. The industrial edge computing server calculates the assembly gap value of key dimensions based on the registration result. The assembly gap values are compared with the tolerance range specified in the standard operating procedure. When the industrial edge computing server determines that all assembly gap values fall within the tolerance range, the assembly quality conformity inspection is confirmed to have passed.
[0015] As a preferred embodiment of the intelligent assembly guidance and error correction system based on human-machine collaboration described in this invention, the entire chain data of this collaborative assembly process is encapsulated into experience fragments, including the following steps: Industrial edge computing servers collect multi-source time-series data throughout the entire process, from generating assembly task instances aligned with the physical environment to determining assembly quality compliance. The multi-source time-series data is then timestamped and structured to form structured process data. The structured process data is associated and labeled with the assembly quality compliance inspection results. The industrial edge computing server then packages the associated and labeled structured process data into a full-link data experience fragment of the collaborative assembly process in a predefined encapsulation format.
[0016] The beneficial effects of this invention are as follows: the task instantiation module generates assembly task instances that are precisely aligned with the physical environment; the cognitive intent state module collects and analyzes the operator's multimodal physiological and behavioral signals in real time to identify their cognitive intent; the human-machine collaboration strategy module dynamically generates an adapted collaboration strategy based on the intent and process requirements; the operation mode pre-configuration module generates guidance information and pre-configures the robot mode accordingly; the assembly module controls the intelligent end effector to grasp the workpiece in an adaptive manner and physically buffers the initial pose error during the grasping process, thereby collaboratively completing precision assembly under the guidance of the collaboration strategy; and the experience fragment module monitors and inspects the entire process and encapsulates the full-link data into experience fragments, thus realizing the upgrade from one-way programmed guidance to two-way adaptive collaboration, effectively improving the intelligence level and collaboration efficiency of the assembly process. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of an intelligent assembly guidance and error correction system based on human-machine collaboration. Detailed Implementation
[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0020] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0021] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0022] Reference Figure 1 This is one embodiment of the present invention, which provides an intelligent assembly guidance and error correction system based on human-machine collaboration, comprising the following steps: The task instantiation module generates assembly task instances that are aligned with the physical environment, based on the assembly order and environmental awareness information.
[0023] The industrial edge computing server calls up the product digital twin model based on the assembly order and drives a high-resolution 3D vision sensor to acquire color point cloud data of the workbench environment.
[0024] Furthermore, the industrial edge computing server receives assembly orders from the production management system. These orders contain the model and batch information of the products to be assembled. Based on the product model in the assembly order, the industrial edge computing server retrieves and calls the corresponding product digital twin model and standard operating procedure (SOP) from the connected process knowledge base server. The product digital twin model is a three-dimensional virtual model containing the geometry, assembly relationships, and physical attributes of all components. The SOP defines the assembly sequence (standard assembly sequence) and the process parameters for each step, such as tightening torque or pressing force range, in structured data format. When calling the product digital twin model... Simultaneously with the creation of the digital twin model and standard operating procedure, the industrial edge computing server sends a trigger command to the high-resolution 3D vision sensor. The high-resolution 3D vision sensor is fixedly installed above the assembly station. After receiving the command, the high-resolution 3D vision sensor quickly scans the workbench area below. The high-resolution 3D vision sensor uses structured light or time-of-flight principles to acquire the three-dimensional spatial coordinates and color information of all objects on the workbench surface, including parts in the hopper, fixtures, and tools, thereby generating color point cloud data of the workbench environment. Color point cloud data is a set of dense spatial points that can characterize the surface morphology of the physical environment.
[0025] Industrial edge computing servers register and label product digital twin models with color point cloud data, and combine standard assembly sequences and process parameters to generate assembly task instances aligned with the physical environment.
[0026] Furthermore, the industrial edge computing server uses point cloud registration algorithms to process the product's digital twin model and color point cloud data. Point cloud registration algorithms, such as the iterative nearest-point algorithm, first extract the theoretical 3D features of the part to be assembled from the product's digital twin model, and simultaneously extract the corresponding physical 3D features from the color point cloud data. Through iterative calculation, the optimal spatial transformation matrix is found to minimize the overall distance error between the theoretical and actual features, thus completing the spatial alignment, or registration, between the product's digital twin model and the color point cloud data. Based on the completed registration, the industrial edge computing server, according to the registration relationship, assigns each virtual part instance in the product's digital twin model... The precise location and orientation of the virtual geometric entity in the physical space corresponding to the color point cloud data are marked, and this annotation information binds the virtual geometric entity to the real-world object one by one. Subsequently, the industrial edge computing server associates and merges the product digital twin model with the marked spatial location with the standard assembly sequence and process parameters parsed from the standard operating procedure. The standard assembly sequence specifies the sequential logic of the parts assembly, while the process parameters define the quality standards of each assembly action. The merged data entity constitutes a complete, executable assembly task instance aligned with the physical environment. This assembly task instance can accurately guide the specific execution of each subsequent assembly operation in the real space.
[0027] The cognitive intent state module collects and processes multimodal physiological and behavioral signals of operators in real time to identify cognitive intent states related to the current assembly task.
[0028] Wearable biosignal acquisition headbands, augmented reality glasses, and directional microphone arrays collect multimodal physiological and behavioral signals from operators in real time, which are then processed and fused by industrial edge computing servers.
[0029] Furthermore, the wearable biosignal acquisition headband worn by the operator incorporates multiple dry electrode contact points. These contact points are in close contact with the operator's forehead and scalp to continuously acquire electroencephalogram (EEG) signals, which reflect the electrophysiological activity of the cerebral cortex. Simultaneously, a miniature infrared camera is integrated into the augmented reality glasses frame to form an eye-tracking component. This camera tracks the operator's pupil movements and calculates the coordinates of the gaze's focal point in the real world, generating gaze focus data. A directional microphone array installed at the workstation captures the operator's voice commands or natural language descriptions during operations, forming speech signals. The wearable biosignal acquisition headband, augmented reality glasses, and directional microphone array all connect to an industrial edge computing server via a wireless network, synchronizing the acquired EEG signals, gaze focus data, and speech signals with millisecond-level timestamps. The audio signals are transmitted to the industrial edge computing server in real time. After receiving these multimodal physiological and behavioral signals, the industrial edge computing server first performs bandpass filtering on the EEG signals to remove power frequency interference and electromyographic noise, performs coordinate transformation on the gaze focus data to unify it to the world coordinate system defined by the assembly task instance, and performs noise reduction and endpoint detection on the speech signals. The preprocessed EEG signals, gaze focus data and speech signals are sent to a feature fusion unit. The feature fusion unit extracts the power spectral density of the EEG signals in a specific frequency band as a feature vector, extracts the dwell time and saccade path of the gaze focus on the key virtual assembly guidance area as a feature vector, and extracts keywords related to doubt or confirmation in the speech signals as feature vectors. All these feature vectors are aligned and spliced according to time windows to form fused multidimensional time series feature data.
[0030] The deep learning model identifies the cognitive intent state related to the current assembly task.
[0031] Furthermore, a deep learning model is pre-trained in the industrial edge computing server. This deep learning model can be, for example, a combination of a long short-term memory network and an attention mechanism. The industrial edge computing server inputs the fused multi-dimensional time-series feature data into the deep learning model for forward inference. The deep learning model first captures the dynamic change patterns of the feature data in the time dimension through the long short-term memory network layer, and then automatically focuses on the feature fragments most relevant to the current assembly task context through the attention mechanism layer. The output layer of the deep learning model is a softmax classifier, which maps the learned high-level feature representation to a probability distribution of a series of discrete categories. These discrete categories are predefined as cognitive intention states closely related to the assembly task, such as a state of focusing on understanding guidance, a state of confusion and hesitation about the current step, a state of actively trying to explore alternatives, or a state of distraction due to fatigue. The deep learning model calculates the confidence probability of the operator being in each cognitive intention state in real time based on the input multi-dimensional time-series feature data. The industrial edge computing server selects the category with the highest confidence probability as the recognition result at the current moment, thereby completing the recognition of the cognitive intention state related to the current assembly task.
[0032] The human-machine collaboration strategy module dynamically generates an appropriate human-machine collaboration strategy based on the recognized cognitive intent state and the process requirements of the current assembly task instance.
[0033] Industrial edge computing servers extract current process requirements from assembly task instances aligned with the physical environment.
[0034] Furthermore, the industrial edge computing server aligns the current progress index of the assembly process with the physical environment to identify assembly task instances. These physically aligned assembly task instances store standard assembly sequences and process parameters. The industrial edge computing server parses the physically aligned assembly task instances, determines the currently executing or soon-to-be-executed sub-steps from the standard assembly sequences, and extracts the specific process requirements associated with these sub-steps from the process parameters. The extracted process requirements are represented in structured data form, which may include the operation type of the current sub-step, such as precision insertion or torque tightening, the three-dimensional spatial coordinates and attitude tolerance range of the operated object, the maximum permissible contact force required for the assembly process, the standard operation duration range, and the complexity level marked for the sub-step in historical data. These specific process requirements constitute the operational constraints and objectives of the current sub-step, providing a clear task context for subsequent dynamic decision-making.
[0035] The cognitive intent state is input into the dynamic decision engine for evaluation and matching with the process requirements of the current sub-step. Based on the evaluation and matching results, the dynamic decision engine outputs an appropriate human-machine collaboration strategy.
[0036] Furthermore, the industrial edge computing server takes the identified cognitive intent state and the extracted process requirements of the current sub-step as input and sends them to the dynamic decision engine. The dynamic decision engine maintains a predefined decision rule base, which consists of a series of condition-action pairs. For example, if the cognitive intent state is focused following and the complexity level of the process requirement is high, then the output human-machine collaboration strategy is a combination of robot master guidance mode and augmented reality detailed guidance. The dynamic decision engine performs an evaluation and matching process by matching the input cognitive intent state and various features of the process requirements with the conditional parts of the decision rule base. This matching process may involve fuzzy logic reasoning or a weighted scoring mechanism. For example, weights are assigned to different categories of cognitive intent states and different levels of complexity of process requirements, and the strategy with the highest score is selected after calculating the comprehensive score; the dynamic decision engine selects and outputs an appropriate human-machine collaboration strategy from a predefined set of strategies based on the evaluation and matching results; the output human-machine collaboration strategy is a structured instruction set that clearly defines the respective roles of the operator and the collaborative robot in the current context, the robot's level of autonomy, the level of detail and presentation of augmented reality guidance information, and the parameter presets for the robot's force control or position control mode. For example, the strategy may specify that the robot performs high-precision visual servo alignment operations, while the augmented reality glasses only need to display a preview of the final assembly result.
[0037] The operation mode pre-configuration module generates operator-oriented guidance information based on the human-machine collaboration strategy and pre-configures the robot's operation mode accordingly.
[0038] The industrial edge computing server extracts relevant guidance content from assembly task instances aligned with the physical environment based on human-machine collaboration strategies, drives the augmented reality rendering engine to convert the guidance content into virtual-real fusion visual cues, and the augmented reality glasses project the virtual-real fusion visual cues into the operator's field of vision.
[0039] Furthermore, the industrial edge computing server parses the received human-machine collaboration strategy, which specifies the type and level of detail of the guidance information required for the current step. Based on the instructions of the human-machine collaboration strategy, the industrial edge computing server accesses the assembly task instance aligned with the physical environment, extracting the guidance content corresponding to the current sub-step from the physically aligned assembly task instance. This guidance content includes, but is not limited to, the highlighted outline of the 3D model of the part to be grasped or assembled, a dynamic path animation arrow from the current position to the target position, a transparent perspective view of key mating surfaces, text labels for the torque values to be read, and a virtual button for operation confirmation. The industrial edge computing server then integrates the extracted guidance content with high-resolution 3D visual sensing... The device combines real-time acquired physical environment spatial coordinates to drive an augmented reality rendering engine running on an industrial edge computing server. The augmented reality rendering engine uses 3D graphics algorithms to perform virtual-real registration calculations based on the spatial coordinates, ensuring that the virtual guidance content can be accurately superimposed on the corresponding objects or locations in the real world. Then, the calculated superimposed image data is encoded into a video stream. The encoded video stream is transmitted wirelessly to the augmented reality glasses worn by the operator. The optical display module built into the augmented reality glasses receives the video stream and decodes it. Finally, the virtual guidance content is projected into the operator's real field of vision with the correct spatial position and posture through a semi-transparent lens, forming a virtual-real fusion visual cues that the operator can directly observe and understand.
[0040] Based on the human-machine collaboration strategy, robot control instructions are generated, and the robot controller adjusts the collaborative robot to the specified operating mode according to the robot control instructions.
[0041] Furthermore, while generating guidance information, the industrial edge computing server generates robot control instructions based on the same human-robot collaboration strategy. This strategy defines the operating modes that the collaborative robot should adopt, such as a high-precision position servo mode for robot-led precision alignment, an admittance control mode allowing the operator to apply force to guide the robot's movement, or a zero-force drag mode for the operator to manually pull the robot. The industrial edge computing server encapsulates the selection of the operating mode and corresponding control parameters, such as servo stiffness, damping coefficient, and maximum speed limit, into structured robot control instructions. These instructions are sent to the robot controller via a real-time industrial Ethernet network. The robot controller, the core control unit of the collaborative robot, receives and parses the control instructions, calls the corresponding internal control algorithm according to the operating mode specified in the instructions, and loads the relevant control parameters into the joint servo drivers of the collaborative robot. The joint servo drivers adjust their response characteristics according to the new control algorithm and parameters, ensuring that the overall dynamic behavior of the collaborative robot conforms to the requirements of the human-robot collaboration strategy. This completes the pre-configuration of the collaborative robot's operating mode, preparing it to execute subsequent collaborative assembly actions.
[0042] The assembly module controls the robot's end effector to grip the workpiece in an adaptive manner and physically buffers the initial pose error during the gripping process. Under the guidance of the human-machine collaboration strategy, the operator and the robot work together to complete the precise alignment and assembly of the workpiece.
[0043] The robot controller drives the robot body to move to the workpiece gripping point. The intelligent end effector receives the gripping command and contacts the workpiece surface in a low-stiffness mode. The flexible contact surface array of the intelligent end effector generates adaptive deformation according to the contact pressure distribution to envelop the workpiece.
[0044] Furthermore, the robot controller analyzes the coordinates of the grasping target from the industrial edge computing server, driving the servo motors of each joint of the robot body to move in coordination, enabling the intelligent end effector installed at the end of the robot to move precisely above the workpiece grasping point. After receiving the grasping command forwarded by the robot controller, the intelligent end effector's internal micro-electromagnetic actuator or electrorheological fluid control unit applies a low electric or magnetic field, causing the flexible contact surface array constituting the gripping surface of the intelligent end effector to be in a state of low elastic modulus and high viscosity coefficient, i.e., low stiffness mode. In low stiffness mode, the flexible contact surface array begins to close and make contact with the workpiece surface. The micro-pressure sensor network distributed below the flexible contact surface array senses and provides feedback on the spatial distribution of contact pressure in real time. Because the flexible contact surface array material has high flexibility in the low stiffness state, each unit in the array will deform non-uniformly according to the local contact pressure it senses. This deformation allows the overall shape of the flexible contact surface array to actively adapt to the actual contour of the workpiece, thereby achieving adaptive envelope for irregular or tilted workpieces.
[0045] During the enveloping process of the workpiece, the intelligent end effector buffers the initial pose error through adaptive deformation physical properties. After the gripping is stable, the intelligent end effector switches the local area of the contact surface to a high stiffness mode according to the process requirements.
[0046] Furthermore, during the adaptive envelopment of the workpiece by the flexible contact surface array, an initial pose error exists between the actual workpiece pose and the theoretical gripping pose, such as a slight offset or tilt. This error leads to uneven distribution of contact pressure on the flexible contact surface array. The non-uniform deformation of the flexible contact surface array based on the pressure distribution allows for slight relative sliding and attitude adjustment between the intelligent end effector and the workpiece at the microscopic level. This process of absorbing relative displacement and angle through the material's own compliance is the mechanism for physically buffering the initial pose error. When the pressure sensor network detects that the contact pressure distribution tends to stabilize and reaches the preset gripping force threshold, it is determined that the gripping is stable. Subsequently, the intelligent end effector, based on the current process requirements obtained from the assembly task instance aligned with the physical environment, such as the need for high-precision insertion operations, changes the control signal applied to specific area units of the flexible contact surface array. This switches the material state of these local areas in contact with the key feature surfaces of the workpiece to a high elastic modulus state, i.e., a high stiffness mode, while other non-critical contact areas can maintain low stiffness to provide damping, thereby forming a reliable force transmission path with local rigidity between the intelligent end effector and the workpiece.
[0047] When the human-machine collaboration strategy is in the human-led - machine compliant assistance mode, the operator manually guides the collaborative robot to move the workpiece to the rough assembly position. The collaborative robot's hand-eye vision system and six-dimensional force / torque sensor collect alignment deviation data and contact force data in real time.
[0048] Furthermore, when the human-machine collaboration strategy is explicitly specified as a human-led, machine-compliant assist mode, the joint servo drives of the collaborative robot switch to a low-impedance state. The operator directly grasps the workpiece firmly held by the intelligent end effector or holds the traction handle equipped on the end of the collaborative robot, applying force to manually guide the collaborative robot and the workpiece to move towards the rough assembly position. During this movement, the hand-eye vision system fixedly installed on the end flange of the collaborative robot continuously captures feature images of the workpiece and the target assembly position. The image processing algorithm calculates in real time the position and angle deviation of the workpiece relative to the target position in three-dimensional space, i.e., the alignment deviation data. At the same time, the six-dimensional force / torque sensor installed between the end flange of the collaborative robot and the intelligent end effector measures and outputs in real time the force and torque vectors applied by the operator, as well as the contact force data that may occur between the workpiece and the surrounding environment.
[0049] Based on the alignment deviation data and contact force data, the robot end effector is calculated to provide fine-tuning compensation instructions. The robot controller then executes these instructions to assist in completing precise alignment and assembly operations.
[0050] Furthermore, the industrial edge computing server receives alignment deviation data transmitted in real time from the hand-eye vision system and contact force data transmitted in real time from the six-dimensional force / torque sensor. The industrial edge computing server runs an impedance control or admittance control algorithm, which uses the alignment deviation data as the position correction target and the contact force data as the compliance constraint to calculate a series of tiny, continuous robot end-effector pose adjustments, i.e., robot end-effector fine-tuning compensation commands. The robot end-effector fine-tuning compensation commands are sent to the robot controller through a real-time communication network. The robot controller parses these robot end-effector fine-tuning compensation commands, converts them into incremental motion commands for each joint motor, and drives the joint servo motors to execute them. The joints of the collaborative robot generate tiny, compliant movements accordingly, thereby assisting the operator in accurately aligning the workpiece and placing it into the target assembly position, realizing precise alignment and assembly actions under human-machine collaboration.
[0051] The experience fragment module monitors key process parameters in real time and provides early warnings of anomalies during collaborative assembly, performs multi-dimensional quality compliance checks on assembled components, and encapsulates the entire collaborative assembly process data into experience fragments.
[0052] The collaborative robot's six-dimensional force / torque sensor collects contact force and torque data in real time during the assembly process; the collaborative robot's hand-eye vision system collects image and position data in real time during the assembly process; and the directional microphone array collects sound data during the assembly process.
[0053] Furthermore, during the continuous precision alignment and assembly operations performed by the collaborative robot, a six-dimensional force / torque sensor installed at the robot's end flange continuously measures and outputs the force components and torque components of the assembly contact interface in three axes at a high frequency. These dynamically changing six-dimensional data constitute the contact force and torque data. Simultaneously, a hand-eye vision system fixed at the robot's end captures image sequences of the workpiece and the assembly target area at a fixed frame rate, and uses visual algorithms to calculate the workpiece's three-dimensional spatial coordinates and attitude angles in real time from the image sequences. This information constitutes image and position data. A directional microphone array deployed around the workstation synchronously collects audio signals generated in the assembly environment, especially the characteristic sounds generated when the workpiece contacts, presses in, or tightens. These audio signals constitute sound data.
[0054] The industrial edge computing server integrates and analyzes contact force and torque data, image and location data, and compares them with standard process parameters. When any parameter in the contact force and torque data, image and location data, or sound data exceeds a preset threshold, it triggers an anomaly warning.
[0055] Furthermore, the industrial edge computing server synchronously receives contact force and torque data from a six-dimensional force / torque sensor, image and position data from a hand-eye vision system, and sound data from a directional microphone array via a high-speed data bus. The industrial edge computing server runs a multi-sensor data fusion algorithm, first aligning the contact force and torque data, image and position data, and sound data in a unified time coordinate system. Then, it extracts the standard process parameters corresponding to the current assembly step from the assembly task instance aligned with the physical environment. These standard process parameters define, for example, the acceptable range of pressing force, the maximum allowable value of alignment deviation, and the baseline characteristics of the sound spectrum. The fusion analysis process includes calculating the real-time contact force curve. The peak value and slope of the line are compared with a standard range to calculate the deviation between the real-time position and the theoretical path. Correlation analysis is performed on the real-time sound signal after performing a fast Fourier transform and comparing it with the reference spectrum. The industrial edge computing server sets preset thresholds for each type of parameter, such as the maximum allowable contact force threshold, position deviation threshold, and sound spectrum difference threshold. Once the real-time analysis results show that any measurement value of the contact force and torque data, image and position data, or sound data exceeds its corresponding preset threshold, the industrial edge computing server immediately generates an anomaly warning signal containing the anomaly type, timestamp, and severity level, and issues the anomaly warning signal through an audible and visual alarm device or the operator's augmented reality glasses interface.
[0056] The assembled components are 3D scanned to obtain point cloud data. The obtained point cloud data is then precisely registered with the standard 3D model in the product's digital twin model. The industrial edge computing server calculates the assembly gap value of key dimensions based on the registration result.
[0057] Furthermore, after the assembly action is confirmed, the collaborative robot controls the hand-eye vision system to perform a fine, multi-angle 3D scan of the assembled components. The hand-eye vision system uses structured light or laser triangulation principles to obtain the coordinates of dense spatial points on the component surface, thereby generating high-precision point cloud data representing the actual shape of the component. After acquiring this point cloud data, the industrial edge computing server calls the point cloud registration algorithm to match the acquired point cloud data with the standard 3D model of the corresponding component in the product's digital twin model under ideal conditions. The point cloud registration algorithm finds the optimal spatial transformation through iterative calculation to minimize the overall distance error between the actual point cloud data and the standard 3D model, thereby achieving accurate registration between the two. Based on the accurate registration result, the industrial edge computing server calculates the actual distance between key mating parts, such as specific measurement point pairs between the mating surfaces of two parts, in 3D space according to a predefined detection program. These calculated distance values are the assembly gap values.
[0058] The assembly gap values are compared with the tolerance range specified in the standard operating procedure. When the industrial edge computing server determines that all assembly gap values fall within the tolerance range, the assembly quality conformity inspection is confirmed to have passed.
[0059] Furthermore, the industrial edge computing server reads the tolerance ranges for assembly gap values of each critical dimension specified for the current assembly step from the standard operating procedure associated with the assembly task instance aligned with the physical environment. The tolerance ranges define the acceptable intervals in the form of minimum and maximum values. The industrial edge computing server compares each calculated assembly gap value with its corresponding tolerance range to determine whether each assembly gap value is greater than or equal to the minimum value and less than or equal to the maximum value of its tolerance range. Only when all calculated assembly gap values meet the condition of falling within their respective tolerance ranges does the industrial edge computing server determine that the assembly quality conformity inspection of the current component has passed, and generates a Boolean logic flag indicating that the inspection has passed.
[0060] Industrial edge computing servers collect multi-source time-series data throughout the entire process, from generating assembly task instances aligned with the physical environment to determining assembly quality compliance. The multi-source time-series data is then timestamped and structured to form structured process data.
[0061] Furthermore, the industrial edge computing server initiates a data collection process, which begins with the initial event of generating an assembly task instance aligned with the physical environment and continues recording until the assembly quality compliance inspection result is generated. During this process, the industrial edge computing server collects and caches multi-source time-series data from all participating stages. This multi-source time-series data includes color point cloud data generated by high-resolution 3D vision sensors, EEG signal data generated by wearable biosignal acquisition headbands, gaze focus data generated by augmented reality glasses, cognitive intent state recognition results generated by the industrial edge computing server itself, human-machine collaboration strategy instructions, robot control instructions, contact force and torque data generated by six-dimensional force / torque sensors, image and position data generated by hand-eye vision systems, sound data generated by directional microphone arrays, and assembly gap value calculation results. The industrial edge computing server uses a global high-precision clock to assign a unified timestamp to all incoming multi-source time-series data, and then cleans, sorts, and converts these timestamped data streams according to a predefined data structure template to form a structured process data document containing a timeline, data source identifiers, and specific values.
[0062] The structured process data is associated and labeled with the assembly quality compliance inspection results. The industrial edge computing server then packages the associated and labeled structured process data into a full-link data experience fragment of the collaborative assembly process in a predefined encapsulation format.
[0063] Furthermore, the industrial edge computing server reads the structured process data document and the final assembly quality compliance inspection result. The server adds a field to the metadata header of the structured process data document to record and associate the final assembly quality compliance inspection result, such as marking the final result of this assembly task as "qualified" or "unqualified," and can also associate any triggered abnormal warning information. After completing the association and annotation, the industrial edge computing server serializes and packages the entire annotated structured process data according to a predefined, universal data encapsulation format, such as JSON or Protocol Buffers. This serialized data package contains end-to-end information from perception, decision-making, execution to inspection, forming a complete, independently storable, retrieval, and analyzable end-to-end data experience fragment of the collaborative assembly process.
[0064] In summary, this invention generates assembly task instances that are precisely aligned with the physical environment through a task instantiation module, collects and analyzes multimodal physiological and behavioral signals of operators in real time to identify their cognitive intentions through a cognitive intent state module, dynamically generates an adaptive collaboration strategy based on the intention and process requirements through a human-machine collaboration strategy module, generates guidance information and pre-configures robot modes based on this information through a pre-configuration module, controls the intelligent end effector to grasp the workpiece in an adaptive manner and physically buffers the initial pose error during the grasping process, and then completes precision assembly collaboratively under the guidance of the collaboration strategy, and the experience fragment module monitors and inspects the entire process and encapsulates the full-link data into experience fragments, thereby realizing an upgrade from one-way programmed guidance to two-way adaptive collaboration, effectively improving the intelligence level and collaboration efficiency of the assembly process.
[0065] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. An intelligent assembly guidance and error correction system based on human-machine collaboration, characterized in that: This includes a task instantiation module, which generates assembly task instances aligned with the physical environment based on assembly orders and environmental awareness information. The cognitive intent state module collects and processes multimodal physiological and behavioral signals of operators in real time to identify cognitive intent states related to the current assembly task. The human-machine collaboration strategy module dynamically generates an appropriate human-machine collaboration strategy based on the recognized cognitive intent state and the process requirements of the current assembly task instance. The operation mode pre-configuration module generates operator-oriented guidance information based on the human-machine collaboration strategy and pre-configures the robot's operation mode accordingly. The assembly module controls the robot's end effector to grip the workpiece in an adaptive manner and physically buffers the initial pose error during the gripping process. Under the guidance of the human-machine collaboration strategy, the operator and the robot work together to complete the precise alignment and assembly of the workpiece. The experience fragment module monitors key process parameters in real time and provides early warnings of anomalies during collaborative assembly, performs multi-dimensional quality compliance checks on assembled components, and encapsulates the entire collaborative assembly process data into experience fragments.
2. The intelligent assembly guidance and error correction system based on human-machine collaboration as described in claim 1, characterized in that: Based on the assembly order and combined with environmental awareness information, an assembly task instance aligned with the physical environment is generated, including the following steps: The industrial edge computing server calls the product digital twin model according to the assembly order and drives the high-resolution 3D vision sensor to acquire color point cloud data of the workbench environment; Industrial edge computing servers register and label product digital twin models with color point cloud data, and combine standard assembly sequences and process parameters to generate assembly task instances aligned with the physical environment.
3. The intelligent assembly guidance and error correction system based on human-machine collaboration as described in claim 2, characterized in that: Real-time acquisition and processing of multimodal physiological and behavioral signals from operators to identify cognitive intention states related to the current assembly task includes the following steps: Wearable biosignal acquisition headbands, augmented reality glasses, and directional microphone arrays collect multimodal physiological and behavioral signals from operators in real time, which are then processed and fused by industrial edge computing servers; The deep learning model identifies the cognitive intent state related to the current assembly task.
4. The intelligent assembly guidance and error correction system based on human-machine collaboration as described in claim 3, characterized in that: Based on the recognized cognitive intent state and the process requirements of the current assembly task instance, an adaptive human-machine collaboration strategy is dynamically generated, including the following steps: Industrial edge computing servers extract current process requirements from assembly task instances aligned with the physical environment; The cognitive intent state is input into the dynamic decision engine for evaluation and matching with the process requirements of the current sub-step. Based on the evaluation and matching results, the dynamic decision engine outputs an appropriate human-machine collaboration strategy.
5. The intelligent assembly guidance and error correction system based on human-machine collaboration as described in claim 4, characterized in that: Based on the human-machine collaboration strategy, guide information for the operator is generated, and the robot is pre-configured with the corresponding operating mode, including the following steps: The industrial edge computing server extracts relevant guidance content from assembly task instances aligned with the physical environment based on human-machine collaboration strategies, drives the augmented reality rendering engine to convert the guidance content into virtual-real fusion visual cues, and the augmented reality glasses project the virtual-real fusion visual cues into the operator's field of vision. Based on the human-machine collaboration strategy, robot control instructions are generated, and the robot controller adjusts the collaborative robot to the specified operating mode according to the robot control instructions.
6. The intelligent assembly guidance and error correction system based on human-machine collaboration as described in claim 5, characterized in that: Controlling the robot's end effector to adaptively grasp a workpiece and physically buffering initial pose errors during the grasping process includes the following steps: The robot controller drives the robot body to move to the workpiece gripping point. The intelligent end effector receives the gripping command and contacts the workpiece surface in a low-stiffness mode. The flexible contact surface array of the intelligent end effector generates adaptive deformation according to the contact pressure distribution to envelop the workpiece. During the enveloping process of the workpiece, the intelligent end effector buffers the initial pose error through adaptive deformation physical properties. After the gripping is stable, the intelligent end effector switches the local area of the contact surface to a high stiffness mode according to the process requirements.
7. The intelligent assembly guidance and error correction system based on human-machine collaboration as described in claim 6, characterized in that: Guided by the human-machine collaboration strategy, the operator and the robot work together to complete the precise alignment and assembly of the workpiece, including the following steps: When the human-machine collaboration strategy is in the human-led - machine compliant assistance mode, the operator manually guides the collaborative robot to move the workpiece to the rough assembly position. The collaborative robot's hand-eye vision system and six-dimensional force / torque sensor collect alignment deviation data and contact force data in real time. Based on the alignment deviation data and contact force data, the robot end effector is calculated to provide fine-tuning compensation instructions. The robot controller then executes these instructions to assist in completing precise alignment and assembly operations.
8. The intelligent assembly guidance and error correction system based on human-machine collaboration as described in claim 7, characterized in that: During collaborative assembly, real-time monitoring and early warning of key process parameters are conducted, including the following steps: The collaborative robot's six-dimensional force / torque sensor collects contact force and torque data in real time during the assembly process; the collaborative robot's hand-eye vision system collects image and position data in real time during the assembly process; and the directional microphone array collects sound data during the assembly process. The industrial edge computing server integrates and analyzes contact force and torque data, image and location data, and compares them with standard process parameters. When any parameter in the contact force and torque data, image and location data, or sound data exceeds a preset threshold, it triggers an anomaly warning.
9. The intelligent assembly guidance and error correction system based on human-machine collaboration as described in claim 8, characterized in that: Perform multi-dimensional quality conformity inspection on assembled components, including the following steps: The assembled components are 3D scanned to obtain point cloud data. The obtained point cloud data is then precisely registered with the standard 3D model in the product's digital twin model. The industrial edge computing server calculates the assembly gap value of key dimensions based on the registration result. The assembly gap values are compared with the tolerance range specified in the standard operating procedure. When the industrial edge computing server determines that all assembly gap values fall within the tolerance range, the assembly quality conformity inspection is confirmed to have passed.
10. The intelligent assembly guidance and error correction system based on human-machine collaboration as described in claim 9, characterized in that: The entire data from this collaborative assembly process is encapsulated into experience fragments, including the following steps: Industrial edge computing servers collect multi-source time-series data throughout the entire process, from generating assembly task instances aligned with the physical environment to determining assembly quality compliance. The multi-source time-series data is then timestamped and structured to form structured process data. The structured process data is associated and labeled with the assembly quality compliance inspection results. The industrial edge computing server then packages the associated and labeled structured process data into a full-link data experience fragment of the collaborative assembly process in a predefined encapsulation format.