Intelligent gear assembly laminating system
By constructing a high-dimensional, robust world model using four-dimensional Gaussian sputtering technology and observable space operator theory, and combining multi-source information fusion and predictive trajectory planning, the robustness and self-optimization problems of visual alignment technology in complex environments are solved, enabling efficient and safe automated assembly of gear components.
Patent Information
- Application Number
- CN202511043534.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-10-21
AI Technical Summary
Existing visual alignment technologies are not robust enough in complex and dynamic industrial environments, cannot effectively integrate information from multiple sensors, lack self-optimization capabilities, and are difficult to cope with unexpected events, resulting in a decline in assembly accuracy and reliability.
A dynamic environment perception layer is constructed using four-dimensional Gaussian sputtering technology. Combined with multi-source heterogeneous information fusion and observable space operator theory, a high-dimensional and robust world model is realized. The system then performs trajectory planning through the cognition and prediction layer and real-time decision-making and execution through the decision and execution layer, forming a closed-loop intelligent control system.
It enables high-precision, safe, and efficient fully automated assembly in dynamic and ever-changing industrial environments, possesses the ability to proactively detect and self-optimize against unknown anomalies, and enhances the robustness and adaptability of the system.
Smart Images

Figure CN120816296A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of gear assembly bonding, and in particular to an intelligent gear assembly bonding system. Background Art
[0002] In the field of automated assembly, existing visual alignment technologies have achieved some success, particularly for the alignment of precision components such as gears. These technologies typically use a 2D industrial camera to capture component images, calculate the component's position and orientation deviation on a plane using a template matching algorithm, and then drive a robotic arm to make corrections. Some advanced solutions even incorporate prior knowledge, such as gear ratios, to establish mathematical models to accurately compensate for angular errors, thereby achieving static alignment.
[0003] However, these mainstream automation solutions are essentially reactive control models, with their technical frameworks based on relatively idealized working environments. When faced with the complexity and dynamics of real industrial sites, their inherent limitations become apparent, primarily in the following areas:
[0004] 1. Single and discrete perception dimensions: The system simplifies the complex assembly process into a series of independent static snapshots. It gradually approaches the target through a "take-calculate-move" cycle, lacking a continuous understanding and time-series modeling of the entire dynamic process (e.g., the robot's motion trajectory, subtle vibrations of the component in the gripper, changes in ambient light, etc.).
[0005] 2. Insufficient environmental robustness: The system's strong reliance on 2D images makes it extremely sensitive to changes in the working environment. Common lighting variations in industrial sites, high light reflections on metal parts, shadows, airborne dust, and subtle differences between parts from different batches can all lead to image feature extraction failures or reduced template matching accuracy, thus affecting alignment reliability.
[0006] 3. Insufficient information utilization: Solutions typically rely solely on visual information, forming an open-loop perception-execution chain. This fails to effectively integrate the robot's own proprioceptive information (such as joint angles and end-torque sensor feedback) with the component's design information (such as 3D computer-aided design models). This creates information silos and limits the system's ability to conduct deeper analysis and decision-making.
[0007] 4. Limited intelligence: Existing technologies solve the pre-defined problem of "how to align," but their control logic is rigid and lacks the ability to respond to unexpected events. For example, when there are loading errors, surface defects on the workpiece, or foreign objects in the work area, the system cannot autonomously identify and handle them, let alone explain its decision-making process or self-optimize.
[0008] Therefore, the core technical bottleneck currently facing the field of precision assembly automation lies in how to upgrade from a "reactive" system that executes fixed programs in an ideal environment to a "cognitive-level" intelligent system that can predict, adapt and self-optimize in a dynamic, changeable and uncertain real industrial environment. Summary of the Invention
[0009] The intelligent gear assembly bonding system provided in this application can realize high-robustness, high-efficiency and high-safety full-process automated assembly of high-precision components represented by gear bonding in a dynamic, changeable and uncertain industrial environment.
[0010] The present application provides an intelligent gear assembly fitting system, which includes: a dynamic environment perception layer, which is used to use four-dimensional Gaussian sputtering technology to model the elements in the entire working area to obtain a four-dimensional Gaussian model, and receive real-time data streams from various sensors to update the four-dimensional Gaussian model in real time during the automated assembly process; a cognitive and prediction layer, which is used to input the initial state and target state of the fitting task into a fitting trajectory prediction model when a fitting task needs to be performed to obtain a predicted fitting trajectory; the fitting trajectory prediction model is obtained based on dynamic linearization training of observable space operator theory; a decision and execution layer, which is used to instruct the robotic arm to perform the fitting task according to the predicted fitting trajectory, and in the process of performing the fitting task, obtain real-time sensor data from the dynamic environment perception layer in real time, make inference decisions when the real-time sensor data is abnormal, and request the cognitive and prediction layer to re-plan a new predicted fitting trajectory when a new decision is generated.
[0011] Among them, the dynamic environment perception layer includes: an initialization module, which is used to use one or more cameras to perform multi-perspective scanning of the entire working area when the intelligent gear assembly bonding system is first deployed or the work station is changed, to obtain a multi-perspective image stream, and based on the multi-perspective image stream, use four-dimensional Gaussian sputtering technology to model the elements in the entire working area to obtain a four-dimensional Gaussian model; an operation module, which is used to receive real-time data streams from various sensors to update the four-dimensional Gaussian model during the automated assembly process; wherein, the camera image stream is used to correct the color and position of the Gaussian basis element in the four-dimensional Gaussian model, and the motion data of the robotic arm is used to drive the Gaussian basis element representing the robotic arm part in the four-dimensional Gaussian model to move.
[0012] Among them, the dynamic environment perception layer also includes: a visual data acquisition module, which is used to collect original images using single-frame high dynamic range imaging technology and perform real-time enhancement on the original images using a neural network model. The enhanced original images are used to update the four-dimensional Gaussian model in real time, and the four-dimensional Gaussian model is updated in real time using the data stream collected by the event camera; a collaborative perception module for multi-source heterogeneous information fusion, which is used to unify the visual stream, ontological stream, and prior stream into a collaborative representation space; wherein, the visual stream is provided by the data collected by the HDR camera and the event camera, the ontological stream comes from the joint angle and speed provided by the robot encoder, and the real-time feedback of the force / torque sensor on the end gripper, and the prior stream comes from the original three-dimensional computer-aided design model of the component to be assembled; a cross-platform feature compatibility module, which is used to train a feature encoder for each sensor in the intelligent gear assembly fitting system, and map the original features output by each into a unified, semantically aligned shared latent space; information from different devices but pointing to the same physical feature in the shared latent space can be directly compared and fused.
[0013] Among them, the dynamic environment perception layer also includes: a feature quantization module, which is used to convert the key geometric features extracted from the four-dimensional Gaussian model into a standardized, fixed-length feature vector; the feature vector has error correction capabilities and can tolerate geometric deviations caused by manufacturing tolerances, slight wear or sensor noise.
[0014] Among them, the cognition and prediction layer includes: a learning module, which is used to train the fitting trajectory prediction model using the real or simulated assembly process data provided by the dynamic environment perception layer, so that the fitting trajectory prediction model can map the state of the physical world to a point in a high-dimensional virtual observable space, and learn a finite-dimensional approximation matrix of a Koopman operator, which describes the linear evolution law of the points in the virtual observable space over time in the virtual observable space; a prediction module, which is used to input the initial state and target state of the fitting task into the fitting trajectory prediction model when a fitting task needs to be performed, so as to obtain a predicted fitting trajectory; the fitting trajectory prediction model is obtained based on the dynamic linearization training of the observable space operator theory.
[0015] Among them, the fitting trajectory prediction model includes: an encoder, an evolution unit and a decoder; the prediction module is also used to input the initial state and target state of the fitting task into the encoder when the fitting task needs to be performed, map the initial state into the virtual observable space, and use the evolution unit to parse the complete, continuous and smooth optimal trajectory from the initial state to the target state in the virtual observable space; and use the decoder to decode the optimal trajectory to obtain the predicted fitting trajectory.
[0016] Among them, the decision-making and execution layer includes: an initial decision module, which is used to instruct the robotic arm to perform the fitting task according to the predicted fitting trajectory; a continuous monitoring and evidence collection module, which is used to obtain real-time sensor data from the dynamic environment perception layer in real time during the execution of the fitting task; an inference update module, which is used to make inference decisions when the real-time sensor data is abnormal; and a conclusion abolition and decision correction module, which is used to request the cognition and prediction layer to re-plan a new predicted fitting trajectory when a new decision is made.
[0017] Among them, the decision-making and execution layer also includes: an out-of-distribution detection module, which is used to perform anomaly detection on the four-dimensional Gaussian model during operation and trigger the safety decision-making mechanism when an out-of-distribution object is detected; among them, anomaly detection is used to identify objects that do not belong to any known normal object category.
[0018] The decision-making and execution layer also includes: a human-computer interaction module, which is used to receive the operator's natural language and use the visual language model to perform tasks corresponding to the natural language. When the conclusion abolition and decision correction module makes decision adjustments, the visual language model is used to translate the internal state and decision logic into natural language output that can be understood by humans. Among them, the visual language model adopts model consistency assurance technology to make consistent responses to natural languages with the same semantics but different expressions.
[0019] Among them, the intelligent gear component fitting system selects an action sequence that can maximize the future expected utility at each decision moment.
[0020] The beneficial effects of the present application are: different from the existing technology, the intelligent gear assembly fitting system provided by the present application includes: a dynamic environment perception layer, which is used to use four-dimensional Gaussian sputtering technology to model the elements in the entire working area to obtain a four-dimensional Gaussian model, and in the automated assembly process, receives real-time data streams from various sensors to update the four-dimensional Gaussian model in real time; a cognitive and prediction layer, which is used to input the initial state and target state of the fitting task into the fitting trajectory prediction model when the fitting task needs to be performed to obtain a predicted fitting trajectory; the fitting trajectory prediction model is obtained based on the dynamic linearization training of the observable space operator theory; a decision and execution layer, which is used to instruct the robotic arm to perform the fitting task according to the predicted fitting trajectory, and in the process of performing the fitting task, obtain real-time sensor data from the dynamic environment perception layer in real time, and make inference decisions when the real-time sensor data is abnormal, and when a new decision is generated, request the cognitive and prediction layer to re-plan a new predicted fitting trajectory. It is possible to build a closed-loop intelligent control system that integrates multimodal dynamic perception, high-dimensional process modeling, cross-platform feature compatibility, predictive motion planning and explainable autonomous decision-making, so as to achieve high-robustness, high-efficiency and high-safety full-process automated assembly of high-precision components represented by gear fitting in dynamic, changeable and uncertain industrial environments, and enable it to have active perception of unknown anomalies, explainability of its own decisions and adaptability to new tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. Among them:
[0022] Figure 1 This is a structural diagram of an embodiment of the intelligent gear assembly bonding system provided by the present application;
[0023] Figure 2 This is a structural diagram of an embodiment of a dynamic environment perception layer provided by this application;
[0024] Figure 3 This is a schematic diagram of the structure of an embodiment of the cognition and prediction layer provided by this application;
[0025] Figure 4 It is a structural diagram of an embodiment of the decision-making and execution layer provided by this application. DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It will be understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for ease of description, only some, rather than all, structures related to the present application are shown in the drawings. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0027] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0028] See Figure 1 , Figure 1 This is a schematic diagram of the structure of an embodiment of the intelligent gear assembly bonding system provided in this application. The intelligent gear assembly bonding system comprises a dynamic environmental perception layer 10, a cognition and prediction layer 20, and a decision-making and execution layer 30. This system utilizes a layered, decoupled, systematic architecture, implementing a complete technical chain from environmental perception and process prediction to intelligent decision-making.
[0029] The dynamic environment perception layer 10 is used to model the elements in the entire working area using four-dimensional Gaussian sputtering technology to obtain a four-dimensional Gaussian model, and during the automated assembly process, receives real-time data streams from various sensors to update the four-dimensional Gaussian model in real time.
[0030] In some embodiments, see Figure 2 The dynamic environment perception layer 10 includes: an initialization module 11, an operation module 12, a visual data acquisition module 13, a collaborative perception module 14 for multi-source heterogeneous information fusion, a cross-platform feature compatibility module 15 and a feature quantification module 16.
[0031] The initialization module 11 is used to use one or more cameras to perform multi-perspective scanning of the entire working area when the intelligent gear assembly bonding system is first deployed or the work station is changed, to obtain a multi-perspective image stream, and based on the multi-perspective image stream, use four-dimensional Gaussian sputtering technology to model the elements in the entire working area to obtain a four-dimensional Gaussian model.
[0032] The operation module 12 is used to receive real-time data streams from various sensors to update the four-dimensional Gaussian model during the automated assembly process; the camera image stream is used to correct the color and position of the Gaussian primitives in the four-dimensional Gaussian model, and the motion data of the robotic arm is used to drive the Gaussian primitives representing the robotic arm part in the four-dimensional Gaussian model to move.
[0033] The visual data acquisition module 13 is used to collect original images using single-frame high dynamic range imaging technology, and to enhance the original images in real time using a neural network model. The enhanced original images are used to update the four-dimensional Gaussian model in real time, and the four-dimensional Gaussian model is updated in real time using the data stream collected by the event camera.
[0034] The collaborative perception module 14 of multi-source heterogeneous information fusion is used to unify the visual flow, ontological flow, and prior flow into a collaborative representation space; wherein, the visual flow is provided by the data collected by the HDR camera and the event camera, the ontological flow comes from the joint angle and speed provided by the robot arm encoder, and the real-time feedback of the force / torque sensor on the end gripper, and the prior flow comes from the original three-dimensional computer-aided design model of the component to be assembled.
[0035] The cross-platform feature compatibility module 15 is used to train a feature encoder for each sensor in the intelligent gear assembly fitting system, mapping the original features output by each into a unified, semantically aligned shared latent space; in the shared latent space, information from different devices but pointing to the same physical feature can be directly compared and fused.
[0036] The feature quantization module 16 is used to convert the key geometric features extracted from the four-dimensional Gaussian model into a standardized, fixed-length feature vector; the feature vector has error correction capabilities and can tolerate geometric deviations caused by manufacturing tolerances, slight wear or sensor noise.
[0037] In one application scenario, the dynamic environment perception layer 10 can construct a high-dimensional, robust world model.
[0038] The core goals and concepts are as follows:
[0039] The core goal of the Dynamic Environment Perception Layer 10 is to completely abandon the outdated concept of simplifying the physical world into a two-dimensional, static image, as is common in traditional approaches. Its mission is to build an internal, digital, and fully synchronized world model for the entire intelligent system. This model must possess the following four key characteristics:
[0040] 1. High dimensionality: It is not just a two-dimensional plane, but includes complete three-dimensional geometric structure, materials, lighting, and dynamic changes in the time dimension.
[0041] 2. Robustness: Ability to work stably under common interferences in real industrial environments (such as sudden changes in illumination, noise, vibration, and occlusion).
[0042] 3. Multimodal: The information source is not limited to vision, but integrates multiple sensor data.
[0043] 4. Real-time: Ability to capture and update environmental changes at a high enough frequency to meet control requirements.
[0044] The technical composition and detailed description are as follows:
[0045] 1. Core technology: Upgrading the representation from two-dimensional images to four-dimensional dynamic scenes.
[0046] This is the cornerstone of building a "world model", and its ideas are derived from the latest advances in computer graphics and 3D vision, especially dynamic scene reconstruction technology.
[0047] Technology origin and introduction:
[0048] A technique called 4D Gaussian Splatting (4D-GS) is used to reconstruct dynamic 3D scenes that change over time. Traditional 3D models, such as point clouds or meshes, are inefficient or ineffective in representing dynamic objects. 4D-GS, on the other hand, represents the entire scene as a collection of millions of tiny 3D Gaussian ellipsoids with color and opacity. Each Gaussian ellipsoid has not only its 3D spatial position (x, y, z) and shape (scale, rotation), but also its trajectory over time (t).
[0049] Extended applications in this solution:
[0050] This application draws on this idea and no longer regards the gear or robotic arm to be bonded as a pixel area on a two-dimensional image, but instead models the entire working unit (including robotic arm, fixture, gear, workbench, conveyor belt, etc.) as a four-dimensional Gaussian scene.
[0051] Initialization phase (corresponding to initialization module 11): When the system is first deployed or when a workstation is changed, a rapid multi-view scan of the work area is performed using one or more cameras (which can be fixed or handheld). Through this process, the system constructs an initial, high-fidelity 4D Gaussian model offline, including all static and dynamic elements. This model forms the skeleton of the "world model" of this application.
[0052] During the automated assembly process, real-time data streams from various sensors continuously and efficiently update the model. For example, the camera image stream corrects the color and position of the Gaussian primitives, while the robot arm's motion data directly drives the movement of the Gaussian primitives representing the robot arm in the model.
[0053] Result: This approach creates a continuous, dynamic, and differentiable scene representation within the system. It's no longer a discrete image frame, but a complete digital representation of the entire process of a gear moving from point A to point B. This provides unprecedented information density and spatiotemporal continuity for subsequent prediction and planning.
[0054] 2. Technical support 1: High-fidelity and all-weather visual data acquisition (corresponding to visual data acquisition module 13).
[0055] In order to provide high-quality "nutrients" for the four-dimensional Gaussian model, the collection of raw visual data must overcome the challenges of the industrial environment.
[0056] Dealing with complex lighting:
[0057] The reflections on metal surfaces and shadows from equipment at industrial sites can cause partial overexposure or darkening of images, resulting in a loss of a lot of details. Traditional multi-frame exposure synthesis HDR technology can produce artifacts due to object movement. Therefore, this application introduces single-frame high dynamic range (HDR) imaging technology. This type of technology uses special sensor design or optical encoding to capture a wide brightness range of the scene in a single shutter exposure. At the same time, drawing on the idea of "image dehazing / deblurring", this application uses a lightweight neural network model to perform real-time enhancement on the collected original images, further improving the image clarity and signal-to-noise ratio, ensuring that the four-dimensional model can obtain accurate texture and color information even in harsh lighting conditions.
[0058] Capturing high-speed dynamics:
[0059] When the robotic arm moves at high speed or the gears rotate quickly, traditional cameras will produce motion blur due to exposure time limitations. For this reason, this application introduces an event camera as a supplement. The event camera simulates the biological retina. It does not capture complete image frames, but asynchronously records the brightness change events of each pixel. Its advantages lie in extremely high temporal resolution (microsecond level) and extremely low data redundancy. In this solution, the data stream of the event camera is used to provide an ultra-high frequency motion update signal for the four-dimensional Gaussian model, especially for accurately capturing the instantaneous posture of those high-speed moving parts, thereby greatly improving the dynamic fidelity of the world model.
[0060] 3. Technical support 2: Collaborative perception of multi-source heterogeneous information fusion (corresponding to collaborative perception module 14 of multi-source heterogeneous information fusion).
[0061] After all, single visual information is one-sided, and a robust "world model" must be the result of joint verification of multiple information sources.
[0062] Fusion framework: Information from different sources can verify and complement each other, thus obtaining more reliable conclusions than a single source.
[0063] In this proposal, a multimodal fusion framework is constructed to unify the following information flows into a collaborative representation space:
[0064] 1. Visual stream: data provided by the HDR camera and event camera for updating the 4D Gaussian model.
[0065] 2. Proprioceptive flow: The robot arm encoders provide real-time feedback on joint angles and speeds, as well as the force / torque sensors on the end grippers. This information directly reflects the physical state of the system.
[0066] 3. Prior flow: The original 3D computer-aided design (CAD) model of the component to be assembled. This model provides the most accurate and noise-free geometry, size, and theoretical fit relationship of the object.
[0067] Implementation effect: For example, when the vision system observes that two gears are about to contact, the CAD model can provide an accurate tooth profile for comparison, while the force sensor can monitor in real time whether the contact force is within a safe range. When the three pieces of information are consistent, the system has a very high "confidence" in the current state. If a contradiction occurs (for example, there is no visual contact yet, but the force sensor has a reading), the system can immediately determine that there may be model errors or unforeseen physical interference. This cross-validation mechanism of multi-source information is a key guarantee for the robustness of the "world model."
[0068] 4. Technical support three: cross-platform feature compatibility and real-world adaptability.
[0069] The "world model" must not only respond to environmental changes, but also adapt to the diversity of the hardware itself and the imperfections of the physical world.
[0070] Feature quantization and fault tolerance (corresponding to feature quantization module 16):
[0071] One of the core challenges in the field of biometrics is to deal with "fuzzy" inputs (such as fingerprints where the position and strength of each press are different). They improve the fault tolerance of the system through feature quantization and error correction code technology. This application draws on this idea and does not require the system to perform pixel-level precise matching of the features of the gears. Instead, this application converts the key geometric features extracted from the four-dimensional model (such as the tooth profile curve, root circle diameter, etc.) into a standardized, fixed-length feature vector through a learned feature quantization network. This vector has a certain error correction capability and can tolerate small geometric deviations caused by manufacturing tolerances, slight wear or sensor noise.
[0072] Heterogeneous sensor compatibility (corresponding to cross-platform feature compatibility module 15):
[0073] In large or scalable automation systems, it is common to use cameras or sensors from different vendors. The image processing and feature extraction algorithms inside these devices are different. By matching the feature extraction methods of different algorithms through a shared "embedding space", a feature encoder is trained for each sensor in the system (such as camera of brand A, depth sensor of brand B), and the raw features output by each of them are mapped to a unified, semantically aligned shared latent space. In this space, information from different devices but pointing to the same physical feature can be directly compared and fused. This makes the perception layer of this application have "plug and play" hardware compatibility, greatly improving the flexibility and deployment convenience of the system.
[0074] The summary is as follows:
[0075] In summary, the dynamic environmental perception layer 10 successfully constructs a high-dimensional, robust "world model" through a series of interconnected technical approaches. Core to this approach is four-dimensional Gaussian sputtering, which digitizes the physical world into a dynamic, continuous internal representation. High-fidelity imaging and event vision ensure the quality and real-time nature of input data. Multimodal fusion enables cross-validation and robustness enhancement of information. Finally, feature quantization and cross-platform compatibility address the imperfections of the real world and hardware heterogeneity. This layer provides an unprecedented, high-quality, and high-dimensional information foundation for higher-level cognition, prediction, and decision-making, and is the fundamental prerequisite for the entire intelligent system to realize its advanced capabilities.
[0076] In some embodiments, the cognition and prediction layer 20 is used to input the initial state and target state of the fitting task into the fitting trajectory prediction model when a fitting task needs to be performed to obtain a predicted fitting trajectory; the fitting trajectory prediction model is obtained based on dynamic linearization training of observable space operator theory.
[0077] In some embodiments, see Figure 3The cognition and prediction layer 20 includes: a learning module 21 and a prediction module 22.
[0078] The learning module 21 is used to train the fitting trajectory prediction model using the real or simulated assembly process data provided by the dynamic environment perception layer 10, so that the fitting trajectory prediction model can map the state of the physical world to a point in a high-dimensional virtual observable space, and learn a finite-dimensional approximation matrix of a Koopman operator, which describes the linear evolution law of the points in the virtual observable space over time in the virtual observable space.
[0079] The prediction module 22 is used to input the initial state and target state of the fitting task into the fitting trajectory prediction model when a fitting task needs to be performed to obtain a predicted fitting trajectory; the fitting trajectory prediction model is obtained based on dynamic linearization training of observable space operator theory.
[0080] The fitting trajectory prediction model includes an encoder, an evolution unit, and a decoder. When a fitting task is required, the prediction module 22 inputs the initial and target states of the fitting task into the encoder, maps the initial state into a virtual observable space, and uses the evolution unit to parse a complete, continuous, and smooth optimal trajectory from the initial state to the target state in the virtual observable space. The decoder then decodes the optimal trajectory to obtain a predicted fitting trajectory.
[0081] In one application scenario, the cognitive and prediction layer 20 can achieve a leap from "reaction" to "prediction".
[0082] The core goals and concepts are as follows:
[0083] The control logic of traditional automation systems follows an "if-then" reactive model: a sensor detects a state (If), and the controller executes an action (Then) based on preset rules. The fundamental flaw of this model lies in its short-sightedness and suboptimality. It can only process the current situation and cannot foresee the state several steps ahead. As a result, its decisions are often partial and suboptimal, and they become clumsy and inefficient when faced with complex dynamic processes.
[0084] The core goal of the Cognition and Prediction Layer 20 is to radically overturn this reactive model. It aims to empower the system with the ability to "foresee the future." This means using a deep understanding of the current world (provided by the Dynamic Environment Perception Layer 10) to calculate and deduce the optimal evolutionary path for the system over the next period of time. The core concept of this layer is to shift from "determining the next move" to "planning the optimal path for the entire process all at once," achieving a fundamental shift from "reaction" to "foreknowledge."
[0085] The technical composition and detailed description are as follows:
[0086] To achieve "precognition" of the future, the key is to solve a difficult mathematical problem: how to efficiently and accurately solve the future evolution of a complex nonlinear dynamic system.
[0087] Challenges faced: Complexity of nonlinear dynamics.
[0088] Take gear fitting as an example. A gear held at the end of a robotic arm must be precisely inserted into the tooth groove of another gear along a complex spatial curve, at a specific posture and speed. This process is influenced by numerous factors, including the robotic arm's multi-joint coupling, motor characteristics, gravity, friction, and possible micro-collision forces between components. The resulting dynamics form a highly nonlinear system of ordinary differential equations.
[0089] The traditional solution: numerical solution. This involves dividing time into extremely small steps and iteratively calculating the system's state at the next instant. This method is computationally intensive and time-consuming, making it difficult to meet the requirements of real-time control. Furthermore, it only yields a specific solution and cannot guarantee that this path is the global optimal one.
[0090] To overcome this challenge, the Cognition and Prediction Layer 20 introduces revolutionary tools derived from modern dynamical systems theory and machine learning.
[0091] 1. Core technology: Dynamic linearization prediction based on observable space operator theory.
[0092] This technology is the key to achieving "precognition" capabilities at this level, and its ideas have profoundly changed the way this application handles nonlinear problems.
[0093] Technology origin and introduction:
[0094] The core idea is to use the Koopman operator theory to "unfold" and "straighten out" complex nonlinear flows. The Koopman operator theory states that any nonlinear dynamical system can be represented as a completely linear dynamical system in a higher-dimensional, infinite-dimensional space composed of "observable functions."
[0095] Popular understanding: Imagine a river winding in three-dimensional space (nonlinear dynamics). It is very complicated to directly describe the direction of this river. However, if this application can find a kind of "magic glasses" (a nonlinear mapping), after wearing it, this winding river will look like a straight highway (linear dynamics). On this highway, predicting the future position of any point becomes extremely simple, just use "speed × time". The Koopman operator theory is to help
[0096] This application systematically searches for and constructs the mathematical tools for this pair of “magic glasses”.
[0097] Extended applications in this solution:
[0098] This application uses deep learning to learn this mapping (i.e., encoder) that "straightens" the nonlinear dynamics of the physical world, as well as the evolution law in this linear space (i.e., a fixed linear operator, usually a matrix).
[0099] a. Learning phase (corresponding to learning module 21): During training, the system observes a large amount of real or simulated assembly process data (including the movement of the robot arm, changes in the position and posture of the components, etc.) provided by the perception layer. Its goal is to learn a neural network encoder g(x) that can map the state x of the physical world (such as a vector consisting of the robot arm joint angles, the six-dimensional position and posture of the gear, etc.) to a point g(x) in a high-dimensional virtual "observable space." At the same time, the system also learns a matrix L (a finite-dimensional approximation of the Koopman operator), which describes the linear evolution of g(x) over time in this virtual space.
[0100] b. Prediction phase (i.e., the implementation of "prediction" corresponds to prediction module 22): When a fitting task needs to be performed, the system no longer needs to perform complex numerical integration. Instead, it performs an extremely efficient "encoding-evolution-decoding" three-step process:
[0101] a. Encoding: The system first obtains the initial state x_initial (the upper gear is at point A) and the target state x_target (the upper gear and the lower gear fit perfectly together). Then, it uses the learned encoder g() to map the initial state x_initial to a linear observable space, obtaining g(x_initial).
[0102] b. Linear Evolution: In observable space, the evolution from g(x_initial) to g(x_target) is linear. The system can directly and analytically compute the complete, continuous, and smooth optimal trajectory g(x(t)) from the initial state to the target state through a single matrix exponential operation (e^(Lt)). This is as simple and quick as calculating a route on a straight highway. This step is the core of "precognition," instantly completing the entire future prediction.
[0103] c. Decoding: The encoder design of this application makes the decoding process very straightforward. The predicted high-dimensional trajectory g(x(t)) contains the original physical coordinates. The system can directly "read" the optimal position x(t) of the robot arm and gear in the physical world at each time t, thereby forming a complete, executable motion instruction sequence.
[0104] Achieved results and major breakthroughs:
[0105] By using the Koopman operator theory, this application transforms the problem of solving complex nonlinear differential equations into a problem of learning nonlinear mappings and performing matrix operations. This has a revolutionary effect:
[0106] A leap in efficiency: From time-consuming iterative numerical solutions to near-instantaneous analytical calculations, this makes online real-time optimal path planning possible.
[0107] Guarantee of optimality: Since the evolution occurs in linear space, the resulting path naturally possesses favorable properties such as smoothness and energy optimization, thus avoiding the error accumulation and local optimality problems that may be caused by numerical integration.
[0108] Enhanced interpretability: Decomposing the Koopman operator matrix L into eigenvalues and eigenvectors reveals the inherent "modalities" of the system's dynamics. For example, the magnitude of the eigenvalues can reflect the speed and stability of certain motion patterns, while the eigenvectors correspond to the specific physical form of these patterns. This provides a powerful mathematical tool for understanding and diagnosing system behavior.
[0109] Summarize
[0110] The cognitive and prediction layer 20 is the key link in the evolution of the entire intelligent system from a "tool" to its "brain." By creatively introducing and engineering the Koopman operator theory, it successfully overcomes the challenge of real-time, optimal prediction of complex nonlinear dynamic systems.
[0111] The core contribution of this layer is that it builds a bridge from a deep understanding of the present (the high-dimensional world model provided by the perception layer) to precise foresight of the future (deducing the global optimal path through linearized dynamics). This enables the system's behavior to move beyond passive reactions to external stimuli and instead enable proactive, planned action based on a foresighted vision of the future. This foresight provides a solid foundation for adaptive, explainable, and highly secure decision-making at the upper decision-making layer, ultimately achieving a qualitative leap for the entire system from traditional automation to cognitive intelligence.
[0112] In some embodiments, the decision and execution layer 30 is used to instruct the robotic arm to perform the fitting task according to the predicted fitting trajectory, and in the process of performing the fitting task, obtain real-time sensor data from the dynamic environment perception layer 10 in real time, and make inference decisions when the real-time sensor data is abnormal, and when a new decision is made, request the cognition and prediction layer 20 to re-plan a new predicted fitting trajectory.
[0113] In some embodiments, see Figure 4 The decision-making and execution layer 30 includes: an initial decision module 31, a continuous monitoring and evidence collection module 32, a reasoning and updating module 33, an out-of-distribution detection module 34 and a human-computer interaction module 35.
[0114] The initial decision module 31 is used to instruct the robot arm to perform the bonding task according to the predicted bonding trajectory.
[0115] The continuous monitoring and evidence collection module 32 is used to obtain real-time sensor data from the dynamic environment perception layer 10 in real time during the execution of the fitting task.
[0116] The reasoning update module 33 is used to make reasoning decisions when the real-time sensor data is abnormal.
[0117] The conclusion abolition and decision correction module is used to request the cognition and prediction layer 20 to re-plan a new prediction fitting trajectory when a new decision is generated.
[0118] The out-of-distribution detection module 34 is used to perform anomaly detection on the four-dimensional Gaussian model during operation and trigger a safety decision mechanism when an out-of-distribution object is detected; wherein, anomaly detection is used to identify objects that do not belong to any known normal object category.
[0119] The human-computer interaction module 35 is used to receive the operator's natural language and use the visual language model to perform tasks corresponding to the natural language. When the conclusion abolition and decision correction module makes decision adjustments, the visual language model is used to translate the internal state and decision logic into natural language output that can be understood by humans. Among them, the visual language model adopts model consistency assurance technology to make consistent responses to natural languages with the same semantics but different expressions.
[0120] In one application scenario, the decision-making and execution layer 30 can give the system the ability to "think" and "adapt".
[0121] The core goals and concepts are as follows:
[0122] If the dynamic environmental perception layer 10 is the system's "five senses," and the cognitive and predictive layer 20 is its "left brain" for logical reasoning, then the decision-making and execution layer 30 is the "right brain" and "central nervous system" responsible for coordinating the overall situation, handling unexpected situations, and communicating with the outside world. Traditional automation systems have rigid execution logic. Once activated, they follow a single path and are unable to adapt to any unplanned changes.
[0123] The core goal of the Decision and Execution Layer 30 is to break this rigidity and endow the system with cognitive flexibility. It aims to enable the system, after possessing the ability to "see the world clearly" (the perception layer) and "foresee the future" (the prediction layer), to make intelligent, reasonable, safe, and explainable decisions based on ever-changing situations and real-time feedback, and to translate these decisions into precise physical execution. The core concept of this layer is to enable machines to evolve from tools that only "execute" to intelligent entities that can "think," "reflect," and "adapt."
[0124] The technical composition and detailed description are as follows:
[0125] To achieve this advanced cognitive capability, the decision-making and execution layer 30 integrates a series of technologies from the cutting-edge fields of artificial intelligence to build a complete closed loop that includes reasoning, diagnosis, interaction and execution.
[0126] 1. Core technology: dynamic decision engine based on defeasible reasoning.
[0127] This is the fundamental mechanism by which the system achieves its "thinking" and "adaptation" capabilities. It draws on a very important ability in human cognition: revising existing conclusions based on new information.
[0128] Technology origin and introduction:
[0129] The core concept is defeasible reasoning, which challenges the model to think like a skeptic. Traditional logical reasoning is monotonic: once a conclusion is established, no amount of new information will change it. However, defeasible reasoning is non-monotonic: new evidence can overturn (or invalidate) the previous conclusion at any time. For example, "birds can fly" is a universal conclusion, but when new evidence ("this is a penguin") is introduced, the conclusion "it can fly" is invalidated.
[0130] Extended applications in this solution:
[0131] This application builds this reasoning model into the system's decision-making engine. Each decision of the system is no longer final, but a temporary and challengeable "hypothesis."
[0132] Example workflow:
[0133] i. Initial Decision (corresponding to Initial Decision Module 31): The Cognition and Prediction Layer 20 predicts an optimal alignment trajectory based on the current state and goal. The Decision Layer adopts this trajectory and forms an initial decision: "Hypothesis 1: Executing alignment along Path A is safe and optimal."
[0134] ii. Continuous monitoring and evidence collection (corresponding to the continuous monitoring and evidence collection module 32): When the robot arm starts to execute path A, all sensors in the perception layer (force, vision, encoders, etc.) are working continuously to collect new "evidence" in real time.
[0135] iii. Reasoning Update (corresponding to Reasoning Update Module 33): Suppose that halfway through the execution, the end force sensor detects a small, unexpected force feedback that exceeds the normal friction range. This new evidence is fed into the decision engine.
[0136] iv. Conclusion Revocation and Decision Correction (corresponding to the Conclusion Revocation and Decision Correction module): The engine infers: "New evidence: An abnormal force of 0.5N appears in the Z-axis direction. This evidence contradicts the assumption that 'Path A is safe.'" Therefore, "Hypothesis 1" is revoked. The system immediately generates a new decision: "Revised Decision: Immediately pause the current motion, slightly back off 0.1mm to relieve stress, and request the prediction layer to replan Path B, avoiding potential interference points."
[0137] Effect: Through defeasible reasoning, the system acquires the core capability to cope with dynamic uncertainty. Instead of blindly executing predetermined plans, it constantly compares them with reality, and possesses closed-loop control capabilities to self-correct when the two diverge. This "think twice before acting, and think twice before acting" approach is a key manifestation of advanced intelligence.
[0138] 2. Technical support 1: Active security defense against unknown risks.
[0139] In order for a defeasible inference engine to work effectively, the system must be able to perceive those “unplanned” anomalies.
[0140] Technology origin and introduction:
[0141] Through out-of-distribution (OOD) detection, which identifies novel samples that the model has never seen during training and do not belong to any known categories, it leverages the powerful prior knowledge of the Visual Language Model (VLM) and a carefully designed graph propagation algorithm to effectively distinguish between "known" and "unknown".
[0142] Extended applications in this solution:
[0143] This application applies OOD detection technology to real-time monitoring of a three-dimensional workspace. During training, the system learns all "normal" object categories (e.g., type A gear, type B gripper, standard workbench, etc.). During operation, the four-dimensional world model constructed by the perception layer is continuously fed into the OOD detection module (corresponding to the out-of-distribution detection module 34).
[0144] Anomaly detection: When an incorrect part is placed on a workstation, a tool is forgotten in the work area, or even a broken drill bit falls on a gear, the OOD module will immediately recognize that these objects do not belong to any known "normal" category.
[0145] Triggering safety decisions: Once an OOD object is detected, the system will immediately mark it as "high-priority abnormal evidence" and trigger the highest level of safety decisions, such as immediately stopping all mechanical movements, sounding an audible and visual alarm, and highlighting the location and shape of the abnormal object on the human-computer interaction interface, requesting manual intervention.
[0146] Result: OOD detection empowers the system with the ability to proactively discover unknown risks. It transcends traditional passive security mechanisms that rely on predefined fault lists and effectively protects against the infinite number of possible anomalies in the open world.
[0147] 3. Technical support 2: Human-computer collaboration and model consistency assurance based on natural language.
[0148] An advanced intelligent system ultimately needs to work efficiently and reliably with humans.
[0149] Technology origin and introduction:
[0150] Improves the reasoning capabilities and output consistency of the Large Language Model (LLM) and Visual Language Model (VLM). Test-time consistency technology ensures that the model responds consistently to commands with the same semantics but different expressions. Multi-image contrast-enhanced visual reasoning enhances the model's ability to discern subtle visual differences through contrastive learning.
[0151] Extended applications in this solution (corresponding to human-computer interaction module 35):
[0152] a. Natural Language Interaction: The system integrates a virtual machine language interface (VLM) as the core of human-machine interaction. Operators can use natural language to issue complex queries and commands, such as, "Compare the current positions of gears A and B and calculate their axial and radial deviations," or "Reduce the insertion speed at the end of the fitting process by 15% and then continue."
[0153] b. Explainable Output: More importantly, when the system makes decisions based on its internal defeasible reasoning engine, it can leverage the VLM to translate complex internal states and decision logic into natural language understandable to humans. For example, in the aforementioned abnormal force feedback event, the system could display the following on the console: "Warning: An unexpected collision force of 0.5N was detected in the Z-axis direction, halting operation. Preliminary assessment suggests possible interference with the positioning pin. It is recommended to check the workpiece positioning or manually fine-tune the robot arm's posture."
[0154] c. Consistency and Reliability: By introducing model consistency assurance technology, this application ensures the stability and reliability of the VLM when performing these interactive tasks. Whether the operator asks "What is the deviation?" or "Calculate the error," the system understands the core intent and gives the same accurate answer.
[0155] Result: This technology transforms the system from a cold machine into an "intelligent colleague" that can communicate and explain its behavior. This transparency and explainability are crucial for building trust in machines and enabling efficient troubleshooting and process optimization in complex industrial scenarios.
[0156] In some embodiments, the intelligent gear assembly fitting system selects an action sequence that maximizes future expected utility at each decision moment.
[0157] In some embodiments, the intelligent gear assembly bonding system has a unified decision optimization formula for a systematic intelligent bonding system. To describe how the entire system makes optimal decisions in a dynamic, uncertain environment, we construct an expected utility maximization framework based on risk perception and time-varying optimal control. The ultimate goal of the system is to select an action sequence (t) that maximizes future expected utility at each decision moment τ. This optimization problem can be expressed as:
[0158]
[0159] This formula is the top-level mathematical abstraction of the entire system. Below, we will break down each of its core components in detail and explain how they integrate the technical essence of each level.
[0160] Detailed explanation of the core components of the formula:
[0161] 1. World model (τ): mathematical expression of the dynamic environment perception layer 10.
[0162] The world model is the system's internal cognitive representation of the entire operating environment at the decision time τ. It is not a static parameter, but a set of multimodal probability distributions that include uncertainty. Its mathematical form is:
[0163] World model (τ) = {P(4D scene (τ) | observation (≤τ)), P(entity state (τ) | entity feedback (≤τ)), P(component prior)}.
[0164] P(4D scene (τ) | observation (≤τ)): This is a probabilistic estimate of the current state of the physical world.
[0165] The four-dimensional scene (τ) is a collection of massive three-dimensional Gaussian primitives, each of which contains position, shape, color, and time evolution information, fully describing the dynamic form of the robotic arm, components, etc. at time τ.
[0166] Observations (≤τ) represent all visual data streams collected from system startup to the current time τ, including high dynamic range images and event camera data streams.
[0167] This conditional probability distribution P(...) reflects the uncertainty of perception. For example, in areas with occlusion or poor lighting, the variance of the position distribution of the four-dimensional scene will be larger.
[0168] P(proprioceptive state (τ) | proprioceptive feedback (≤τ)): This is a probabilistic estimate of the robot's own state, integrating data from joint encoders, force / torque sensors, and other data.
[0169] P (component prior): This is the prior probability distribution about the component's exact geometry and ideal fit extracted from its 3D CAD model.
[0170] Creative fusion point: The construction of this world model deeply integrates the ideas of four-dimensional Gaussian sputtering, HDR imaging, event vision and multimodal information fusion, forming a world representation that is far richer, more dynamic and more probabilistic than traditional two-dimensional images. It shows that the system's decision is the average optimal choice after considering all these perceptual uncertainties.
[0171] 2. Evolution of state (t): mathematical expression of cognitive and prediction layer 20.
[0172] State(t) is the predicted state of the system at the future time t (t>τ). Its evolution is no longer predicted through complex numerical integration, but rather through the theory of Koopman operators, which is the core of the "precognition" of this solution.
[0173] d / dt g(state(t))=L g(state(t)).
[0174] in:
[0175] g(·) is a learned nonlinear encoder that maps the state of the physical world (t) to a high-dimensional “observable space” where the dynamics are linear.
[0176] L is the learned Koopman generator (a constant matrix) that describes the evolution of the state in the linear observable space.
[0177] The solution to this linear ordinary differential equation is: g(state(t))=e L(t-τ) g(state(τ)).
[0178] Creative integration point: Through this formula, the problem of solving complex nonlinear dynamics is transformed into an elegant matrix exponential operation. This makes long-term and accurate prediction of future states efficient and feasible, and is the key to achieving The basis for this future integration.
[0179] 3. Utility function (...): mathematical expression of the decision-making and execution layer 30.
[0180] The utility function is the core of decision-making. It quantifies the "benefits" that can be obtained by the system being in state (t) and performing action (t) at a certain future time t. It is a carefully designed composite function that includes multiple objectives:
[0181] Utility function = task benefit - execution cost - risk penalty.
[0182] Task benefit: describes the benefits of completing the final task.
[0183] Task benefit = weight_fit * fit(state(t))
[0184] Fit is a function that measures how close two gears are to perfect fit, reaching its maximum value when the fit is perfect. It is calculated by comparing the current state with the ideal state of the components known a priori. The calculation process incorporates robust feature quantization and cross-platform feature matching techniques to ensure accuracy in the presence of noise and hardware variations.
[0185] Execution cost: describes the resources consumed to execute an action.
[0186] Execution cost = weight_energy||action(t)|| 2 +weight_time.
[0187] It penalizes excessive arm movements (energy consumption) and excessive time consumption.
[0188] Risk penalty: This is the key to embodying the system's advanced intelligence, integrating the ideas of defeasible reasoning and out-of-distribution anomaly detection.
[0189] Risk penalty = weight_collision * collision probability (state(t)) + weight_anomaly * anomaly score (state(t)).
[0190] The collision probability comes from Monte Carlo sampling of the 4D scene representation of the robot arms and components in the world model, calculating the probability of their interference in the future.
[0191] The anomaly score is the output of the out-of-distribution detection module 34. When any unknown object that does not belong to the normal working range appears in the world model, this score will increase sharply, resulting in a large penalty term.
[0192] Creative Integration: The utility function is no longer simply "the closer to the target, the better." It is a complex function that comprehensively considers task progress, resource consumption, and future risks. In particular, the risk penalty term quantifies the perception of "unknown risks" as part of decision-making, a rare feature in traditional optimal control theory and reflects the system's proactive risk avoidance intelligence.
[0193] 4.γ (discount factor) and
[0194] γ is a future benefit discount factor between 0 and 1. The smaller γ is, the more the system focuses on immediate benefits; the larger γ is, the more "foresighted" the system is, and the more it values the long-term overall optimality.
[0195] The operator represents the complex optimization problem the system must solve to find the sequence of actions that maximizes total future expected utility. Thanks to the introduction of the Koopman operator, this seemingly complex solution becomes computationally feasible due to the linearization of state evolution.
[0196] Summary: A creative unified mathematical framework.
[0197] This formula creatively links three core aspects together through a unified decision-making optimization goal:
[0198] It uses the probabilistic world model constructed by the perception layer as the starting point for decision-making and the source of uncertainty.
[0199] It uses the Koopman linearized dynamics of the prediction layer as the core engine (e L(t-τ) ).
[0200] It uses the multi-objective utility function of the decision-making layer as the criterion for judging the future, and incorporates the avoidance of unknown risks (risk penalty) into mathematical considerations.
[0201] Ultimately, by solving this optimization problem, the system outputs no longer a simple "next step" but a comprehensive, optimal trajectory of action that carefully considers efficiency, cost, and all known and unknown future risks. This formula itself is the mathematical embodiment of the intelligent system's complete logical closed loop, from "reaction" to "foresight" to "deliberate action."
[0202] Summarize
[0203] The decision-making and execution layer 30 is the "commander" of the entire intelligent system. Its core decision-making mechanism, revocable reasoning, enables flexible adaptation and self-correction to dynamic environments. It uses object-oriented (OD) anomaly detection as a security sentinel, building a powerful capability for proactively discovering unknown risks. And with interpretable natural language interaction as a communication bridge, it fosters unprecedented trust and collaboration between humans and machines.
[0204] This layer of development ultimately transforms the perception layer's ability to "see the world clearly" and the prediction layer's ability to "foresee the future" into practical action capabilities that enable safe, efficient, and intelligent completion of tasks in the complex and uncertain real world. It provides machines with a logical framework for "thinking" and an action guide for "adapting," and is the final and most critical link in achieving advanced cognitive intelligence.
[0205] The unexpected technical effects of this application are as follows:
[0206] Through the systematic integration of the above three levels and multiple technologies, this solution not only solves the preset complex technical problems, but also brings a series of unexpected and paradigm-changing value-added technical effects, which together constitute the creative value of the system.
[0207] 1. The leap from "automated execution" to "digital twin simulation":
[0208] The initial goal of this solution was to complete assembly tasks in the physical world. However, due to the introduction of four-dimensional dynamic scene modeling and motion prediction technology based on observable space, the system inevitably constructs a high-fidelity digital twin that is highly synchronized with the physical world and can accurately predict its future dynamic evolution. This digital twin model not only serves real-time equipment control but also becomes a valuable byproduct in its own right. It can be used independently for offline process simulation, motion path optimization, equipment fault diagnosis, and even virtual reality operator training, greatly expanding the application boundaries and value dimensions of automation systems.
[0209] 2. The leap from “black box control” to “explainable decision-making”:
[0210] Traditional automated systems are "black boxes" that cannot communicate, and their behavior is difficult to understand. This solution, by integrating a language model and a defeasible reasoning engine, enables the system to "replay" and "explain" its decision-making process. When the system adjusts its strategy due to an anomaly, it can generate human-readable logs or voice prompts: "An unexpected force feedback of 0.5 Newtons was detected in the Z-axis. To avoid potential component damage, the current fitting speed has been reduced by 20%, and the insertion angle has been replanned to avoid potential interference points." This interpretability of decisions has revolutionized human trust in complex automated systems and improved the efficiency of collaborative work.
[0211] 3. From "dedicated equipment" to "general capability platform":
[0212] The initial technical solution was designed for the specific task of gear assembly. However, this solution, through a modular architecture, decouples and abstracts perception, prediction, and decision-making capabilities, ultimately resulting in a set of generalized capabilities for high-precision dynamic process control. In theory, simply by replacing the 3D prior model of the component to be assembled and the definition of the task objective, this system can be rapidly adapted to other precision assembly tasks, such as precise press-fitting of miniature bearings, high-speed placement of semiconductor chips, and flexible docking of complex connectors. This demonstrates strong task generalization capabilities, transforming it from a specialized device into a universal intelligent assembly platform.
[0213] 4. The emerging “active defense” security paradigm:
[0214] Traditional safety mechanisms are passive, relying on pre-set safety fences or emergency stop buttons. This solution achieves proactive awareness and defense against unknown risks through out-of-distribution anomaly detection and multimodal information fusion (particularly force feedback). Rather than relying on a known catalog of hazards, it continuously monitors the environment for anomalies. This shift from "passive safety" to "active defense" elevates industrial production safety to a whole new level.
[0215] In summary, this solution, through the systematic and creative integration of multiple cutting-edge technologies, elevates a specific engineering problem into the construction of a new intelligent system paradigm. The ultimate technical effect achieved far exceeds the simple addition of its individual components, forming a logically rigorous and valuable closed loop of innovation.
[0216] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another system, or ignoring or not implementing certain features.
[0217] If the integrated units in the above other embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0218] The above description is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. An intelligent gear assembly bonding system, characterized in that: The intelligent gear assembly fitting system includes: The dynamic environment perception layer is used to model the elements in the entire working area using four-dimensional Gaussian sputtering technology to obtain a four-dimensional Gaussian model. During the automated assembly process, the four-dimensional Gaussian model is updated in real time by receiving real-time data streams from various sensors; The cognition and prediction layer is used to input the initial state and target state of the fitting task into the fitting trajectory prediction model when the fitting task needs to be performed to obtain a predicted fitting trajectory; the fitting trajectory prediction model is obtained based on dynamic linearization training of observable space operator theory; The decision-making and execution layer is used to instruct the robotic arm to perform the fitting task according to the predicted fitting trajectory, and in the process of executing the fitting task, obtain real-time sensor data from the dynamic environment perception layer in real time, and make inference decisions when the real-time sensor data is abnormal, and when a new decision is generated, request the cognition and prediction layer to re-plan a new predicted fitting trajectory.
2. The intelligent gear assembly bonding system according to claim 1, characterized in that: The dynamic environment perception layer includes: an initialization module, configured to scan the entire working area from multiple perspectives using one or more cameras when the intelligent gear assembly bonding system is first deployed or a workstation is changed, thereby obtaining a multi-perspective image stream, and to model elements within the entire working area using a four-dimensional Gaussian sputtering technique based on the multi-perspective image stream to obtain the four-dimensional Gaussian model; An operating module is used to receive real-time data streams from various sensors to update the four-dimensional Gaussian model during the automated assembly process; wherein the camera image stream is used to correct the color and position of the Gaussian primitives in the four-dimensional Gaussian model, and the motion data of the robotic arm is used to drive the Gaussian primitives representing the robotic arm parts in the four-dimensional Gaussian model to move.
3. The intelligent gear assembly bonding system according to claim 2, characterized in that: The dynamic environment perception layer also includes: a visual data acquisition module, configured to acquire original images using single-frame high dynamic range imaging technology, and enhance the original images in real time using a neural network model, wherein the enhanced original images are used to update the four-dimensional Gaussian model in real time, and the four-dimensional Gaussian model is updated in real time using the data stream acquired by the event camera; A collaborative perception module that integrates multi-source heterogeneous information into a unified visual, ontological, and priori data streams into a collaborative representational space. The visual stream is provided by data collected by an HDR camera and an event camera. The ontological data stream comes from the joint angles and velocities provided by the robot arm encoder, as well as real-time feedback from the force / torque sensor on the end gripper. The priori data stream is derived from the original 3D computer-aided design model of the component to be assembled. A cross-platform feature compatibility module is used to train a feature encoder for each sensor in the intelligent gear assembly fitting system, mapping the raw features output by each into a unified, semantically aligned shared latent space; in the shared latent space, information from different devices but pointing to the same physical feature can be directly compared and fused.
4. The intelligent gear assembly bonding system according to claim 2, characterized in that: The dynamic environment perception layer also includes: A feature quantization module is used to convert key geometric features extracted from the four-dimensional Gaussian model into a standardized, fixed-length feature vector; the feature vector has error correction capability and can tolerate geometric deviations caused by manufacturing tolerances, slight wear or sensor noise.
5. The intelligent gear assembly bonding system according to claim 1, characterized in that: The cognitive and predictive layer includes: a learning module, configured to train the fitting trajectory prediction model using real or simulated assembly process data provided by the dynamic environment perception layer, so that the fitting trajectory prediction model can map the state of the physical world to points in a high-dimensional virtual observable space, and learn a finite-dimensional approximation matrix of a Koopman operator, wherein the finite-dimensional approximation matrix describes the linear evolution law of the points in the virtual observable space over time in the virtual observable space; The prediction module is used to input the initial state and target state of the fitting task into the fitting trajectory prediction model when a fitting task needs to be performed to obtain a predicted fitting trajectory; the fitting trajectory prediction model is obtained based on dynamic linearization training of observable space operator theory.
6. The intelligent gear assembly bonding system according to claim 5, characterized in that: The fitting trajectory prediction model includes: an encoder, an evolution unit, and a decoder. The prediction module is further used to input the initial state and target state of the fitting task into the encoder when a fitting task needs to be performed, map the initial state into the virtual observable space, and use the evolution unit to parse a complete, continuous, and smooth optimal trajectory from the initial state to the target state in the virtual observable space; and use the decoder to decode the optimal trajectory to obtain a predicted fitting trajectory.
7. The intelligent gear assembly bonding system according to claim 1, characterized in that: The decision-making and execution layer includes: an initial decision module, configured to instruct the robotic arm to perform the bonding task according to the predicted bonding trajectory; A continuous monitoring and evidence collection module, configured to obtain real-time sensor data from the dynamic environment perception layer in real time during the execution of the fitting task; An inference update module, configured to make an inference decision when the real-time sensor data is abnormal; The conclusion abolition and decision correction module is used to request the cognitive and prediction layer to re-plan a new prediction fitting trajectory when a new decision is generated.
8. The intelligent gear assembly bonding system according to claim 7, characterized in that: The decision-making and execution layer also includes: The out-of-distribution detection module is used to perform anomaly detection on the four-dimensional Gaussian model during operation and trigger a safety decision mechanism when an out-of-distribution object is detected; wherein the anomaly detection is used to identify objects that do not belong to any known normal object category.
9. The intelligent gear assembly bonding system according to claim 7, characterized in that: The decision-making and execution layer also includes: The human-computer interaction module is used to receive the operator's natural language and use the visual language model to perform the task corresponding to the natural language. When the conclusion abolition and decision correction module makes decision adjustments, the visual language model is used to translate the internal state and decision logic into natural language output that can be understood by humans. The visual language model adopts model consistency assurance technology to make consistent responses to natural languages with the same semantics but different expressions.
10. The intelligent gear assembly bonding system according to any one of claims 1 to 9, characterized in that: The intelligent gear assembly fitting system selects an action sequence that maximizes future expected utility at each decision moment.