Parameter self-learning algorithm based on dynamic Gaussian clustering consistency

Through the parameter self-learning algorithm of dynamic Gaussian cluster consistency, the problem of robotic arms lacking unified expression and adaptation in multimodal information processing is solved, real-time perception and stable control of complex environments is realized by the robot, and the execution ability of industrial assembly tasks is improved.

CN120470341APending Publication Date: 2025-08-12HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510616511.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art lacks online adaptability to dynamic environments in autonomous manipulation and industrial assembly scenarios, multimodal perception results lack uniform semantic expression, control strategies are separated from environmental understanding, and deep learning models are difficult to update internal representations in real time to achieve generalization.

Method used

A parameter self-learning algorithm based on dynamic Gaussian cluster consistency is adopted, and multimodal information is fused through dynamic Gaussian semantic fields, and a control strategy and causal graph are generated using a dual-branch network, combining a closed-loop bias correction mechanism to achieve real-time update and adaptation of the model.

Benefits of technology

It has achieved the real-time perception ability of robots for complex dynamic environments, improved the stability and autonomy of task execution, and enhanced the flexibility and automation level of industrial production lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005401189240000021
    Figure BDA0005401189240000021
  • Figure BDA0005401189240000051
    Figure BDA0005401189240000051
  • Figure BDA0005401189240000059
    Figure BDA0005401189240000059
Patent Text Reader

Abstract

The invention relates to the cross technical field of robot control, computer vision and artificial intelligence, in particular to a parameter self-learning algorithm based on dynamic Gaussian clustering consistency. Aiming at the problems that multi-modal information is difficult to express uniformly, dynamic self-adaptability is lacked and a control strategy and task understanding are disjointed in industrial assembly of an existing mechanical arm, the invention provides a method for constructing a uniform time sequence semantic expression by utilizing a dynamic Gaussian clustering consistency mechanism; a multi-head Kolmogorov-Arnold network and a differential attention network are adopted to realize multi-modal information fusion and stable clustering, a causal module and a strategy module double-branch structure are adopted to simultaneously generate an adaptive control track, action and abstraction, a strategy causal graph and supervise strategy feedback consistency, and parameter self-learning is realized in combination with closed-loop correction. According to the invention, the autonomy and generalization ability of the mechanical arm in an industrial environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the interdisciplinary field of robot control, computer vision, and artificial intelligence, and in particular to a parameter self-learning algorithm based on dynamic Gaussian clustering consistency, which is suitable for robot arm manipulation and industrial assembly tasks. Background Art

[0002] In autonomous robotic arm control and industrial assembly scenarios, robots often need to simultaneously process multimodal information, including visual (RGB-D), point clouds, language commands, and their own state, to accurately understand the operating environment and task requirements. Parameter self-learning algorithms, which adaptively optimize internal control system parameters through real-time or near-real-time feedback, are considered a key means of improving robot flexibility and generalization capabilities.

[0003] However, existing technologies still have the following prominent deficiencies in utilizing multimodal data to drive robot control: (1) lack of online adaptation to dynamic environments, and lack of unified and consistent semantic expression between multimodal perception results; (2) control strategies are separated from environmental understanding, and fail to explicitly model task causal dependencies; (3) deep learning models have difficulty in updating their internal representations in real time according to new scenarios after deployment to achieve generalization.

[0004] In general, current parameter self-learning methods can be divided into two categories: traditional fixed parameter optimization methods and deep learning-based self-learning methods. However, both have limitations in terms of unified representation of multimodal semantics and real-time adaptation. Existing representative research can be roughly summarized into the following categories:

[0005] 1) Method based on traditional fixed parameter control model

[0006] These methods rely on precise dynamic models and offline tuning of control gains (such as PID adaptation, model predictive control (MPC), and gain scheduling). They perform stably in predictable environments or where changes are minimal. However, when faced with the frequently changing material postures, interfering forces, or task switching found in assembly sites, fixed parameters make timely adjustments difficult, resulting in a significant decrease in control accuracy and robustness.

[0007] 2) Multimodal perception and fusion methods

[0008] Existing work typically processes RGB-D images, point clouds, and text instructions separately, lacking cross-modal alignment and constraints, making it difficult to construct a unified 3D semantic field. Some research attempts to align vision and language, but these efforts often focus on 2D images or static 3D spaces, ignoring the scene's evolution over time. This results in a lack of consistent, long-term understanding of the robot's environment.

[0009] 3) Parameter self-learning algorithm based on reinforcement learning

[0010] Reinforcement learning uses a reward function to drive robots to optimize their strategies during interaction, achieving high generalization capabilities in unknown dynamic environments. However, it requires a large number of training samples, converges slowly, and is sensitive to reward design. This is particularly true in multi-step assembly tasks, where a lack of task causal modeling makes it difficult to adjust subsequent strategies online if a step fails to achieve the expected result.

[0011] 4) Parameter self-learning algorithm based on imitation learning

[0012] Imitation learning quickly acquires initial strategies through expert demonstrations and is easy to implement. However, in real-world operations, covariate drift can occur: the robot may enter a state space not covered by the demonstration data, resulting in a sudden drop in performance. Furthermore, this method lacks real-time error correction and adaptive adjustment mechanisms, making it difficult to cope with continuously changing operating conditions.

[0013] 5) Multimodal control network based on attention mechanism and Transformer structure

[0014] Transformers, with their global modeling capabilities, have become a hot topic in multimodal information fusion. However, single backbone networks often only output a single form of result, unable to simultaneously generate a structured semantic field for environmental representation and a policy map for decision-making. Model parameters are largely fixed after deployment, lacking a mechanism for dynamic updates based on feedback, limiting adaptability.

[0015] Furthermore, existing control strategies generally map sensor states directly to actions, lacking explicit representation of the causal dependencies between task steps. When a step deviates, the robot lacks a basis for adjusting subsequent plans, making it difficult to ensure the continuity and success rate of the assembly process.

[0016] In summary, the existing technology has not yet provided a method that can: (1) construct a multimodal scene representation that is adaptively updated over time, (2) explicitly model task causality and deeply couple it with the control strategy, and (3) dynamically adjust the model's internal representation according to changes in environmental instructions after the algorithm model is deployed, thereby achieving rapid generalization to new task scenarios beyond the training data. These shortcomings result in insufficient strategic flexibility and accuracy when the robot arm performs complex and changeable industrial assembly tasks. In order to solve the above problems, the present invention aims to provide a parameter self-learning control algorithm based on dynamic Gaussian clustering consistency. Summary of the Invention

[0017] To overcome the shortcomings of the existing technical problems, the present invention proposes a parameter self-learning algorithm based on dynamic Gaussian clustering consistency, which uses dynamic Gaussian clustering and multiple consistency constraints to provide an accurate and unified multimodal scene representation that evolves over time, eliminating interference and conflict between modalities; through a dual-branch network structure, low-latency action execution and long-context-based policy causal graph updates are simultaneously achieved; when new changes or consistency deviations occur in the environment, by triggering incremental updates or global re-clustering of the model, the internal model is always matched with the current environment and task requirements, thereby significantly improving the stability and autonomy of long-term operation.

[0018] The dynamic Gaussian of the present invention is represented by the following 13 parameters: The position μ i , direction R i , scale s i , existence probability α i , spherical harmonic color Material parameter m i , semantic vector First-order kinetics (v i ,ω i ), second-order dynamics (a i ,β i ) and 4D time window Together they characterize the geometry, appearance, semantics, and dynamic properties of objects or regions in the environment.

[0019] The present invention achieves a unified representation of cross-modal information by constructing a dynamic Gaussian semantic field (DGSF) that is updated over time, and uses the "Gaussian cluster consistency" constraint to ensure that the same real object is continuously represented by the same cluster of Gaussians in time series, thereby avoiding the problems of frequent addition and deletion of components and model jitter that occur in traditional GMM.

[0020] To achieve the above objectives, the present invention proposes the following technical solution, the main steps of which include multimodal data acquisition, Gaussian semantic field generation, dynamic Gaussian clustering, control strategy and abstract causal graph generation, robot real-time decision-making and execution, and closed-loop correction mechanism, as described below.

[0021] (1) Multimodal data acquisition and synchronization: Utilizing RGB-D cameras, structured light / laser point cloud sensors, industrial DSL language instructions, and robotic arm joint encoders, five types of input data are acquired: color images, depth images, dense point clouds, task text, and self-status. All modalities are synchronized via the IEEE 1588PTP global clock to ensure timestamp accuracy within 200ns.

[0022] (2) Multimodal projection to dynamic Gaussian semantic field: Multimodal feature extraction is performed using a pre-trained model, and then a multi-head Kolmogorov-Arnold network (KAN) projection layer is used to obtain a dynamic Gaussian semantic field:

[0023] Among them, the visual encoder uses ResNet to extract multi-scale appearance features and combines the depth map back projection to obtain three-dimensional coordinates;

[0024] Among them, the point cloud encoder uses PointNet++ to extract local geometric features;

[0025] Among them, the language encoder is based on lightweight BERT, which parses industrial instructions into semantic vectors;

[0026] Among them, the robot state encoder normalizes the joint angle and end load;

[0027] Among them, the multi-head KAN projection layer obtains a unified dynamic Gaussian semantic field by performing spatial-semantic-dynamic feature projection on four types of multimodal features.

[0028] (3) Dynamic Gaussian clustering: Dynamic Gaussian clustering is performed using a cross-modal differential attention layer to calculate the composite distance between Gaussian functions. Dynamic clustering is then performed based on the calculation results to form dynamic Gaussian clusters with consistent spatiotemporal semantics.

[0029] (4) Control strategy and abstract causal graph generation: A dual-branch deep neural network is used to simultaneously generate the trajectory action of the current control strategy and the abstract strategy causal graph that takes into account the context of scene changes based on dynamic Gaussian clusters. This ensures that the decision is based on complete multimodal scene information and that the causal relationship is abstractly summarized.

[0030] The backbone network of the action module is a Mamba structure, which leverages the long-term temporal modeling capabilities of the Mamba sequence model and ultimately outputs the correct strategy with the policy feedback consistency obtained through collaboration with the causal module.

[0031] The backbone network of the causal module is a graph neural network (GNN), which uses the dynamic Gaussian clusters that complete clustering to generate an abstract causal graph, and uses the trajectory + action output by the action module to generate a strategy causal graph.

[0032] (5) Real-time decision-making and execution of robots: The current dynamic Gaussian clustering results are input into the action module to generate the trajectory and action of the current control strategy. The current dynamic Gaussian clustering results are input into the semantic + edge projection layer to generate an abstract causal graph. The strategy trajectory and action input trajectory + action encoding layer generate the corresponding strategy causal graph. The strategy trajectory and action are executed and feedback is obtained. The algorithm updates the dynamic Gaussian clustering, abstract causal graph and strategy causal graph based on the feedback.

[0033] (6) Closed-loop correction: The algorithm executes the strategy trajectory and actions, and calculates the strategy feedback consistency based on the feedback update. When the result does not meet the expectation, the algorithm inputs the strategy feedback consistency loss back to the strategy module for strategy fine-tuning. If the strategy feedback still does not meet the expectation after fine-tuning, correction is completed through dynamic Gaussian reclustering.

[0034] The present invention adopts a multi-threaded parallel architecture design: the perception thread is 30Hz; the decision thread is 10Hz; and the learning thread is 1Hz. The three threads exchange dynamic Gaussian clustering scene representation and abstract causal graph through zero-copy shared memory. The mutex lock waiting time is less than 0.2ms, ensuring industrial real-time performance.

[0035] The present invention implements a method for parameter self-learning: using geometric consistency, semantic consistency and temporal consistency to constrain the dynamic Gaussian clustering process, retaining the dynamic Gaussian scene representation with spatiotemporal continuous semantics; using an abstract causal graph and a strategy causal graph with global causal relationships combined with the real-time strategy feedback consistency constraints during strategy execution to achieve online strategy self-learning and correction without manual debugging and recalibration.

[0036] Among them, based on the abstract causal graph and policy causal graph generated above, the algorithm plans and adjusts the task strategy in real time: the robot arm's current execution action interacts with the policy causal graph to determine the key states and executable actions required for the next action decision. At the same time, the abstract causal graph records the different usage scenarios of each task strategy and guides the robot arm to make strategy selections.

[0037] The algorithm framework described in the present invention is designed according to a modular concept. The six units of multimodal data acquisition and synchronization, Gaussian semantic field generation, dynamic Gaussian clustering, control strategy and abstract causal graph generation, and robot real-time decision-making and execution and closed-loop correction have clear logic and standardized interfaces, which facilitates engineering deployment and subsequent functional expansion.

[0038] Compared with the prior art, the present invention has the following advantages or beneficial effects:

[0039] Multimodal input fuses vision, language, point cloud and robot state information through dynamic Gaussian semantic fields, enabling the robot to achieve deep integrated understanding at the geometric, appearance and semantic levels, avoiding errors caused by conflicts and interference between modalities.

[0040] The dynamic Gaussian clustering consistency mechanism realizes the spatiotemporal and temporal continuous multimodal environment semantic expression, significantly improving the robot's real-time perception capability of complex dynamic environments.

[0041] The causal module explicitly models task causal constraints, providing an explainable planning basis for the robotic arm and improving the stability and recovery capability of task execution.

[0042] The policy module integrates abstract causal information of the environment with current observations into a latent state sequence, and then generates future trajectories and action plans through sequential reasoning capabilities. Finally, it co-evolves with the causal module under the constraint of policy feedback consistency to improve decision-making accuracy.

[0043] Through the self-supervised loss design composed of multiple consistency constraints, the algorithm has the ability of online parameter self-learning, can automatically adjust the internal model without stopping the line, and significantly enhance the flexibility of the industrial production line.

[0044] Closed-loop error correction and adaptive decision-making systems reduce manual intervention and improve the automation level and economic benefits of scenarios such as industrial assembly and warehousing sorting. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0046] Figure 1 This is a flow chart of a parameter self-learning algorithm based on dynamic Gaussian clustering consistency according to an embodiment of the present invention;

[0047] Figure 2 Schematic diagram of a Gaussian semantic field generation module according to an embodiment of the present invention;

[0048] Figure 3 Schematic diagram of a dynamic Gaussian clustering module in an embodiment of the present invention;

[0049] Figure 4 This is a schematic diagram of a strategy module in an embodiment of the present invention;

[0050] Figure 5 This is a schematic diagram of a cause-effect module in an embodiment of the present invention;

[0051] Figure 6 Schematic diagram of the parameter self-learning method in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solutions and advantages of this application more clearly understood, this application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0053] This embodiment relies on the attached Figure 1As shown in the flowchart, a dynamic Gaussian semantic field (DGSF) that is updated over time is introduced in the perception layer. The semantic field is in the form of a Gaussian mixture. Sparse expression is performed. Each Gaussian component G i Parameterization is performed using 13 parameter representations: Location information The (x, y, z) center coordinates are referenced to the manipulator base coordinate system; the direction matrix R i ∈SO(3) uses quaternion storage; scale Represents the isotropic radius and stores the anisotropic ε in the form of Cholesky decomposition i ; Existence probability α i ∈(0,1), satisfying ε i α i =1; spherical harmonic color RGB coefficients; material parameters Reflectivity, roughness, etc.; semantic vector Fusion of visual-linguistic features; first-order dynamics (v i ,ω i ), linear velocity and angular velocity; second-order dynamics (a i ,β i ), linear acceleration and angular acceleration; time window Gaussian lifetime center and variance.

[0054] The DGSF has three functions: (1) unified storage of multimodal geometry-semantics; (2) providing differentiable clustering consistency constraints; and (3) providing object-level indexing for subsequent strategic causal graphs.

[0055] During the system initialization phase, the multi-head KAN projection layer processes the four-way multimodal information and generates the initial N0 Gaussian components. After that, whenever the perception thread receives a new frame of multimodal input, it executes the E / M microstepping. Incremental updates enable smooth tracking and adaptive addition and deletion.

[0056] To prevent component jitter, this embodiment calculates the responsibility in the E step and performs μ in the M step. i ,R i ,ε i Perform a Kalman smoothing: where K μ ,K ε is the adaptive gain, is the observed mean.

[0057] When the component is responsible i Continuous T del =30 frames below θ del = 0.01, the background thread will i Mark as pending deletion and remove smoothly in the idle slot of the strategy thread; if a high residual point set appears Meet Maxia ij <θ new =0.05, then select the seed through K-Means++ and add the component G new .

[0058] The dynamic addition and deletion mechanism ensures that N always remains floating (matching the experimental scenario), and the GPU memory peak is controlled within a reasonable range without sacrificing representation accuracy.

[0059] The present invention uses the sparsity of DGSF to compress point cloud points into N Gaussians, and then compresses them into k dynamic Gaussian clusters through dynamic Gaussian clustering, realizing spatiotemporal continuous multimodal environment semantic expression.

[0060] The dynamic Gaussian cluster directly provides a dual-branch structure consisting of a policy module and a causal module to generate trajectories, actions, and policy causal graphs for entity matching, ensuring zero-copy and non-duplicate encoding in the perception-decision link. It also implements parameter self-learning and saves task strategies based on policy feedback consistency constraints, task execution, and closed-loop correction.

[0061] As attached Figure 2 As shown, Gaussian semantic field generation first adopts four-way multimodal synchronous sensing:

[0062] Visual encoder f I : Use improved ResNet-FPN to output multi-scale features Among them, l=3 layers are aligned with the depth map, and TSDF back projection is used to generate 3D feature points

[0063] Point Cloud Encoderf P : Use PointNet++ to output feature point cloud

[0064] Language Encoderf L :Based on CLIP, encode instruction L to get e lang ∈R 384 At the same time, the lexical parser is used to extract the object-action-constraint triples and store them in the temporary knowledge table.

[0065] Robot state encoder f S :The 20-dimensional signals including 6 joint angles, end force-torque, TCP speed, etc. are normalized and then output as state vector e by Dual-MLP state∈R 64 .

[0066] The multi-head KAN projection layer, based on the Kolmogorov-Arnold network (KAN), effectively projects multimodal, high-dimensional features from vision, point cloud, language, and robot state into a unified dynamic Gaussian semantic field through multi-head nonlinear mapping. KAN networks have powerful function approximation capabilities and can accurately express complex nonlinear mapping relationships with a relatively small number of parameters, significantly improving projection efficiency and reducing the model's demand for computational resources.

[0067] Multi-head KAN projection layer network structure: Multi-head KAN projects the multi-scale features, point cloud features, language features, and robot state features generated by the multimodal encoder to generate a Gaussian semantic field. KAN-1 and KAN-2 use geometric consistency loss to ensure that the μ predicted by the RGB-D branch and the point cloud branch is consistent. i ,R i ,s i , To maintain consistency, the geometric consistency calculation expression is:

[0068] KAN-3 ensures that visual and language task instructions are semantically aligned and embedded in the same parameters based on semantic consistency. The semantic consistency calculation expression is:

[0069] KAN-4 uses the temporal consistency of first-order differences and physical prior constraints to ensure that the parameters describing the state of the robot are smooth and accurate. The temporal consistency calculation expression is:

[0070] As attached Figure 2 As shown, the input representation of the cross-modal differential attention clustering layer is for each Gaussian component Select the specific features required for clustering, and then form a unified vector after L2-norm and scale matching. The vector expression is:

[0071] Using the DGSF label prior clustering in the vector to avoid blind global calculation of the attention mechanism, the adjacent constructed table is:

[0072] Sparse adjacency construction calculation expression:

[0073] Sparse neighborhood aggregation to obtain sparse mask M ij ,like When M ij =1, otherwise 0. Different linear heads use corresponding sparse neighborhoods, Query / Key / Value expressions:

[0074] The differential logits are only calculated at the position where the mask is 1, and the calculation expression is:

[0075] The local soft-max calculation obtains the weight, and the calculation expression is:

[0076] Integrate the weights obtained from the attention heads of different sparse neighborhoods and calculate the expression:

[0077] The process of transforming the initial sparse neighborhood into a soft cluster is to let c k is the central feature of the k-th cluster, and the Gaussian feature is calculated using the differential attention weight obtained in the previous step:

[0078] Use Gaussian features as queries to calculate soft cluster membership. The calculation expression is:

[0079] Dynamic Gaussian clustering to obtain Gaussian clusters Cluster-level responsibility ρ k <10 -4 The entire cluster is marked as pending deletion, and the ID is recycled during the next run. The calculation expression is:

[0080] Calculate responsibility i =ε j a ij If r i <10 -3 Freeze G i And it is replaced by the background candidate list in the next cycle.

[0081] Update the calculation expressions of the center, semantic and dynamic parameters:

[0082] The covariance takes the diagonal-Cholesky form: The same applies to the other coordinates, ∈=10 -6 .

[0083] To avoid Gaussian instantaneous jumps within the cluster, exponential smoothing is used where k(i) = argmax k p ik , the remaining parameters R i ,s i , m i Press a for both ij Weighted average update. This operation is completed in one go via a custom CUDA kernel in <3.2ms.

[0084] All update steps are encapsulated as a differentiable "dynamic Gaussian clustering layer", and the gradients can be automatically backpropagated to the cross-modal differential attention weights in PyTorch to achieve end-to-end clustering learning.

[0085] For steady-state training, cluster consistency that considers geometric clustering, semantic clustering, dynamic clustering, and inter-cluster separation is used as a constraint and backpropagated to calculate the expression:

[0086] As attached Figure 4 As shown in Figure 3, the backbone network of the policy module adopts the timestamp + cluster coding encoding layer and vectorization layer + Mamba architecture.

[0087] The first input branch of the dynamic Gaussian cluster is embedded through the timestamp + cluster coding layer: the dynamic Gaussian cluster input is sorted by parameter priority to obtain a feature sequence φ of length m k .

[0088] Timestamp + sinusoidal encoding: where Δt k is the relative time (in seconds) from the current frame to the center of the execution window; P=16.

[0089] The cluster code is obtained by the cross-modal differential attention layer calculation process to obtain the cluster IDc k Decide.

[0090] The second input branch transforms the abstract causal graph and policy causal graph output by the causal module into ψ k , the semantic neighbor edge embedding layer calculation expression:

[0091] Before inputting Mamba, the first and second input branches need to concatenate all embeddings and normalize them:

[0092] The first part of Mamba is a six-layer stacked architecture that implements timing context aggregation: For initial input have:

[0093] The hidden state is pooled using Gated Mean:

[0094] The second part of Mamba Stack 6 more layers of Mamba to get Then the decision decoding is realized through the trajectory head and action head:

[0095] As attached Figure 5 As shown in the figure, the causal module backbone network adopts the architecture of semantic + adjacent edge projection layer, semantic + adjacent edge encoding layer combined with GNN.

[0096] Node semantic projection is to re-take the key fields of the Gaussian parameters of cluster i and implement it through two layers of MLP:

[0097] Neighbor projection selects neighbors based on the distance difference calculation results to generate the projection map:

[0098] Each projected edge carries edge features The calculation formula is as follows:

[0099] The semantic + adjacent edge encoding layer converts each terminal node p into t With the nearest cluster v π(t) Align and obtain the strategy node features:

[0100] Policy edge encoding generates directed edges in trajectory order:

[0101] The first part of the GNN backbone uses the L1 layer Edge-Conditioned GAT to implement abstract causal reasoning:

[0102] The second part of the GNN backbone uses the L2 layer FiLM-GAT to implement strategic causal correction with feedback modulation:

[0103] The new node hidden vector h output by the second part of the GNN backbone * Input semantics + adjacent edge projection layer updates the abstract causal graph, and the last layer of attention Update the policy causal graph as new weights.

[0104] Policy feedback consistency achieves the purpose of calculating loss and backpropagation by aligning constraint trajectories, constraint actions, task results, and policy reasoning with causal reasoning. The calculation expression is:

[0105] As attached Figure 6 As shown, this embodiment completes the phase division and gradient flow of closed-loop self-learning in a simulation environment: "perception-projection-clustering-strategy causal reasoning-decision-making and execution-closed-loop error correction." The entire process consists of six steps: ① multimodal data acquisition and synchronization, ② multimodal projection to a dynamic Gaussian semantic field, ③ dynamic Gaussian clustering, ④ control strategy and abstract causal graph generation, ⑤ real-time robot decision-making and execution, and ⑥ closed-loop error correction.

[0106] During training in step 1, Gazebo-Ignition + NVIDIA Isaac Sim is used to generate multimodal data (RGB-D, LiDAR, DSL commands, joint states). Lighting, material, and dynamic perturbations are randomly inserted to ensure domain randomization.

[0107] In step ②, the KAN projection layer performs a tensor-to-Gaussian differentiable projection on the graphics card and writes the geometric-semantic-temporal consistency constraints into the loss cache in real time; this layer provides a sparse representation for subsequent modules.

[0108] In step ③, cross-modal differential attention uses language query to strengthen clustering weighting; cluster consistency constraints are calculated here.

[0109] The outputs of the policy module and causal module in step ④ are used to calculate the policy feedback consistency loss with the execution feedback results of step ⑤. The loss result will be used as a constraint signal to adjust the attention allocation in Mamba along with the gradient backpropagation.

[0110] In step ⑥, closed-loop correction is to refresh the node completion probability based on the strategy execution result; if the deviation between prediction and execution exceeds the threshold, the difference is written into the total loss function as a "feedback update" and dynamic Gaussian re-clustering is performed.

[0111] The task judgment in step ⑥ is to give a "success / failure" label to each causal path, which is used as the learning reward benchmark for the next epoch.

[0112] Total loss:

[0113] Geometric consistency Center-covariance dual metric,adding photometric-material error.

[0114] Semantic consistency Through visual, semantic pairing (e vis ,e lang ),(e vis ,e lang )(e vis ,e lang ) comparison to achieve semantic alignment.

[0115] Temporal consistency Temporal smoothing using first-order differences and physical prior constraints.

[0116] Cluster consistency Comprehensive consideration of the geometric center μ i , semantic vector Motion parameter v i ,w i ,a i ,β i The difference between the parameters and their cluster centers is used to implement the internal consistency constraint of the cluster through the weighted Euclidean distance penalty; and the inter-cluster distinction is achieved through the separation term of the inter-cluster geometric-semantic distance.

[0117] Strategy feedback consistency It consists of four losses, which respectively measure the deviation between trajectory prediction and execution, action decision error, the accuracy of action effectiveness, and the degree of attention alignment between the policy module (Mamba) and the causal module (GNN), thereby achieving consistency constraints on node execution confidence and edge weight reasoning logic.

[0118] In the early stages of training, to prevent clustering collapse due to initialization perturbations, the system monitors clustering consistency loss. The gradient response of a parameter Exceeding 0.3×||g|| of the current average gradient modulus avg When the training is performed, gradient clipping is performed automatically to suppress the propagation of abnormally amplified clustering errors to subsequent network layers and improve the overall training stability.

[0119] A hybrid training strategy, combining simulation and real-world data, is used to shorten training time and improve system generalization. An 80% simulation and 20% real-world training strategy is employed. The first 50 epochs are trained entirely on simulation data. Every five epochs, 500 frames of real-world data are sampled and added to the training set for lightweight fine-tuning. Real-world data is directly injected into the differentiable computation graph via the KAN projection layer, eliminating the need for manual labeling. The system automatically generates weak labels by combining DSL task instructions with the robot's end-user force / torque feedback, enabling efficient data self-labeling and end-to-end deployment.

[0120] Multi-threaded scheduling and real-time assurance divide perception, decision-making, and execution into three ROS2 real-time threads:

[0121] Perception (P) thread: Responsible for the acquisition and processing of multimodal sensory data, including camera images (RGB-D), point cloud data, the robot's own state, and the construction of an environmental model. For example, the P thread uses deep learning models to extract multi-scale appearance features and local geometric features, fuses them to generate a Gaussian semantic field representation, and performs perception calculations such as dynamic Gaussian clustering. The perception thread typically runs at the sensor frame rate or as high a frequency as possible, with the soft real-time goal of providing timely and up-to-date environmental perception results.

[0122] Control strategy thread C (Control): responsible for real-time decision-making and control strategy execution, it is the core of the system's soft real-time closed loop. The C thread runs at a fixed cycle (10Hz, set according to the servo control requirements), reads the latest perception thread output (such as the current environment Gaussian model and clustering results) in each cycle, and generates control strategies and control instructions based on this. The control strategy includes trajectory planning and action command calculation, such as using the "strategy module" network (such as the attached Figure 4 The C thread's soft real-time goal is to complete decision calculations and command transmission within a strict cycle timeframe, ensuring servo control is not interrupted even with slight lags in sensor data. This ensures stable operation of the control loop, with outputs generated in every cycle and consistent robot motion. If new environmental perception results do not arrive in a timely manner, the control thread will continue execution based on the data or policy predictions from the previous cycle to ensure control continuity.

[0123] Background Learning Thread (L): Responsible for background computing tasks such as policy optimization and causal reasoning, updating the system's cognitive and decision-making models without affecting real-time control. The L thread runs at a lower frequency or during idle periods (e.g., several times per second or triggered on demand), performing incremental learning using data generated by the Perception and Control Thread. Its responsibilities include updating the abstract causal graph (causal relationship module) to reflect the latest environmental changes and robot feedback, adjusting control policy model parameters (such as reinforcement learning parameters and adaptive gains), and monitoring the consistency of policy feedback. Gaussian parameters and decision node indexes are only transferred between threads through a zero-copy ring buffer to avoid jitter caused by large-size images being transferred on the bus.

[0124] Safety fallback mechanism: When the C thread detects that the external force exceeds 40N or the current joint current exceeds 60% of the rated value, the "safety stop" service is immediately triggered; DGSF additions and deletions are frozen; the policy causal graph is cleared and switched to the conservative path set; and a pre-compiled "return to zero" trajectory is executed. After exiting the safety stop state, the system restores the last stable Gaussian set snapshot And restart the clustering layer, which only takes ≈120ms.

[0125] Global re-clustering hot switching strategy: When the background detects the cluster drift index When the frame rate is > 0.8 for 200 consecutive frames, perform full E / M in the GPU spare memory to obtain a new set A remapping table is calculated to ensure node ID stability. A double-ended queue is used to smoothly replace the node within three frames, and Control-T is unaware of the switch.

[0126] The learning thread scans the Replay Buffer every 30 seconds and randomly extracts 32 × 16 = 512 samples to perform incremental gradient analysis: learning rate 1e-5; weight decay 1e-4; if GPU utilization is > 90%, the mini-batch size is automatically reduced to 16. Incremental learning takes ≈140ms and does not block other threads.

[0127] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present invention, and these modifications or replacements should all be included in the scope of protection of the present invention.

Claims

1. A parameter self-learning algorithm based on dynamic Gaussian clustering consistency, characterized in that: The algorithm comprises the following steps: S1. Multimodal Data Acquisition and Synchronization: Using RGB-D cameras, structured light or laser point cloud sensors, industrial DSL language commands, and robotic arm joint encoders, the system acquires real-time and synchronous color images, depth images, dense point cloud data, task text information, and robotic arm status data of the environment. The multimodal data is precisely synchronized using a global clock, ensuring that the synchronization error between modalities is controlled to the nanosecond level. S2. Multimodal Feature Projection and Unified Representation: A multi-level Kolmogorov-Arnold Network (KAN) projection layer is used to extract visual features, point cloud features, language features, and robot state features. These high-dimensional multimodal features are efficiently projected into a unified low-dimensional Dynamic Gaussian Semantic Field (DGSF) through multi-level nonlinear mapping. The dynamic Gaussian semantic field is constrained by both geometric consistency and semantic consistency to ensure that the multimodal feature space within the dynamic Gaussian semantic field maintains a high degree of geometric and semantic consistency. S3. Dynamic Gaussian Clustering with Cross-Modal Differential Attention: This algorithm calculates the composite distance between dynamic Gaussian functions through a differential multi-head attention mechanism to highlight subtle differences between cross-modal features and form dynamic Gaussian clusters with consistent spatiotemporal semantics, thereby achieving stable and continuous representation of objects and environments in dynamic environments. S4. Control Strategy and Abstraction, Policy Causal Graph Generation: Based on dynamic Gaussian clusters, a dual-branch deep neural network architecture is used to simultaneously generate control strategy execution actions and abstraction, policy causal graphs. Specifically, it includes: S41, Action Module: Taking dynamic Gaussian clusters as input, through the Mamba network structure with long-term context modeling capabilities, combined with the results of policy feedback consistency comparison, it generates real-time control strategy trajectories and action sequences; S42, Causal Module: This module uses a graph neural network (GNN) structure and a semantic and edge projection layer to encode the node and edge features of dynamic Gaussian clusters. By dynamically updating the causal graph, it explicitly represents the causal dependencies between task steps, providing a clear and interpretable planning basis during task execution. S5. Real-time Robot Decision-Making and Execution: The robot makes real-time decisions and executes actions based on the control strategies generated by the action modules. The robot also uses execution feedback to update the dynamic Gaussian clusters, abstract causal graphs, and policy causal graphs in real time to ensure that the model continuously and accurately reflects the actual environment state. S6. Closed-loop error correction and online self-learning adjustment: Closed-loop error correction is performed based on the policy feedback consistency loss. When the feedback result of the robot's action does not meet expectations, the policy module is triggered to perform real-time fine-tuning and dynamically adjust the internal model parameters to ensure that the robot adapts to environmental and task changes in real time during task execution.

2. The algorithm according to claim 1, characterized in that The dynamic Gaussian semantic field uses 13 parameters to represent the dynamic Gaussian. Position μ i , direction R i , scale s i , existence probability α i , spherical harmonic color Material parameter m i , semantic vector First-order kinetics (v i ,ω i ), second-order dynamics (a i ,β i ), 4D time window These parameters are projected through a multi-head Kolmogorov-Arnold network to comprehensively and accurately describe the geometry, appearance, semantics, and dynamic characteristics of objects or regions in the environment.

3. The multi-head Kolmogorov-Arnold network according to claim 2, characterized in that Consists of four different KAN heads, KAN-1 and KAN-2 utilize Simultaneously minimize the geometric consistency loss of related projection parameters of RGB-D and 3D point clouds: KAN-3 uses semantic consistency to ensure that visual and language task instructions maintain semantic alignment in projected parameters: KAN-4 ensures smoothness and accuracy of the parameters describing the state of the robot body by using the temporal consistency of first-order differences and physical prior constraints:

4. The algorithm according to claim 1, characterized in that The cross-modal differential attention mechanism is implemented through multiple attention heads, each of which performs differential calculations of attention weights on visual and point cloud modal features, effectively improving sensitivity to changes in dynamic environments and multimodal information.

5. The algorithm according to claim 4, characterized in that The attention kernel function of the four-head sparse attention clustering satisfies: Local soft-max weight: Use Gaussian features as Query to calculate soft cluster membership p ik : For steady-state training, cluster consistency that considers geometric clustering, semantic clustering, dynamic clustering, and inter-cluster separation is used as a constraint and backpropagated to calculate the expression:

6. The algorithm according to claim 1, characterized in that The backbone network Mamba structure of the control strategy execution action realizes accurate prediction and control of the strategy action sequence through temporal context aggregation and normalization processing, based on the timestamp information and cluster encoding information of dynamic Gaussian clusters and combined with the strategy feedback consistency results.

7. The algorithm according to claim 6, characterized in that The recursive gating block of the first phase of the Mamba network satisfies: For initial input have: The hidden state is pooled using Gated Mean: The second part of Mamba Stack 6 more layers of Mamba to get Then the decision decoding is realized through the trajectory head and action head:

8. The algorithm according to claim 1, characterized in that The graph neural network (GNN) structure of the causal module performs semantic encoding and neighboring edge feature encoding on dynamic Gaussian clusters, uses execution feedback to update the causal graph in real time, and explicitly expresses the causal and dependency relationships of each stage in the task to support the robustness and interpretability of strategic decision-making.

9. The algorithm according to claim 8, characterized in that The first part of the GNN backbone uses the L1 layer Edge-Conditioned GAT to implement abstract causal reasoning: The second part of the GNN backbone uses the L2 layer FiLM-GAT to implement strategic causal correction with feedback modulation: The new node hidden vector h output by the second part of the GNN backbone * Input semantics + adjacent edge projection layer updates the abstract causal graph, and the last layer of attention Update the policy causal graph as new weights.

10. The algorithm according to claim 1, characterized in that The closed-loop correction mechanism is based on real-time strategy feedback consistency judgment, performs real-time detection and identification of deviations between feedback and expectations, and automatically triggers fine-tuning and dynamic updating of the control strategy to improve the robustness and adaptability of the strategy.

11. The algorithm according to claim 10, characterized in that The policy feedback consistency achieves the purpose of computing loss and backpropagation by aligning the constraint trajectory, constraint action, task results, and policy reasoning with causal reasoning:

12. The algorithm according to claim 1, characterized in that The implementation of the algorithm is based on a multi-threaded parallel architecture, specifically including a perception thread, a control strategy thread, and a background learning thread. Each thread performs efficient data interaction through zero-copy shared memory, ensuring the real-time and stability of the system in an industrial environment.

13. The algorithm according to claim 1, characterized in that The dynamic Gaussian clustering consistency constraint mechanism ensures that real objects in the environment are continuously represented by the same Gaussian cluster at different times, thereby avoiding jitter problems caused by frequent additions and deletions of components in the model and ensuring the continuity and stability of the environment representation.

14. The algorithm according to claim 1, characterized in that The algorithm uses geometric consistency, semantic consistency, temporal consistency, clustering consistency and policy feedback consistency to construct a composite self-supervised loss function, realizes online dynamic adjustment and optimization of the parameter self-learning mechanism, and can adaptively optimize the control strategy and model representation without human intervention.

15. The algorithm according to claim 14, characterized in that The self-supervised loss function is the total loss function:

16. The algorithm according to claim 1, characterized in that The algorithm adopts a hybrid training strategy, alternating training in simulation environments and real industrial environments. It generates rich and diverse simulation data through Domain-Randomization technology, and supplements it with a small amount of real data for online fine-tuning, effectively improving the robot's ability to quickly generalize and adapt to new task scenarios.

Citation Information

Cited By

  • Humanoid robot operation behavior learning method based on natural language guidance and knowledge reasoning

    CN121625081A

  • A humanoid robot operation behavior learning method based on natural language guidance and knowledge reasoning

    CN121625081B