An ai-based general-purpose work robot intelligent control method and system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SICHUAN LVSHU CONSTR ENG CO LTD
- Filing Date
- 2026-05-21
- Publication Date
- 2026-08-07
AI Technical Summary
该现象出现概率低、持续时间短,却可能使末端越过相邻放空阀触发误操作
1、本发明过在接触前引入预接触微探测机制,构建局部扰动指纹Bi,并结合任务语义与历史异常作业样本建立隐性扰动拓扑图G,使原本难以观测的微小扰动及其传播关系得到显式建模;进一步将初始状态特征φ0与隐性扰动拓扑图G耦合,构建包含动态安全包络与补偿时间窗的数字孪生控制模型M,从而在控制决策阶段引入空间约束与时间约束的协同作用,实现对潜在扰动的提前预测与主动规避。相比现有技术中基于单一状态反馈的控制方式,本发明能够在扰动尚未显著放大之前进行干预,有效降低末端操作偏差及底盘失稳风险,提高复杂环境下作业的安全性与稳定性。
Smart Images

Figure CN122210666B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot intelligent control technology, specifically to an AI-based intelligent control method and system for general-purpose work robots. Background Technology
[0002] In the pressurized sampling valve area of an offshore chemical platform, when a general-purpose robot opens and closes a stainless steel handwheel, condensation and oil film on the valve surface can easily form a momentary lubrication layer. This causes a sudden drop in the handwheel's starting torque, while the chassis compensation lags behind, resulting in millisecond-level reverse torque swing. Although this phenomenon has a low probability of occurrence and a short duration, it can still cause the terminal to overshoot adjacent vent valves, triggering malfunctions. Existing methods mostly follow a segmented control approach of "identification-planning-execution," making it difficult to predict such hidden disturbances using minute pre-contact responses and to collaboratively suppress them at the moment of contact. Summary of the Invention
[0003] The purpose of this invention is to provide an AI-based intelligent control method and system for general-purpose work robots to address the shortcomings in the prior art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: an AI-based intelligent control method for general-purpose work robots, comprising: S100, acquire the initial state features φ0 of the robot, including body posture, joint torque, chassis attachment parameters and visual features of target components; S200 controls the end effector to perform a pre-contact micro-probing action on the target component, and collects the displacement response, vibration response, image blurring phase difference and contact force strain anomaly coefficient during the micro-probing process to generate the local disturbance fingerprint Bi corresponding to the target component. S300, Bi is coupled with the semantics of the current task and historical abnormal task samples for analysis to construct the implicit perturbation topology map G corresponding to the target area; S400, based on φ0 and G, establish a digital twin control model M that includes a dynamic safety envelope and a compensation time window; S500 inputs the real-time status characteristics during the operation into M, and outputs the chassis anchoring compensation amount, the robotic arm anti-torsion compensation amount and the end trajectory correction amount through AI calculation, and performs synchronous control according to the compensation time window. S600 updates Bi, G, and M online based on the measured displacement deviation and contact response after execution until the target operation is completed.
[0005] Preferably, the method for obtaining the contact force strain anomaly coefficient includes the following steps: The contact force and strain time series collected during the micro-probe process are segmented and modeled to construct a corresponding family of probability distribution models, and the distribution parameters are uniformly normalized. The probability distribution model family is embedded into the information geometry space, and a Riemannian metric tensor is constructed based on the Fisher information matrix to form the corresponding statistical manifold structure. By selecting the reference distribution corresponding to the historical normal contact process as the benchmark point, the geodesic distance of the current distribution on the statistical manifold is calculated to obtain the distribution offset. The distribution offset is mapped to a preset stability threshold to generate a contact force strain anomaly coefficient.
[0006] Preferably, generating the local perturbation fingerprint Bi corresponding to the target component includes the following steps: The displacement response, vibration response, image blur phase difference, and contact force strain anomaly coefficient collected during the micro-detection process are time-synchronized and multi-source aligned to construct a unified temporal feature matrix; Based on the unified temporal feature matrix, a multi-scale time window is used to decompose each feature and extract the set of transient perturbation sub-features at different time scales. The transient perturbation sub-feature set is input into a non-Euclidean feature embedding model for structure mapping to obtain a perturbation structure vector that reflects the multi-source coupling relationship. The perturbation structure vector is encoded, compressed, and mapped to generate a local perturbation fingerprint Bi with a unique identifier.
[0007] Preferably, constructing the latent perturbation topology map G corresponding to the target region includes the following steps: The semantics of the current task are structured and parsed to extract the task action sequence, contact object attributes and operation constraint parameters, and the task semantics are encoded into semantic feature vectors. The local perturbation fingerprint Bi is aligned and mapped with the semantic feature vector, a joint feature sequence is constructed based on the temporal correlation, and the perturbation features corresponding to the historical abnormal operation samples are used as the reference sequence. Based on the similarity relationship between the joint feature sequence and the reference sequence, node connection weights are constructed, and a multi-node association structure is established using the local perturbation fingerprint Bi as a node to form an initial topology graph. The initial topology graph is iteratively updated with edge weights and path filtering is performed to extract high-risk disturbance propagation paths and generate a hidden disturbance topology graph G.
[0008] Preferably, the step of establishing a digital twin control model M containing a dynamic safety envelope and a compensation time window based on φ0 and G includes the following steps: The initial state features φ0 are mapped to the robot's kinematics and dynamics parameter space to construct a basic set of state constraints. High-risk disturbance nodes and their propagation paths are extracted based on the implicit disturbance topology graph G to form a set of disturbance constraints. The basic state constraint set and the disturbance constraint set are fused together, and a dynamic safety envelope that varies with time is generated by the constraint boundary expansion method, wherein the envelope boundary is adaptively adjusted according to the disturbance propagation intensity. Based on the dynamic security envelope, the temporal delay characteristics corresponding to the disturbance propagation path are calculated, and the compensation time interval is divided accordingly to construct the compensation time window; The dynamic safety envelope and the compensation time window are coupled and embedded into the control model to form a digital twin control model M for real-time control.
[0009] Preferably, the step of inputting the real-time status features during the operation into M, and outputting the chassis anchoring compensation amount, the robotic arm anti-torsion compensation amount, and the end-effector trajectory correction amount through AI calculation includes the following steps: The real-time state characteristics during the operation process are continuously collected and uniformly time-series aligned to form a real-time state characteristic sequence, which is then mapped to the state variable space corresponding to the digital twin control model M. The real-time state feature sequence is input into the digital twin control model M. Based on the dynamic safety envelope constraint and the disturbance propagation path of the implicit disturbance topology graph G, the state at future time is predicted, and the chassis anchoring compensation, robotic arm anti-torsion compensation, and end-point trajectory correction are generated. Based on the compensation time window, each compensation quantity is synchronized and aligned in time to determine the effective time and duration of each compensation quantity, thus forming a time-series compensation control sequence. The timing compensation control sequence is applied to the robot's execution control process to achieve coordinated compensation control between the chassis and the robotic arm.
[0010] Preferably, the online update of Bi, G, and M based on the measured displacement deviation and contact response after execution includes the following steps: The measured displacement deviation and contact response during the execution process are synchronously acquired and time-aligned to construct a feedback feature sequence, and the feedback feature sequence is mapped to the feature space corresponding to the local disturbance fingerprint Bi. The difference between the feedback feature sequence and the current local perturbation fingerprint Bi is calculated, and the local perturbation fingerprint Bi is incrementally updated based on the difference result to generate the updated local perturbation fingerprint. The updated local perturbation fingerprint is mapped onto the latent perturbation topology graph G, and the node weights and path connectivity are adjusted to obtain the updated latent perturbation topology graph. The updated implicit disturbance topology map and feedback feature sequence are input into the digital twin control model M. The dynamic safety envelope and compensation time window in the model are modified to form the updated digital twin control model M, which is then used for subsequent operation process control.
[0011] This invention also provides an AI-based general-purpose intelligent control system for work robots, comprising: Initial state perception module: acquires the initial state features φ0 of the robot, including body posture, joint torque, chassis attachment parameters and visual features of target components; Disturbance fingerprint generation module: controls the end effector to perform pre-contact micro-probing action on the target component, collects displacement response, vibration response, image blur phase difference and contact force strain anomaly coefficient during the micro-probing process, and generates local disturbance fingerprint Bi corresponding to the target component; Topology construction module: Couples Bi with the semantics of the current job task and historical abnormal job samples for analysis to construct the implicit perturbation topology graph G corresponding to the target region; Safety constraint generation module: Establishes a digital twin control model M containing a dynamic safety envelope and a compensation time window based on φ0 and G; Collaborative compensation control decision module: Input the real-time status characteristics during the operation into M, and output the chassis anchoring compensation amount, robotic arm anti-torsion compensation amount and end trajectory correction amount through AI calculation, and perform synchronous control according to the compensation time window; Model Adaptive Module: Updates Bi, G, and M online based on the measured displacement deviation and contact response after execution until the target task is completed.
[0012] The technical effects and advantages provided by the present invention in the above technical solution are as follows: 1. This invention introduces a pre-contact micro-detection mechanism before contact to construct a local disturbance fingerprint Bi, and establishes a latent disturbance topology G by combining task semantics and historical abnormal operation samples, enabling explicit modeling of previously difficult-to-observe minute disturbances and their propagation relationships. Furthermore, it couples the initial state feature φ0 with the latent disturbance topology G to construct a digital twin control model M containing a dynamic safety envelope and a compensation time window. This introduces the synergistic effect of spatial and temporal constraints during the control decision-making stage, achieving early prediction and proactive avoidance of potential disturbances. Compared to existing control methods based on single-state feedback, this invention can intervene before disturbances significantly amplify, effectively reducing end-operation deviations and chassis instability risks, and improving the safety and stability of operations in complex environments.
[0013] 2. This invention generates chassis anchoring compensation, robotic arm anti-torsion compensation, and end-effector trajectory correction by inputting real-time state characteristics into a digital twin control model M. These are then combined with a compensation time window for synchronous control. Simultaneously, based on the measured displacement deviation and contact response after execution, the local disturbance fingerprint Bi, the latent disturbance topology G, and the digital twin control model M are updated online, forming a closed-loop control process of "perception-modeling-prediction-control-feedback-adaptation." This closed-loop update mechanism enables the model to continuously adapt to environmental changes and disturbance evolution characteristics, avoiding the problems of static models and lag response in traditional control methods. This significantly improves the system's ability to suppress low-probability, high-risk transient disturbances, enhances control accuracy, and has outstanding practical value. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0015] Figure 1 This is a flowchart of the method of the present invention.
[0016] Figure 2 This is a flowchart of the system modules of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Example 1, please refer to Figure 1 As shown in this embodiment, a general-purpose intelligent control method for operational robots based on AI includes: In one embodiment of the present invention, step S100 is used to obtain the initial state features φ0 corresponding to the working robot, including body posture, joint torque, chassis attachment parameters and visual features of target components. The specific steps are as follows: After the robot enters the target work area, the body inertial measurement unit, joint encoder, torque sensor, chassis drive feedback module and vision acquisition module first synchronously collect the multi-source state information of the robot at the start of the operation. The multi-source state information is then processed by time alignment, coordinate unification and normalization to form the initial state feature φ0, which provides a unified input for subsequent micro-detection, latent disturbance identification and compensation control.
[0019] Taking the pressurized sampling valve area of an offshore chemical platform as an example, the operating robot needs to travel along a 0.9m wide metal grid channel to the sampling valve assembly to perform opening and closing operations on the 180mm diameter stainless steel handwheel. The ambient humidity in this area is typically greater than 92%, with condensation on the valve assembly surface and a localized oil film thickness of 0.1mm to 0.3mm. The platform equipment operates with continuous mechanical vibrations of 10Hz to 35Hz, and there is significant specular reflection on the surfaces of adjacent pipelines. In this scenario, the robot's attitude can be obtained by an inertial measurement unit, including roll angle, pitch angle, heading angle, and their rates of change; for example, when the robot is docked in front of the valve, the roll angle is 1.8°, the pitch angle is 2.3°, and the heading deviation is 0.6°. The joint torques are obtained from the drive feedback of each joint of the robotic arm. For example, in the operating posture, the static holding torques of each joint of a six-axis robotic arm can be 12.4 N·m, 18.7 N·m, 9.6 N·m, 4.2 N·m, 2.8 N·m, and 1.9 N·m, respectively, to characterize the load distribution and structural preload state under the current configuration. The chassis adhesion parameters are jointly determined by the drive wheel speed, encoder displacement, drive current, and ground contact estimation model, and may include wheel-to-ground adhesion coefficient, slip ratio, and contact stability index. For example, the adhesion coefficient measured on a wet and slippery grid surface is 0.41, and the instantaneous slip ratio is 3.2%. The visual features of the target component are acquired by a visible light camera and a structured light module, including the target component's outline size, edge sharpness, reflectivity, surface texture sparsity, and spatial pose information; for example, the target handwheel's center coordinates are identified as being 1.26m in front of the robot base, 0.18m to the left, and 0.94m high, with the average reflectivity of the handwheel's edge reaching 218 levels of image grayscale, and the number of effective feature points of the local texture being less than 35.
[0020] In this invention, the initial state feature φ0 is not a simple concatenation of individual parameters, but rather a mapping of the robot's posture, joint torques, chassis attachment parameters, and visual features of the target component into a state vector that can be invoked by the intelligent control model according to a unified time scale. This allows the local disturbance fingerprint obtained in subsequent steps to be established based on the actual initial working conditions. In other words, if the initial state feature indicates that the robot already has a slight tilt, insufficient chassis attachment margin, and a highly reflective surface on the target component, the system can predict in advance that the robot is more likely to experience a hidden risk of stable end-effector recognition but amplified chassis micro-slippage when contacting the target component.
[0021] S200 controls the end effector to perform a pre-contact micro-probing action on the target component, and collects the displacement response, vibration response, image blurring phase difference and contact force strain anomaly coefficient during the micro-probing process to generate the local disturbance fingerprint Bi corresponding to the target component.
[0022] In one embodiment of the present invention, the control end effector performs a pre-contact micro-probing action on the target component. Specifically, before a stable contact is formed, the end effector gradually approaches the surface of the target component with a preset micro-displacement step size, and simultaneously collects displacement response, vibration response, image blur phase difference and contact force strain anomaly coefficient during each micro-displacement step, thereby forming multi-source time-series data.
[0023] Furthermore, the contact force and strain time series collected during the micro-probe process are segmented and modeled to construct a family of probability distribution models. Specifically, the time series is collected at a sampling frequency of 1000 Hz and segmented according to a time window length of 20 milliseconds and a sliding step size of 5 milliseconds. For each time segment, a Gaussian mixture distribution is used for fitting, and its probability density function is expressed as: the probability density functions of each sub-component are linearly weighted and summed according to their weights; where each sub-component is determined by the mean and variance; the parameters are solved using the maximum likelihood estimation method, with the objective of maximizing the product of the probability densities of all sampling points or the sum of their logarithms; after modeling, the mean parameter is processed using a linear normalization method to make it fall into the interval between 0 and 1, and the variance parameter is logarithmically transformed and then normalized.
[0024] Furthermore, the family of probability distribution models is embedded in the information geometric space to construct a statistical manifold structure. Specifically, let the probability density function be p(x|θ), where θ is a parameter vector. The elements of the Fisher information matrix are calculated as follows: the partial derivatives of the logarithmic probability density function with respect to the parameters are multiplied, and the expectation is obtained by integrating the variable x over its domain, thus obtaining the values of each element of the matrix. This matrix is used as a Riemannian metric tensor to define the distance between two points in the parameter space, thereby forming a statistical manifold.
[0025] Furthermore, a reference distribution corresponding to a historical normal contact process is selected as the benchmark point to calculate the geodesic distance of the current distribution. Specifically, the path is discretized into multiple parameter point sequences. For any two adjacent points, the distance is calculated by the quadratic form of the corresponding parameter difference vector and the Fisher information matrix, i.e., the parameter difference vector is multiplied by the metric tensor, then multiplied by its own transpose and the square root is taken. The path length is obtained by summing the distances of all adjacent points on the path. The path is continuously adjusted using the gradient descent method to minimize the path length, and finally the geodesic distance is obtained as the distribution offset.
[0026] Furthermore, the distribution offset is mapped to a contact force strain anomaly coefficient. Specifically, let the distribution offset be d, the baseline threshold be d0, and the upper limit threshold be d1. First, the distribution offset d is normalized as follows: the normalized value is equal to "the distribution offset minus the baseline threshold d0, then divided by the difference between the upper limit threshold d1 and the baseline threshold d0", limiting its value range to between 0 and 1. Based on this, the normalized value is input into an exponential function for transformation. The exponential function is an exponential function with the natural constant as the base. Its output value is normalized again by subtracting 1 and then dividing by the exponent of the natural constant minus 1, so that the final result is limited to the interval between 0 and 1. When the distribution offset is close to the baseline threshold d0, the output result approaches 0; when the distribution offset is close to or exceeds the upper limit threshold d1, the output result approaches 1.
[0027] In one embodiment of the present invention, the generation of the local perturbation fingerprint Bi includes the following steps: First, time synchronization processing is performed on displacement response, vibration response, image blur phase difference, and contact force strain anomaly coefficient. The sampling frequency is uniformly set to 500 Hz, and the time offset between different signals is compensated by linear interpolation method to construct a unified time series feature matrix.
[0028] Secondly, a multi-scale time window is used to decompose the unified temporal feature matrix. The first-order difference and the second-order difference are calculated in each time window. The first-order difference is the difference between the eigenvalues of adjacent time points, and the second-order difference is the difference between the first-order differences. The local energy is obtained by summing the squares of all eigenvalues in the window. The features under different time scales are combined to form a set of transient perturbation sub-features.
[0029] Furthermore, the transient perturbation sub-feature set is mapped to a graph structure, where nodes represent sub-features and edge weights are calculated using cosine similarity, specifically the ratio of the dot product of two feature vectors to the product of their respective magnitudes. After normalizing the adjacency relation matrix, graph convolution is performed. The output of each layer is obtained by multiplying the adjacency matrix and the feature matrix and applying a nonlinear function. After three layers of propagation, all node features are converged to obtain the perturbation structure vector.
[0030] Finally, the perturbation structure vector is compressed and encoded. Principal component analysis is used to extract principal components with a cumulative contribution rate of 95%. The reduced vector is then multiplied by the random projection vector, and the result is binarized according to the sign. 0 greater than or equal to 1 is mapped to 1, and less than 0 is mapped to 0, thus generating a fixed-length binary code as the local perturbation fingerprint Bi.
[0031] In one embodiment of the present invention, step S300 is used to couple and analyze Bi with the semantics of the current task and historical abnormal task samples to construct a latent perturbation topology map G corresponding to the target region, specifically including the following steps: The semantics of the current task are structured and parsed to extract the task action sequence, contact object attributes, and operation constraint parameters, and the task semantics are encoded into a semantic feature vector.
[0032] Specifically, the task is input as text or instruction sequence, and the task content is decomposed into multiple ordered sub-actions to form a task action sequence, such as "approach the target", "position the handwheel", "apply rotational force", etc. Each sub-action is arranged in the order of execution. At the same time, the contact object attributes are extracted from the task description, including the geometric dimensions, surface material, surface roughness and spatial position parameters of the target part. The geometric dimensions are represented by length, diameter or angle values, and the surface material is converted into a numerical identifier through a preset material coding table. Furthermore, the operation constraint parameters are extracted, including the maximum allowable contact force, the maximum displacement deviation and the operating speed range.
[0033] After extracting the above information, the task action sequence is converted into a vector representation using sequential encoding, where each action corresponds to a fixed-dimensional code. The encoding method maps the action number to a unit vector. The contact object attributes and operation constraint parameters are normalized using the max-min normalization method, which linearly maps each parameter to the interval between 0 and 1. Finally, the action encoding vector, object attribute vector, and constraint parameter vector are concatenated to form a semantic feature vector with a unified dimension for subsequent coupling analysis.
[0034] The local perturbation fingerprint Bi is aligned and mapped with the semantic feature vector, a joint feature sequence is constructed based on the temporal correlation, and the perturbation features corresponding to the historical abnormal operation samples are introduced as a reference sequence.
[0035] Specifically, the local perturbation fingerprints Bi are arranged in the order of their generation time to form a perturbation fingerprint time series. At the same time, the semantic feature vectors are expanded according to the task action sequence so that each time segment corresponds to the semantic feature of the currently executed action. The perturbation fingerprint time series and the semantic feature vector sequence are synchronously matched through a timestamp alignment method. The time alignment is achieved through the minimum time difference principle, that is, the semantic feature with the smallest time difference is selected as the corresponding item, thereby constructing a joint feature sequence.
[0036] Furthermore, sample data consistent with the current task type are extracted from historical abnormal operation samples, and corresponding perturbation fingerprint sequences and semantic feature sequences are generated in the same way; this historical sequence is introduced as a reference sequence; to ensure comparability, the current joint feature sequence and the reference sequence are processed to have the same length, and linear interpolation is used to pad the shorter sequence so that the two have the same length.
[0037] Based on the similarity relationship between the joint feature sequence and the reference sequence, node connection weights are constructed, and a multi-node association structure is established using the local perturbation fingerprint Bi as a node to form an initial topological relationship graph.
[0038] Specifically, each local perturbation fingerprint Bi in the joint feature sequence is taken as a node in the graph structure; the connection weight between any two nodes is obtained by calculating the similarity between their corresponding feature vectors. The similarity is calculated using cosine similarity, that is, the dot product of the two vectors is divided by the product of their respective magnitudes; at the same time, the corresponding node in the reference sequence is introduced to calculate the similarity between the current node and the historical abnormal node, which is used to correct the connection weight.
[0039] To enhance anomaly sensitivity, the connection weights are adjusted by weighting. When a node has a high similarity to a historical anomaly sample, its connection weight is increased. The adjustment method is to multiply the original similarity by an amplification factor, which is calculated based on the frequency of historical anomalies. Finally, the weighted connection relationships between nodes are obtained, and an adjacency relationship matrix is constructed based on this matrix. An initial topological relationship graph is established based on the adjacency relationship matrix, where nodes represent local perturbation fingerprints Bi, and edges represent the association relationships between perturbations.
[0040] The initial topology graph is iteratively updated with edge weights and path filtering is performed to extract high-risk disturbance propagation paths and generate a hidden disturbance topology graph G.
[0041] Specifically, the edge weights in the initial topology graph are iteratively updated. In each iteration, the edge weights are adjusted according to the global connectivity strength of the node, where the global connectivity strength of the node is calculated by the sum of the connectivity weights of the node and all its neighboring nodes. The edge weights and node strengths are normalized and then reassigned. The number of iterations is set to 10.
[0042] After completing the weight update, all paths in the topology graph are traversed and calculated. The risk value of each path is obtained by multiplying the weights of each edge on the path. A path screening threshold is set. When the risk value of a path is greater than the preset threshold, the path is marked as a high-risk disturbance propagation path. The threshold is determined by the 90th percentile of the risk values of historical abnormal paths.
[0043] Finally, all high-risk disturbance propagation paths and their corresponding node sets are extracted to construct a latent disturbance topology graph G. The latent disturbance topology graph G is used to characterize the potential propagation relationships and risk evolution trends of disturbances within the target area.
[0044] In one embodiment of the present invention, step S400 is used to establish a digital twin control model M containing a dynamic safety envelope and a compensation time window based on φ0 and G, specifically including the following steps: The initial state features φ0 are mapped to the robot's kinematics and dynamics parameter space to construct a basic set of state constraints. High-risk disturbance nodes and their propagation paths are extracted based on the implicit disturbance topology graph G to form a set of disturbance constraints.
[0045] Specifically, the initial state feature φ0 includes the body posture, joint torques, chassis attachment parameters, and visual features of the target component. First, the body posture is converted into a homogeneous transformation matrix using Euler angles to describe the position and orientation of the robot base in space. The joint torques and joint angles are input into the robot's kinematic equations to obtain the pose and mechanical response relationship of the end effector. The chassis attachment parameters are converted into contact constraint coefficients, where the adhesion coefficients and the maximum driving force are determined by a linear relationship, i.e., the maximum driving force is equal to the product of the adhesion coefficients and the normal load. The visual features of the target component are converted into spatial coordinate constraints.
[0046] After completing the above mapping, all constraint parameters are uniformly represented as inequality constraints, that is, each constraint is represented as a limit on the range of variable values, thereby constructing a basic set of state constraints.
[0047] Simultaneously, nodes with edge weights greater than a preset threshold are extracted from the latent perturbation topology graph G as high-risk perturbation nodes. The threshold is taken as the 80th percentile of the edge weights in the latent perturbation topology graph. The connection paths between high-risk perturbation nodes are traversed to extract perturbation propagation paths. The path risk intensity is calculated based on the edge weights in the path, which is obtained by multiplying the edge weights on the path. Paths with path risk intensity greater than the preset path threshold are included in the perturbation constraint set, thereby forming a perturbation constraint set corresponding to the basic state constraint set.
[0048] The basic state constraint set and the disturbance constraint set are fused together, and a dynamic safety envelope that varies with time is generated by the constraint boundary expansion method.
[0049] Specifically, each constraint boundary in the basic state constraint set is used as an initial safety boundary, and it is expanded according to the path risk intensity in the disturbance constraint set. For any constraint boundary, its expansion range is proportional to the corresponding disturbance path risk intensity, that is, the greater the risk intensity, the larger the expansion range of the constraint boundary. The expansion method is to add or subtract an expansion amount based on the original upper and lower bounds of the constraint. This expansion amount is obtained by multiplying the path risk intensity by a proportional coefficient, which is determined based on historical safe operation data.
[0050] To achieve dynamic changes, a time parameter is introduced into the constraint boundary, causing the expansion amount to change over time. The law of change is described by an exponential decay function, that is, the expansion amount gradually decreases as time increases. The function input is the ratio of time to the initial expansion amount. In this way, a set of constraint boundaries that change over time is obtained, and the set of constraint boundaries together constitutes the dynamic safety envelope.
[0051] Based on the dynamic security envelope, the timing delay characteristics corresponding to the disturbance propagation path are calculated, and the compensation time interval is divided accordingly to construct the compensation time window.
[0052] Specifically, for each disturbance propagation path, the propagation time is calculated based on the connection relationship between each node in the path; the propagation time between a single node is obtained by the reciprocal of the weight between the nodes, that is, the larger the weight, the shorter the propagation time; the propagation times between all nodes on the path are summed to obtain the total propagation time of the path; further, combined with the time scale of the current robot's action, the total propagation time of the path is mapped to a time delay feature.
[0053] After obtaining the time delay feature, it is aligned with the time parameters in the dynamic safety envelope, and time intervals are divided. When the time delay feature is in a certain interval, it is considered that there is a potential disturbance effect in that interval. The interval is used as the compensation interval to form a compensation time window. Multiple disturbance paths correspond to multiple compensation time windows, and a set of compensation time windows is formed by time sorting.
[0054] The dynamic safety envelope and the compensation time window are coupled and embedded into the control model to form a digital twin control model M.
[0055] Specifically, the dynamic safety envelope is embedded as a state constraint in the control variable space, that is, the range of values of the control variables is restricted from exceeding the dynamic safety envelope during the solution process; at the same time, the compensation time window is introduced as a time constraint into the control process, and the control strategy is adjusted within the compensation time window so that the control output meets the disturbance suppression requirements.
[0056] In practical implementation, the control variable is represented as a time function, and the constraints corresponding to the dynamic safety envelope are transformed into restrictions on the range of values of the function. Within the compensation time window, the rate of change of the control function is adjusted to advance or delay the response to disturbance effects. The adjustment amount is determined jointly based on the disturbance path risk intensity and time delay characteristics.
[0057] Ultimately, by simultaneously considering spatial and temporal constraints, a digital twin control model M is formed to simulate the robot's operating state in the actual environment in virtual space and to provide a constraint basis for the generation of subsequent control commands.
[0058] In one embodiment of the present invention, step S500 is used to input the real-time status characteristics during the operation into M, output the chassis anchoring compensation amount, the robotic arm anti-torsion compensation amount, and the end-effector trajectory correction amount through AI calculation, and perform synchronous control according to the compensation time window, specifically including the following steps: The real-time state characteristics during the operation are continuously collected and uniformly time-series aligned to form a real-time state characteristic sequence, which is then mapped to the state variable space corresponding to the digital twin control model M.
[0059] Specifically, during the robot's operation, the robot continuously collects body posture, joint angle, joint torque, chassis drive status, and end effector pose information at a sampling frequency of 500 Hz to form multi-source state data. All types of data are resampled according to a unified time reference, and linear interpolation is used to compensate for the time offset between different sensor data, thereby constructing a time-consistent real-time state feature sequence.
[0060] After time alignment is completed, the real-time state feature sequence is mapped to the state variable space of the digital twin control model M. Among them, the body posture is converted into a rotation matrix representation through Euler angles, the joint angles and joint torques are converted into end pose and mechanical response through kinematic equations, and the chassis driving state is converted into motion constraint variables through the relationship between wheel speed and driving force. The above variables are uniformly represented as state vectors for subsequent model calculations.
[0061] The real-time state feature sequence is input into the digital twin control model M. Based on the dynamic safety envelope constraint and the disturbance propagation path of the implicit disturbance topology graph G, the state at future time is predicted, and the chassis anchoring compensation, robotic arm anti-torsion compensation, and end-point trajectory correction are generated.
[0062] Specifically, the state vector is used as the input variable, and the state evolution is calculated in the digital twin control model M. The state evolution is realized by discrete-time recursion, that is, the current state is updated to the state at the next time step through the control input and the influence of disturbance. The influence of disturbance is determined by the disturbance propagation path in the hidden disturbance topology graph G, and its effect intensity is weighted by the path risk intensity.
[0063] During the prediction process, a dynamic safety envelope is introduced as a constraint, which restricts the state variables from exceeding the safety envelope boundary at each prediction time. When the predicted state approaches or exceeds the boundary, the corresponding compensation amount is calculated. Among them, the chassis anchoring compensation amount is achieved by adjusting the chassis driving force, which is calculated as the difference between the current driving force and the driving force allowed by the safety boundary. The robotic arm anti-torsion compensation amount is achieved by adjusting the joint torque, which is calculated as the inverse vector of the difference between the current torque and the constraint torque. The end-effector trajectory correction amount is obtained by calculating the position deviation, which is the spatial difference between the predicted trajectory and the safe trajectory.
[0064] Based on the compensation time window, the time synchronization and alignment of each compensation quantity is performed to determine the effective time and duration of each compensation quantity, thus forming a time-series compensation control sequence.
[0065] Specifically, various compensation quantities are matched with compensation time windows; for each compensation quantity, its effective start time and end time are determined based on the compensation time window associated with its corresponding disturbance propagation path; when the compensation time window covers multiple time segments, the compensation quantity is applied to each time segment in segments.
[0066] To ensure the smoothness of compensation, a transition processing is performed on the compensation amount at the boundary of the time window. A linear interpolation method is used to make the compensation amount gradually change at the beginning and end of the window. Through the above processing, the compensation amounts are arranged in chronological order to form a continuous time-series compensation control sequence.
[0067] By applying a timing compensation control sequence to the robot's execution control process, collaborative compensation control between the chassis and the robotic arm can be achieved.
[0068] Specifically, the chassis anchoring compensation is applied to the chassis drive control variables, and the chassis position is stabilized by adjusting the wheel speed and driving force; the robotic arm anti-torsion compensation is applied to the drive input of each joint, and the attitude is stabilized by changing the joint torque; the end effector trajectory correction is superimposed on the target trajectory, so that the end effector runs along the corrected trajectory.
[0069] During the control process, each compensation quantity is executed sequentially in time sequence, and the status feedback is updated in real time. When the state variable is detected to re-enter the dynamic safety envelope, the compensation quantity is gradually reduced until the normal control state is restored. Through the synergistic effect of the chassis and robotic arm compensation quantities, the latent disturbances are suppressed in advance and dynamically corrected.
[0070] In one embodiment of the present invention, step S600 is used to update Bi, G, and M online based on the measured displacement deviation and contact response after execution until the target operation is completed, specifically including the following steps: The measured displacement deviation and contact response during the execution process are synchronously acquired and time-aligned to construct a feedback feature sequence, which is then mapped to the feature space corresponding to the local disturbance fingerprint Bi.
[0071] Specifically, during the robot's execution control process, the displacement deviation between the actual trajectory and the target trajectory of the end effector is collected at a sampling frequency of 500 Hz, while contact force, contact strain, and vibration response signals are also collected. The displacement deviation is defined as the difference between the actual position and the target position, and its spatial representation is a three-dimensional coordinate difference. The contact response is represented as a combined vector of contact force and strain.
[0072] The aforementioned multi-source data are processed synchronously according to a unified time reference. A linear interpolation method is used to align data from different sampling times to form a continuous time series. Using time as an index, the displacement deviation and contact response are combined to construct a feedback feature sequence.
[0073] To maintain consistency with the local perturbation fingerprint Bi, the feedback feature sequence is mapped to the same feature space. The mapping method involves normalizing the feedback features and reconstructing them according to the encoding dimension of the local perturbation fingerprint Bi, so that they have a consistent vector dimension and distribution range.
[0074] The difference between the feedback feature sequence and the current local perturbation fingerprint Bi is calculated, and the local perturbation fingerprint Bi is incrementally updated based on the difference result to generate the updated local perturbation fingerprint.
[0075] Specifically, for each feature vector at each time point in the feedback feature sequence, the difference between it and the local perturbation fingerprint Bi at the corresponding time point is calculated. The difference is obtained by Euclidean distance, which is the square root of the sum of the squares of the differences in each dimension of the two vectors. The difference values at all time points are weighted and averaged, where the weight is determined according to the distance of the time from the current time, and the closer the time is to the current time, the greater the weight.
[0076] Based on the calculated average difference value, the local perturbation fingerprint Bi is updated. The update method involves linearly fusing the original local perturbation fingerprint with the feedback feature. The fusion ratio is determined by the difference value; when the difference value is large, the weight of the feedback feature is increased. Specifically, the updated local perturbation fingerprint is equal to the original fingerprint multiplied by one weight coefficient plus the feedback feature multiplied by another weight coefficient, with the sum of the two weights being 1. To ensure update stability, the updated local perturbation fingerprint is normalized so that its dimensional values remain within the range of 0 to 1.
[0077] The updated local perturbation fingerprint is mapped onto the latent perturbation topology graph G, and the node weights and path connectivity are adaptively adjusted to obtain the updated latent perturbation topology graph. Specifically, the updated local perturbation fingerprint is used as the new node feature to replace the original node features; the similarity between the node and other nodes is recalculated using cosine similarity, which is the ratio of the dot product of two vectors to the product of their magnitudes; the connection weights between nodes are updated based on the similarity. All edge weights in the topology graph are normalized so that the sum of the connection weights of each node is 1; simultaneously, the risk intensity of the perturbation propagation path is recalculated based on the updated weights, and the risk intensity is obtained by multiplying the weights of each edge on the path. When the change in the connection weight of a node exceeds a preset threshold, its related paths are reconstructed, with the threshold set at a weight change ratio of 0.2; through the above update process, a latent perturbation topology graph reflecting the latest perturbation relationships is obtained.
[0078] The updated implicit disturbance topology and feedback feature sequence are input into the digital twin control model M, and the parameters of the dynamic safety envelope and compensation time window are corrected to form the updated digital twin control model M, which is then used for subsequent operation process control.
[0079] Specifically, the path risk intensity in the updated implicit disturbance topology is used as an input parameter to adjust the boundary of the dynamic safety envelope. The adjustment method is to linearly superimpose the original safety boundary with the risk intensity. The greater the risk intensity, the greater the contraction of the safety boundary, thereby increasing the strictness of the control constraints.
[0080] Meanwhile, the compensation time window is corrected based on the time delay changes in the feedback feature sequence; the time delay is determined by the difference between the peak displacement deviation time and the disturbance propagation time; when a delay change is detected, the start time and duration of the compensation time window are adjusted synchronously.
[0081] The updated dynamic safety envelope and compensation time window are re-embedded into the control model, replacing the original parameters, to form an updated digital twin control model M. In subsequent control processes, the above update steps are repeated until the target operation is completed.
[0082] Example 2, please refer to Figure 2 As shown in this embodiment, a general-purpose intelligent control system for work robots based on AI includes: Initial state perception module: acquires the initial state features φ0 of the robot, including body posture, joint torque, chassis attachment parameters and visual features of target components; Disturbance fingerprint generation module: controls the end effector to perform pre-contact micro-probing action on the target component, collects displacement response, vibration response, image blur phase difference and contact force strain anomaly coefficient during the micro-probing process, and generates local disturbance fingerprint Bi corresponding to the target component; Topology construction module: Couples Bi with the semantics of the current job task and historical abnormal job samples for analysis to construct the implicit perturbation topology graph G corresponding to the target region; Safety constraint generation module: Establishes a digital twin control model M containing a dynamic safety envelope and a compensation time window based on φ0 and G; Collaborative compensation control decision module: Input the real-time status characteristics during the operation into M, and output the chassis anchoring compensation amount, robotic arm anti-torsion compensation amount and end trajectory correction amount through AI calculation, and perform synchronous control according to the compensation time window; Model Adaptive Module: Updates Bi, G, and M online based on the measured displacement deviation and contact response after execution until the target task is completed.
[0083] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A general-purpose intelligent control method for operational robots based on AI, characterized in that, include: S100, acquire the initial state features φ0 of the robot, including body posture, joint torque, chassis attachment parameters and visual features of target components; S200, before stable contact is established, the end effector is controlled by a preset micro-displacement step size to perform a pre-contact micro-probing action on the target component, and the displacement response, vibration response, image blurring phase difference, and contact force strain anomaly coefficient are collected during the micro-probing process to generate a local disturbance fingerprint Bi corresponding to the target component, including the following steps: The displacement response, vibration response, image blur phase difference, and contact force strain anomaly coefficient collected during the micro-detection process are time-synchronized and multi-source aligned to construct a unified temporal feature matrix; Based on the unified temporal feature matrix, a multi-scale time window is used to decompose each feature and extract the set of transient perturbation sub-features at different time scales. The transient perturbation sub-feature set is input into a non-Euclidean feature embedding model for structure mapping to obtain a perturbation structure vector that reflects the multi-source coupling relationship. The perturbation structure vector is encoded, compressed, and mapped to generate a local perturbation fingerprint Bi with a unique identifier; S300, couple Bi with the semantics of the current task and historical abnormal task samples to construct the implicit perturbation topology map G corresponding to the target region, including the following steps: The semantics of the current task are structured and parsed to extract the task action sequence, contact object attributes and operation constraint parameters, and the task semantics are encoded into semantic feature vectors. The local perturbation fingerprint Bi is aligned and mapped with the semantic feature vector, a joint feature sequence is constructed based on the temporal correlation, and the perturbation features corresponding to the historical abnormal operation samples are used as the reference sequence. Based on the similarity relationship between the joint feature sequence and the reference sequence, node connection weights are constructed, and a multi-node association structure is established using the local perturbation fingerprint Bi as a node to form an initial topology graph. The initial topology graph is subjected to edge weight iterative update and path filtering to extract high-risk disturbance propagation paths and generate a hidden disturbance topology graph G. S400, based on φ0 and G, establish a digital twin control model M that includes a dynamic safety envelope and a compensation time window; S500 inputs the real-time status characteristics during the operation into M, and outputs the chassis anchoring compensation amount, the robotic arm anti-torsion compensation amount and the end trajectory correction amount through AI calculation, and performs synchronous control according to the compensation time window. S600 updates Bi, G, and M online based on the measured displacement deviation and contact response after execution until the target operation is completed.
2. The AI-based intelligent control method for a general-purpose work robot according to claim 1, characterized in that: The method for obtaining the contact force strain anomaly coefficient includes the following steps: The contact force and strain time series collected during the micro-probe process are segmented and modeled to construct a corresponding family of probability distribution models, and the distribution parameters are uniformly normalized. The probability distribution model family is embedded into the information geometry space, and a Riemannian metric tensor is constructed based on the Fisher information matrix to form the corresponding statistical manifold structure. By selecting the reference distribution corresponding to the historical normal contact process as the benchmark point, the geodesic distance of the current distribution on the statistical manifold is calculated to obtain the distribution offset. The distribution offset is mapped to a preset stability threshold to generate a contact force strain anomaly coefficient.
3. The AI-based intelligent control method for a general-purpose work robot according to claim 1, characterized in that: The establishment of a digital twin control model M, which includes a dynamic safety envelope and a compensation time window, based on φ0 and G, includes the following steps: The initial state features φ0 are mapped to the robot's kinematics and dynamics parameter space to construct a basic set of state constraints. High-risk disturbance nodes and their propagation paths are extracted based on the implicit disturbance topology graph G to form a set of disturbance constraints. The basic state constraint set and the disturbance constraint set are fused together, and a dynamic safety envelope that varies with time is generated by the constraint boundary expansion method, wherein the envelope boundary is adaptively adjusted according to the disturbance propagation intensity. Based on the dynamic security envelope, the temporal delay characteristics corresponding to the disturbance propagation path are calculated, and the compensation time interval is divided accordingly to construct the compensation time window; The dynamic safety envelope and the compensation time window are coupled and embedded into the control model to form a digital twin control model M for real-time control.
4. The AI-based intelligent control method for a general-purpose work robot according to claim 1, characterized in that: The process of inputting real-time status features during the operation into M, and then outputting chassis anchoring compensation, robotic arm anti-torsion compensation, and end-effector trajectory correction through AI calculation includes the following steps: The real-time state characteristics during the operation process are continuously collected and uniformly time-series aligned to form a real-time state characteristic sequence, which is then mapped to the state variable space corresponding to the digital twin control model M. The real-time state feature sequence is input into the digital twin control model M. Based on the dynamic safety envelope constraint and the disturbance propagation path of the implicit disturbance topology graph G, the state at future time is predicted, and the chassis anchoring compensation, robotic arm anti-torsion compensation, and end-point trajectory correction are generated. Based on the compensation time window, each compensation quantity is synchronized and aligned in time to determine the effective time and duration of each compensation quantity, thus forming a time-series compensation control sequence. The timing compensation control sequence is applied to the robot's execution control process to achieve coordinated compensation control between the chassis and the robotic arm.
5. The AI-based intelligent control method for a general-purpose work robot according to claim 1, characterized in that: The online update of Bi, G, and M based on the measured displacement deviation and contact response after execution includes the following steps: The measured displacement deviation and contact response during the execution process are synchronously acquired and time-aligned to construct a feedback feature sequence, and the feedback feature sequence is mapped to the feature space corresponding to the local disturbance fingerprint Bi. The difference between the feedback feature sequence and the current local perturbation fingerprint Bi is calculated, and the local perturbation fingerprint Bi is incrementally updated based on the difference result to generate the updated local perturbation fingerprint. The updated local perturbation fingerprint is mapped onto the latent perturbation topology graph G, and the node weights and path connectivity are adjusted to obtain the updated latent perturbation topology graph. The updated implicit disturbance topology map and feedback feature sequence are input into the digital twin control model M. The dynamic safety envelope and compensation time window in the model are modified to form the updated digital twin control model M, which is then used for subsequent operation process control.
6. An AI-based intelligent control system for a general-purpose work robot, used to implement the AI-based intelligent control method for a general-purpose work robot as described in any one of claims 1-5, characterized in that: include: Initial state perception module: acquires the initial state features φ0 of the robot, including body posture, joint torque, chassis attachment parameters and visual features of target components; Disturbance fingerprint generation module: controls the end effector to perform pre-contact micro-probing action on the target component, collects displacement response, vibration response, image blur phase difference and contact force strain anomaly coefficient during the micro-probing process, and generates local disturbance fingerprint Bi corresponding to the target component; Topology construction module: Couples Bi with the semantics of the current job task and historical abnormal job samples for analysis to construct the implicit perturbation topology graph G corresponding to the target region; Safety constraint generation module: Establishes a digital twin control model M containing a dynamic safety envelope and a compensation time window based on φ0 and G; Collaborative compensation control decision module: Input the real-time status characteristics during the operation into M, and output the chassis anchoring compensation amount, robotic arm anti-torsion compensation amount and end trajectory correction amount through AI calculation, and perform synchronous control according to the compensation time window; Model Adaptive Module: Updates Bi, G, and M online based on the measured displacement deviation and contact response after execution until the target task is completed.
Citation Information
Patent Citations
Machine fault diagnosis method and device based on information geometry
CN107121975A
Real-time transmission method and system based on internet-of-things perception data in digital twinborn scene
CN120143775A
Force sense feedback control method of intelligent mechanical arm and control system thereof
CN121756348A