Spatial motion attitude modeling method and device of rope-driven flexible arm and medium
Through the adaptive spatial motion posture analysis network and multi-task modeling method, the high-dimensional strong nonlinearity and behavioral coupling problems of the rope-driven flexible arm are solved, the precise modeling of the flexible arm's spatial motion posture is achieved, and the accuracy and generalization ability of the modeling are improved.
Patent Information
- Application Number
- CN202510751305.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-23
AI Technical Summary
The spatial motion posture of a rope-driven flexible arm has high-dimensional strong nonlinear and behavioral coupling characteristics, which are difficult to be effectively solved by existing modeling methods, resulting in inaccurate modeling and insufficient generalization ability.
An adaptive spatial motion posture parsing network is adopted, including a hybrid multi-dimensional scaling module and a spatiotemporal adaptive attention mechanism, combined with a multi-task modeling network. By collecting the spatial motion posture information of the flexible arm and the control signal of the servo motor, a new dataset is constructed and modeled using an adaptive gradient balancing mechanism.
It effectively eliminates the adverse effects of high-dimensional strong nonlinearity and behavioral coupling, improves the accuracy and generalization ability of rope-driven flexible arm modeling, and realizes accurate modeling of the flexible arm's spatial motion posture.
Smart Images

Figure CN120680496A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of rope-driven flexible arms and relates to a method, equipment and medium for modeling the spatial motion posture of a rope-driven flexible arm. Background Art
[0002] Rope-driven flexible arms offer greater degrees of freedom and flexibility, and offer greater compliance and safety than rigid robotic arms in environmental interaction and human-machine collaboration. Their unique drive system reduces their bulk, enabling them to operate effectively in a variety of complex, confined environments. These arms offer significant advantages over rigid robotic arms, with promising applications in areas such as safe inspection of submarine oil pipeline structures and maintenance of narrow pipelines.
[0003] Modeling of rope-driven flexible arms primarily focuses on the relationship between their control input and end position. Existing modeling methods include mathematical mechanism modeling and neural network modeling. Mathematical mechanism modeling uses idealistic assumptions to establish a mathematical expression for the relationship between control input and end position. For example, patent application publication number CN111421529A discloses a control method for a rope-driven flexible arm. This method calculates the length change of the driving rope of the operating-end flexible arm by obtaining the length change of the driving rope of the working-end flexible arm. Based on this length change, the length-time function of the driving rope of the working-end flexible arm is calculated to drive the working-end flexible arm.
[0004] Neural network modeling utilizes a data-driven approach. By collecting control input and end position data pairs, a neural network model is used to train the mapping relationship between control input and end position. Common neural network models include BP neural networks and LSTM neural networks. Prior art, such as the invention patent application publication number CN118493364A, discloses a method for end position control of a rope-driven flexible arm. This method uses a BP neural network trained based on a dataset to predict end position during the dynamics model training phase. During the reinforcement learning agent training phase, the dynamics model predictions and reinforcement learning algorithms are used to train an agent consisting of an actor and critic network. The LSTM network enhances time series processing capabilities to achieve optimal control action learning.
[0005] However, if a rope-driven flexible arm is to be used in complex and confined spaces, modeling its spatial motion is even more crucial. However, the high-dimensional, strong nonlinearity and behavioral coupling of the rope-driven flexible arm's spatial motion make it difficult to effectively model it using mathematical modeling methods. Neural network modeling offers significant advantages, however. However, these characteristics have a significant negative impact on neural network learning. For example, excessively high data dimensionality and behavioral coupling can prevent the network from balancing the importance of different information, leading to negative transfer and other issues. Therefore, it is crucial to eliminate the negative impact of these high-dimensional, strong nonlinearity and behavioral coupling on modeling and accurately model the rope-driven flexible arm's spatial motion. Summary of the Invention
[0006] The technical solution of the present invention is used to solve the problem of how to eliminate the adverse effects of high-dimensional strong nonlinearity and behavioral coupling on modeling of a rope-driven flexible arm, thereby improving the accuracy of modeling of the rope-driven flexible arm.
[0007] The present invention solves the above technical problems through the following technical solutions:
[0008] A method for modeling the spatial motion posture of a rope-driven flexible arm comprises the following steps:
[0009] S1, collects the spatial motion posture information of the flexible arm and the control signal of the servo steering gear to construct the original data set;
[0010] S2, parsing the spatial motion posture information based on an adaptive spatial motion posture parsing network; the adaptive spatial motion posture parsing network includes a hybrid multi-dimensional scaling module and a spatiotemporal adaptive attention mechanism module;
[0011] S3, embedding the control signal based on the encoding module to extract the hidden features in the control signal;
[0012] S4, constructs a new dataset based on the processed spatial motion posture information and control signals, and builds a multi-task modeling network with an adaptive gradient balancing mechanism.
[0013] Furthermore, the S1 includes:
[0014] S11, collect the control signals of the servo motor at different times t in, represents the control signal of the i1th servo actuator at time t, where N is the total number of servo actuators;
[0015] S12, collect the spatial motion posture information of the flexible arm through the motion capture system, and record the spatial motion posture information of the flexible arm at different times t in, represents the position information of the kth point on the i2th metal disk at time t, and P represents the total number of metal disks.
[0016] Furthermore, the S2 includes:
[0017] S21, decoupling and reducing the dimensionality of the spatial motion posture information using a hybrid multidimensional scaling transform; the hybrid multidimensional scaling transform module includes a metric multidimensional scaling transform and a non-metric multidimensional scaling transform;
[0018] S22, uses the spatiotemporal adaptive attention mechanism module to capture the spatial motion characteristics of spatial motion posture information.
[0019] Furthermore, the S21 includes:
[0020] S211, mapping the high-dimensional data to a low-dimensional space based on the metric multidimensional scaling transformation, and performing dimensionality reduction mapping on the relative position information in the original position, specifically:
[0021] The location information of the sample points after dimensionality reduction is recorded as The metric multidimensional scaling transformation is transformed into an optimization problem to solve the following formula:
[0022]
[0023] Among them, D position Represents the Euclidean distance between the original points, D M-MDS Represents the Euclidean distance between points after dimensionality reduction;
[0024] S212, based on non-metric multidimensional scaling transformation, minimizes the difference between the mapped point order and the original point order, and approaches the low-dimensional representation of the original structure, specifically:
[0025] Note the different points between the flexible arms The distance between different points is expressed as M1 represents the total number of points used to collect the flexible arm posture information;
[0026] Identify a set of dissimilarities The non-metric multidimensional scaling transformation is transformed into a minimization problem to solve the following formula:
[0027]
[0028] in, represents the distance between the i-th and j-th dissimilar points in the dimensionality reduction space, Represents the distance between the kth dissimilar point and the lth dissimilar point after dimensionality reduction, represents the distance between the i-th and n-th dissimilar points in the dimensionality reduction space, represents the distance between the jth and nth dissimilar points in the dimensionality reduction space, represents the distance between the kth and nth dissimilar points in the dimensionality reduction space, f(·) represents the loss function, k={1,2,....,M1}, l={1,2,....,M1}, k≠l≠j, represents the position of the i-th different point at time t, represents the position of the jth different point at time t, represents the position of the kth different point at time t;
[0029] S213, fusing the metric multidimensional scaling transformation result and the non-metric multidimensional scaling transformation result, using the following logic representation:
[0030]
[0031] in, It represents the spatial motion posture information of the flexible arm at different times t after dimensionality reduction mapping, and ⊕ represents the exclusive-OR operation.
[0032] Furthermore, the S22 includes:
[0033] S221, the information after dimensionality reduction mapping Divide into S×S non-overlapping areas, and use the adaptive adjustment factor ε to perform differential operations on different sliding window areas to obtain
[0034] S222, obtain the query matrix corresponding to the u-th region at time t through linear mapping Bond Matrix and the value matrix
[0035] S223, considering the influence relationship between different regions, find the influence degree corresponding to different key values through the traction matrix, and find the region that should participate in each given region; specifically, calculate the query vector and key vector of each region, use cosine similarity to calculate the correlation traction matrix, and obtain the traction matrix
[0036] S223, using the traction matrix between regions Calculate the degree of traction between different areas; specifically, With the traction matrix Combine to get in, represents the correlation bond matrix, represents the correlation value matrix, represents the Ronecker product;
[0037] S224, calculate the attention coefficient of each area through the traction matrix Remove the influence of redundant and invalid information on the area of attention and improve the accuracy of attention;
[0038] S225, based on a multi-channel convolutional network, fuses multiple convolution kernels to extract directional local features. Specifically, the square matrix convolution kernel is fused with the row and column vector convolution kernels, using the following logic representation:
[0039]
[0040] in, Represents the result after square matrix convolution kernel processing, Represents the result after column convolution kernel processing, Represents the result after row convolution kernel processing, It represents the result after the dimensionality reduction processing and the attention concentration processing, Indicates k s ×k s is the size of the convolution kernel, and the depth convolution is performed. The number of convolution channels is g. Indicates 1×k c is the size of the convolution kernel, and the depth convolution is performed. The number of convolution channels is g. Indicates k r ×1 is the size of the convolution kernel, depthwise convolution is performed, and the number of convolution channels is g;
[0041] S226, fuses the results of the multi-channel convolutional network, uses regularization and multi-layer perceptron to describe one-dimensional feature information, and outputs the parsing results of the adaptive spatial motion posture parsing network.
[0042] Furthermore, the S3 is specifically:
[0043] Through the linear mapping layer, the control input information is expanded and the nonlinear characteristics contained in the control input information are extracted, which is expressed as The data is encoded and a mask mechanism is added to dynamically shield the redundant information that exists when the information is expanded. The following logic is used to represent it:
[0044]
[0045] in, is the result after feature encoding, is the result after mask processing, Indicates the actual maximum control output of the servo motor, M represents the randomly initialized mask unit, M N,T Indicates the mask unit corresponding to the Nth control information unit at time T, M mask Represents the mask matrix.
[0046] Furthermore, the S4 includes:
[0047] S41, combines the multi-expert network, weight regulation network and output mapping network to build a multi-task modeling network;
[0048] S42, a gradient paradigm for parameter updates of shared parameters defined based on multi-task modeling losses and dynamic weights that fluctuate over time;
[0049] S43, initialize the target gradient, use cosine similarity to calculate whether the target gradient and the original gradient between any tasks conflict, if a conflict occurs, the target gradient will be updated by vector projection according to the conflict relationship between the conflicting gradients, and the following logic is used to represent the update of the target gradient:
[0050]
[0051] in, represents the i3th network parameter update gradient, represents the j3th network parameter gradient, Indicates the initialization target gradient, take
[0052] Furthermore, 60% of all the data in the new dataset was used as a training set, 20% as a validation set, and 20% as a test set, and a multi-task modeling network was constructed based on the torch framework of Python 3.8.
[0053] An electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above-mentioned spatial motion posture modeling method of the rope-driven flexible arm, and the processor is configured to execute the program stored in the memory.
[0054] A storage medium stores a computer program, which, when run by a processor, executes the steps of the above-mentioned method for modeling the spatial motion posture of a rope-driven flexible arm.
[0055] The advantages of the present invention are:
[0056] The present invention provides a method for modeling the spatial motion posture of a rope-driven flexible arm. The method first collects the spatial motion posture information of the flexible arm and the control information corresponding to the servo steering gear to form an original data set. The method then uses a hybrid multidimensional scaling transform and a spatiotemporal adaptive attention mechanism to analyze the spatial motion posture information. The hybrid multidimensional scaling transform decouples and reduces the dimensionality of the spatial motion posture, converting the flexible arm's spatial motion posture information into easily processable two-dimensional coordinate information. The spatiotemporal adaptive attention mechanism further captures the spatial motion characteristics of the flexible arm in both time and space, thereby completing the processing of the flexible arm's spatial motion posture information. The method then embeds and encodes the control information to fully extract the hidden characteristics of the control information. Finally, the processed posture information and control information are combined into a new data set, and the modeling is completed using a multi-task modeling network with an adaptive gradient balancing mechanism. The adaptive gradient balancing mechanism can prevent problems such as negative transfer that occur during network learning. The modeling method provided by the present invention can adaptively analyze the spatial motion posture of the flexible arm, effectively addressing the effects of high-dimensional strong nonlinearity and behavioral coupling, and improving the accuracy and generalization ability of network training. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 This is a flow chart of a method for modeling the spatial motion posture of a rope-driven flexible arm according to a first embodiment of the present invention;
[0058] Figure 2 This is a structural block diagram of spatial motion posture modeling of a rope-driven flexible arm according to the first embodiment of the present invention;
[0059] Figure 3 Schematic diagram of the physical structure of the rope-driven flexible arm according to the first embodiment of the present invention;
[0060] Figure 4 Schematic diagram of the multi-task modeling network training results of the first embodiment of the present invention;
[0061] Figure 5 This is a modeling effect diagram of the "U" posture of the rope-driven flexible arm of the first embodiment of the present invention;
[0062] Figure 6 This is a modeling effect diagram of the "S" posture of the rope-driven flexible arm of the first embodiment of the present invention;
[0063] Figure 7 This is a schematic diagram of various modeling methods under different modeling effect indicators according to the first embodiment of the present invention. DETAILED DESCRIPTION
[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0065] The technical solution of the present invention is further described below with reference to the accompanying drawings and specific embodiments:
[0066] Example 1
[0067] like Figure 3 As shown in the figure, the proposed method for modeling the spatial motion posture of a rope-driven flexible arm is applied to a rope-driven flexible arm entity. The rope-driven flexible arm consists of a silicone rod, 10 metal discs, and four transmission cables. The rope-driven flexible arm is divided into two sections, each equipped with five metal discs. The upper section is towed by four transmission cables, and the lower section is towed by two transmission cables, providing a higher degree of freedom. The motion of the rope-driven flexible arm is coordinated and controlled by four servo motors and four transmission cables. Specifically, the second and third servo motors control the upper section, while the first and fourth servo motors control the lower section.
[0068] The metal disc in the upper part is evenly perforated with 8 cable holes along the circumference. Among them, 4 cable holes are used to pass through the transmission cables of the first section, and the remaining 4 cable holes are used to pass through the 4 transmission cables of the second section. The 4 transmission cables of the first section are fixed to the bottom of the fifth metal disc, and the other ends are fixed to the winding shaft driven by the second servo motor and the third servo motor. The spacing between adjacent transmission cables is set to 1 / 4 of the circumference (90°), and the separation between the two transmission cables on the same winding shaft is 1 / 2 of the circumference (180°). Therefore, when the winding shaft driven by the servo motor rotates, one pair of transmission cables relaxes while the other pair of transmission cables tightens, maintaining the uniformity of their slack and tension lengths.
[0069] The metal disc in the lower half has four evenly spaced upper cable holes along its circumference. One end of the transmission cable is anchored to the 10th metal disc, and the other end is fixed to the winding shaft driven by the first and fourth servo motors. The entire soft manipulator is 31 cm long, and the diameter of each upper cable loop is 4 mm. The specific structural diagram is shown in the figure below. Figure 3 As shown in Part A; Figure 3 Part B in the figure is a schematic diagram of the main structure of the rope-driven flexible arm and the servo drive structure.
[0070] The high-dimensional strong nonlinearity and behavioral coupling characteristics of the rope-driven flexible arm during spatial motion make it impossible for common neural network modeling methods to effectively model the flexible arm's spatial motion posture. This paper proposes a new neural network modeling method that can adaptively analyze the flexible arm's spatial motion posture, effectively solving the impact caused by high-dimensional strong nonlinearity and behavioral coupling.
[0071] Specifically, a dataset consisting of the spatial motion posture information of the flexible arm and the corresponding control information of the servo actuators is collected. This spatial motion posture information is then parsed using a hybrid multidimensional scaling transform and a spatiotemporal adaptive attention mechanism. The hybrid multidimensional scaling transform performs decoupling and dimensionality reduction, converting the flexible arm's spatial motion posture information into easily processable two-dimensional coordinate information. The spatiotemporal adaptive attention mechanism further captures the spatial motion characteristics of the flexible arm in both time and space, thus completing the processing of the flexible arm's spatial motion posture information. The control information is embedded and encoded to fully extract the hidden features within the control information. The processed posture and control information are combined into a new dataset, and modeling is completed using a multi-task modeling network with an adaptive gradient balancing mechanism. The adaptive gradient balancing mechanism prevents recurring negative transfer and other issues during network learning.
[0072] like Figure 1 Specifically, a method for modeling the spatial motion posture of a rope-driven flexible arm is disclosed, comprising the following steps:
[0073] S1, collecting the spatial motion posture information of the flexible arm and the control signal of the servo steering gear to construct an original data set; specifically, S1 includes:
[0074] S11, collect the control signals of the servo motor at different times t in, represents the control signal of the i1th servo actuator at the tth moment. In this embodiment, the control signal of the servo actuator is specifically the rotation angle of the servo actuator; N is the total number of servo actuators.
[0075] In order to prevent the unreasonable shape change of the flexible arm, this embodiment limits the control signal value of the servo steering engine. The maximum value of the control signal is Where T represents the total time for the flexible arm to perform the control task, express The maximum value of .
[0076] S12, collect the spatial motion posture information of the flexible arm through the motion capture system, and record the spatial motion posture information of the flexible arm at different times t in, Represents the three point information on the i2-th metal disk at time t, for example represents the first point information on the i2-th metal disk, P represents the total number of metal disks, k represents the k-th point on the metal disk, and the collected points are on the circumference of the metal disk, with different points separated by 120 degrees.
[0077] Furthermore, the spatial motion posture information of the flexible arm is characterized by the skeleton information of the metal disk. For the skeleton information of each metal disk, the circumscribed circle corresponding to the triangle at the edge can be used to represent the position information of the metal disk corresponding to the triangle. in, Represents the information of the three points in the X-axis direction at time t; Represents the information of the three points in the Y-axis direction at time t; Indicates the information of the three points in the Z-axis direction at time t.
[0078] Furthermore, the information of the circumscribed circle center is solved through the three vertices of the triangle position in space, specifically: Map the triangulated point information to the center information of the metal disk, where v1 and v2 represent the distance vector between two different points; ||v1|| 2 and ||v2|| 2 Represents the Euclidean norm of the two distance vectors, ‖‖[v1×v2]‖‖ 2 Indicates that the cross product of two distance vectors is first performed, and then the Euclidean norm is calculated; in this embodiment, k1, k2, and k3 represent three different points on the metal disk.
[0079] In this embodiment, the motion capture system consists of sixteen NOKOV high-precision motion capture cameras distributed around the flexible arm, completely covering the flexible arm's workspace and performing real-time dynamic data acquisition to collect information about the flexible arm's spatial motion posture. This embodiment uses 30 target balls to capture the spatial position of the flexible arm's overall posture. The motion capture system operates at a frequency of 20 Hz. The servo motor transmits current data in real time at a speed of 40 steps per second. Timestamps are aligned to ensure that the servo motor data matches the motion capture data.
[0080] By systematically traversing the reachable space of the flexible arm's spatial postures, a dataset containing 100,000 valid data points was collected over 28 hours. The resulting task space corresponds to an approximate circle with a radius of 150 mm and a center of (200 mm, 200 mm) along the X and Y axes. The Z-axis range is (100 mm, 320 mm), concentrated in the region between 200 mm and 320 mm. Furthermore, the distribution of the task space along the Z axis is uneven, due to the high-dimensional, nonlinear motion characteristics of the flexible arm.
[0081] S2, parsing the spatial motion posture information based on an adaptive spatial motion posture parsing network; the adaptive spatial motion posture parsing network includes a hybrid multidimensional scaling module and a spatiotemporal adaptive attention mechanism module; the hybrid multidimensional scaling module is used to decouple and reduce the spatial motion posture information, and the spatiotemporal adaptive attention mechanism module is used to capture the spatial motion characteristics of the spatial motion posture information. S2 includes the following steps:
[0082] Based on the spatial motion posture information of the flexible arm and the control signal of the servo motor obtained in step S1, the modeling of the spatial motion posture is converted into finding the mapping relationship between the control signal and the posture information. Since the dimension of the collected flexible arm point information is too high, if modeling is performed directly, the redundancy of irrelevant information between different data will cause negative feedback due to the high data dimension, which will reduce the utilization rate of the information and affect the modeling effect. Therefore, the present invention provides an adaptive spatial posture parsing network that can extract features from multi-dimensional and complex spatial point information, ensure that the information is not over-processed, and realize effective parsing of the flexible arm's spatial motion posture.
[0083] Specifically, the adaptive spatial motion posture parsing network includes a hybrid multidimensional scaling module and a spatiotemporal adaptive attention mechanism module; wherein, the hybrid multidimensional scaling module is used to realize the dimensionality reduction mapping of information in three-dimensional space to two-dimensional space, as well as the representation of the coupling relationship between different segments of the flexible arm, and the spatiotemporal adaptive attention mechanism module is used to highlight the motion characteristics of the flexible arm in the spatial and temporal dimensions.
[0084] S21, decoupling and reducing the dimensionality of spatial motion posture information using hybrid multidimensional scaling transformation; the hybrid multidimensional scaling transformation module includes metric multidimensional scaling transformation and non-metric multidimensional scaling transformation; S21 includes:
[0085] Multidimensional scaling is a data dimensionality reduction method based on correlation analysis between data. It can ensure the distance and position relationships between data while mapping the dimensions, which is consistent with the spatial motion posture characteristics of the flexible arm.
[0086] S211, mapping the high-dimensional data to a low-dimensional space based on the metric multidimensional scaling transformation, and performing dimensionality reduction mapping on the relative position information in the original position, specifically:
[0087] The location information of the sample points after dimensionality reduction is recorded as The metric multidimensional scaling transformation is transformed into an optimization problem to solve formula (1), and the following logic is used to express formula (1):
[0088]
[0089] Among them, Dposition Represents the Euclidean distance between the original points, D M-MDS Represents the Euclidean distance between points after dimensionality reduction. This embodiment introduces the magnification matrix Z position It is used to amplify the information difference between points. T represents the matrix transposition operation. In this embodiment, the following logic is used to represent D position With Z positi0n :
[0090]
[0091] In order to facilitate the calculation and highlight the relative position relationship between samples, the center matrix is introduced Where I represents the unit matrix, e represents the unit vector, and the zero mean processing of the position information is performed to obtain the inner product matrix The optimization problem of formula (1) is transformed into the optimization problem of solving formula (4), and the dimensionality reduction mapping of the relative position information in the original position is completed. Formula (4) is expressed using the following logic:
[0092]
[0093] Among them, B M-MDS Represents the inner product matrix obtained by performing zero mean processing on the matrix after dimension reduction, and takes
[0094] S212, based on non-metric multidimensional scaling transformation, minimizes the difference between the mapped point order and the original point order, and approaches the low-dimensional representation of the original structure, specifically:
[0095] Note the different points between the flexible arms The distance between different points can be expressed as M1 represents the total number of points used to collect the flexible arm posture information.
[0096] At this point, we need to find a set of different points The order of the original points can be preserved as much as possible, and the non-metric multidimensional scaling transformation is transformed into a minimization problem to solve formula (5). The following logic is used to express formula (5):
[0097]
[0098] in, represents the distance between the i-th and j-th dissimilar points in the dimensionality reduction space, Represents the distance between the kth dissimilar point and the lth dissimilar point after dimensionality reduction, represents the distance between the i-th and n-th dissimilar points in the dimensionality reduction space, represents the distance between the jth and nth dissimilar points in the dimensionality reduction space, represents the distance between the kth and nth dissimilar points in the dimensionality reduction space, f(·) represents the loss function, k={1,2,....,M1}, l={1,2,....,M1}, k≠l≠j, represents the position of the i-th different point at time t, represents the position of the jth different point at time t, represents the position of the kth different point at time t.
[0099] In this embodiment, if and only ifΔ ij <Δ kl ,k≠l≠j means that the distance relationship between any two different points after dimensionality reduction remains the same as before dimensionality reduction; and It means that the adjacent different points after dimensionality reduction can maintain the coplanarity of the adjacent different points before dimensionality reduction and the positional relationship of the triangle formed by the three adjacent different points.
[0100] In this embodiment, f(·) uses the maximum entropy loss to minimize the difference between the mapped point order and the original point order, thereby obtaining a low-dimensional representation that is as close as possible to the original structure.
[0101] S213, fusing the metric multidimensional scaling transformation result and the non-metric multidimensional scaling transformation result, using the following logic representation:
[0102]
[0103] in, It represents the spatial motion posture information of the flexible arm at different times t after dimensionality reduction mapping, and ⊕ represents the exclusive-OR operation.
[0104] By fusing the results of metric multidimensional scaling transformation and non-metric multidimensional scaling transformation, the pose information after dimensionality reduction mapping is obtained. Based on hybrid multi-dimensional scaling transformation, the spatial pose of the flexible manipulator is mapped while preserving the position and order relationships between its points to the greatest extent.
[0105] S22, uses the spatiotemporal adaptive attention mechanism module to capture the spatial motion characteristics of spatial motion posture information.
[0106] The commonly used spatial attention mechanism has problems such as insufficient global information capture and excessive loss of local features due to fixed perception and single information capture type, and cannot take into account the changing characteristics of information in the time dimension, which makes it difficult to meet the needs of capturing the motion characteristics of the flexible arm. Therefore, the present invention proposes a spatiotemporal adaptive attention mechanism that depicts the motion characteristics of the manipulator from two dimensions, time and space, and adaptively adjusts the degree of global information capture through a trainable adjustment factor to prevent excessive loss of information and invalid extraction. Specifically, the S22 includes:
[0107] S221, the information after dimensionality reduction mapping Divide into S×S non-overlapping areas of different sizes, and perform differential operations on different sliding window areas using the adaptive adjustment factor ε to obtain The following logic is used:
[0108]
[0109] in, represents the difference operation result of the u-th region at time t, Information of the u-th region at time t, The information of the u-th region at time t-1, ε represents the adaptive adjustment factor.
[0110] In this embodiment, each non-overlapping region contains r×c / S 2 eigenvectors, so In the above formula, u = {1, 2, …, r × c / S} is taken, where S represents the row and column size of the non-overlapping region, r represents the row size of the region before partitioning, c represents the column size of the region before partitioning, and ε is obtained through multi-layer perceptron (MLP) training.
[0111] S222, using the following logic to express the linear mapping to obtain
[0112]
[0113] in, They represent the query matrix, key matrix, and value matrix corresponding to the u-th region at time t, respectively. They represent the parameter vectors of the query matrix, key matrix, and value matrix corresponding to the u-th region at time t.
[0114] S223, considering the influence relationship between different regions, find the influence degree corresponding to different key values through the traction matrix, and find the region that should participate in each given region, specifically:
[0115] First, the query matrix and key matrix of each region are calculated, and then the correlation between different regions is calculated using cosine similarity to obtain the traction matrix The traction matrix is represented by the following logic:
[0116]
[0117] in, They represent the query matrix and key matrix corresponding to the w-th region at time t, w={1,2,....,r×c / S}, u≠w, and ⊙ represents the XOR operation.
[0118] S223, using the traction matrix between regions Calculate the degree of traction between different areas.
[0119] Since the correlation region may be scattered in the global space, it is necessary to With the traction matrix Combine to get in, represents the correlation bond matrix, represents the correlation value matrix, represents the Kronecker product.
[0120] S224, calculate the attention coefficient of each area through the traction matrix, remove the influence of redundant invalid information on the attention area, and improve the accuracy of attention, specifically:
[0121] Through the above steps, useful information in the global scope is extracted, so that the spatiotemporal attention coefficient can be calculated This can remove the influence of redundant and invalid information on the attention area and improve the accuracy of attention, which can be expressed using the following logic:
[0122]
[0123] Among them, π t (u,w) represents the influence between the u-th region and the w-th region at time t, LeakyReLu(·) represents the activation function, represents the attention corresponding to the u-th region at time t, Indicates the area with positive correlation obtained by the traction matrix of area u, π t (u,k1) represents the degree of influence between the uth region and the k1th region at time t. represents the result of attention processing of the u-th region at time t, π t(n,m) represents the degree of influence between the nth region and the mth region at time t, n = {1, 2, ...., r × c / S}, m = {1, 2, ...., r × c / S}, n ≠ m.
[0124] S225, based on a multi-channel convolutional network, fuses multiple convolution kernels to extract directional local features.
[0125] This embodiment utilizes a deep convolutional network to effectively process two-dimensional information and obtain one-dimensional feature information, thereby achieving a complete analysis of the flexible arm's three-dimensional motion posture information. The choice of convolution kernel in the deep convolutional network plays a crucial role in the information processing effect and the usability of the one-dimensional feature information.
[0126] The commonly used convolution kernel is generally a square matrix. Although it can extract global information in the two-dimensional space, it will ignore some edge position information. Since the position of the flexible manipulator in the entire task space is random, the square matrix convolution kernel alone cannot effectively complete the parsing task; while the convolution kernel of the row vector and column vector can characterize the horizontal and vertical directions, extract local features with stronger directionality, and effectively describe the motion characteristics of the flexible arm.
[0127] This embodiment fuses the square matrix convolution kernel with the row and column vector convolution kernel to perform a multi-perspective analysis of the results of the attention mechanism processing, thereby enhancing the comprehensiveness of feature expression. This embodiment uses the following logical representation to fuse the square matrix convolution kernel with the row and column vector convolution kernel:
[0128]
[0129] in, Represents the result after square matrix convolution kernel processing, Represents the result after column convolution kernel processing, Represents the result after row convolution kernel processing, In step S224, the result after the dimensionality reduction processing is processed by the attention concentration processing. Indicates that each area is viewed separately, and the formula (13) in S225 Indicates viewing all areas together; Indicates k s ×k s is the size of the convolution kernel, and the depth convolution is performed. The number of convolution channels is g. Indicates 1×k c is the size of the convolution kernel, and the depth convolution is performed. The number of convolution channels is g. Indicates k r ×1 is the size of the convolution kernel, depthwise convolution is performed, and the number of convolution channels is g.
[0130] S226, fuses the results of the multi-channel convolutional network, uses regularization and multi-layer perceptron to describe one-dimensional feature information, and outputs the parsing results of the adaptive spatial motion posture parsing network.
[0131] The results of the three deep convolutional networks in S225 are fused, and then regularization and multi-layer perceptrons are used to complete the depiction of one-dimensional feature information. Data regularization is used to improve the convergence and training stability of the deep network, while the multi-layer perceptron is used to achieve re-mapping of the one-dimensional information, and to compress or generalize information in the case of information redundancy and over-processing, using the following logic:
[0132]
[0133]
[0134] in, represents the output of the adaptive spatial motion pose parsing network, represents the fusion result after processing by the deep convolutional network of three different convolution kernels. Flatten(·) means flattening the matrix into a vector. represents the result after regularization processing, Var(·) represents the variance of the calculated vector, and Mean(·) represents the mean of the calculated vector.
[0135] S3, embedding the control signal based on the encoding module to extract the hidden features in the control signal, specifically:
[0136] Through the linear mapping layer, the control input information is expanded and the nonlinear characteristics contained in the control input information are extracted, which is expressed as in, It represents the result of the servo control information being processed by the linear mapping layer. Linear(·) represents the linear mapping layer. Represents the control signal of the servo actuator at different times t.
[0137] Since the above process may cause some information to lose its basic characteristics belonging to the servo, this embodiment compensates for this defect by encoding the data. In addition, since the expansion of information may produce more redundant information, affecting the reliability of the control input information, a mask mechanism is added during the data encoding process to dynamically shield possible redundant information, thereby improving the reliability of the network training results. The following logic is used to represent the data encoding and data masking of the control signal:
[0138]
[0139] in, Indicates the characteristic encoding of the control information after linear processing, is the result after feature encoding, Indicates the actual maximum control output of the servo actuator;
[0140] and The result after feature encoding is masked; where M represents a randomly initialized mask unit, M N,T Indicates the mask unit corresponding to the Nth control information unit at time T, M mask represents the mask matrix, is the result after mask processing.
[0141] In this embodiment, the control signal of the servo motor is used as the control input. By embedding the control input into the code, the information in the control information, including hidden features, is fully extracted. Combined with the spatial posture analysis results, the information is used as the input and output of the multi-task modeling network respectively. The gradient update direction during the training process of the multi-task modeling network is regulated to achieve adaptive balance of the multi-task modeling network parameters.
[0142] S4, constructs a new dataset based on the processed spatial motion posture information and control signals, and builds a multi-task modeling network with an adaptive gradient balancing mechanism.
[0143] In order to complete the effective modeling of the flexible arm's spatial motion posture, on the basis of completing the analysis of the flexible arm's three-dimensional spatial posture, it is necessary to further explore the mapping relationship between it and the control input of the flexible arm. The data after the analysis is completed still has a high dimension and still has an association relationship. The one-to-many neural network model directly used for mapping modeling cannot learn the association relationship between the data very well, thereby reducing the accuracy of modeling and affecting the control effect. The present invention adopts a multi-task modeling network, which can take into account the association relationship between different data, but there will be an imbalance in network parameters. A new data set is constructed based on the spatial motion posture information and control signals processed in steps S2 and S3, which can regulate the gradient update direction in the multi-task modeling network training process and realize the adaptive balance of the multi-task modeling network parameters. Therefore, the present invention selects a multi-task modeling network as the flexible arm spatial motion posture modeling neural network model. Specifically, the S4 includes the following steps:
[0144] S41, combining the multi-expert network, weight control network and output mapping network to build a multi-task modeling network, specifically: using the multi-expert network to share information for all control information θ Embed Perform unified feature processing, generate expert weights based on the weight control network, perform final mapping processing on each output through the output mapping network, and use the following logic to represent all control information θ EmbedPerform unified feature processing:
[0145]
[0146] in, Indicates that the expert network outputs the i3th output variable, K indicates the number of expert networks, represents the k1th expert network, k1=1,2,…,K, N′ represents the number of gates, represents a gating network.
[0147] The following logic is used to express the final mapping process for each output through the output mapping network:
[0148]
[0149] in, represents the spatial pose estimated by the neural network, Represents the output mapping network, that is, the artificial neural network.
[0150] In this embodiment, all control information is uniformly processed using information sharing among the K expert networks. Expert weights are generated based on the weight control network. The output mapping network performs final mapping processing on each output. For the i3th output variable, the i3th output variable is estimated using formula (18). The output mapping network then performs final mapping processing on each output, reflecting the differences between different output results. The final estimated value of the i3th output variable is obtained using formula (19).
[0151] S42, defines a gradient paradigm for parameter updates of shared parameters based on multi-task modeling losses with dynamic weights that fluctuate over time.
[0152] The main challenge facing multivariate prediction models is the risk of training imbalance. Specifically, in a multivariate prediction architecture, training shared parameters is not always balanced across different tasks, and the optimizer may overly focus on the dominant task at the expense of other tasks. To address this, this embodiment designs an adaptive gradient balancing module that dynamically adjusts task weights to balance gradients. The multi-task modeling loss L is represented by the following logic:
[0153]
[0154] in, represents the prediction loss of the i3th output variable at different times t, Represents the weight of the corresponding output variable.
[0155] However, the fixed weight It is difficult to determine and requires special parameter adjustment, especially when the predicted output shows non-stationary fluctuations, that is, when the flexible manipulator has a large position change per unit time during the movement. Therefore, this embodiment introduces a dynamic weight that fluctuates over time. The following logic is used:
[0156]
[0157] Among them, L(t) represents the multi-task modeling loss that introduces dynamic weights that fluctuate over time, Represents the weight of the output variable that changes over time.
[0158] According to the multi-task modeling loss and dynamic weights, the gradient paradigm of the parameter update of the shared parameter W can be expressed using the following logic:
[0159]
[0160] in, represents the parameter gradient in the loss function, Express the gradient of the parameters in the loss function, ‖·‖ 2 represents the Euclidean norm.
[0161] S43, initialize the target gradient, use cosine similarity to calculate whether the target gradient and the original gradient between any tasks conflict, where a negative value indicates a conflicting gradient; if a conflict occurs, the target gradient will be updated using vector projection according to the conflict relationship between the conflicting gradients; if no conflict occurs, then
[0162] In this embodiment, the target gradient is initialized The update of the target gradient is expressed using the following logic:
[0163]
[0164] in, represents the i3th network parameter update gradient, represents the j3th network parameter gradient.
[0165] The multi-task learning network contains N′ networks, and the parameters of these N′ networks are updated independently. This will cause the parameters of some neural networks to move in a direction that is not conducive to network learning. To prevent this, this embodiment calculates the direction of the current network gradient and the gradient of other networks. If the directions are opposite, the current gradient is updated using cosine similarity and target gradient. If the directions are the same, the direction of the current network gradient remains unchanged.
[0166] In this example, a multi-task modeling network with an adaptive gradient balancing mechanism was used to effectively model the spatial motion of a rope-driven flexible arm using the pose information obtained from an adaptive spatial motion pose parsing network and embedded coded servo control signals as the dataset. 60% of the new dataset was used as the training set, 20% as the validation set, and 20% as the test set. The model was built using the Torch framework running Python 3.8. The computation was performed using a 2x Intel(R) Xeon(R) Gold 6133 CPU @ 2.50GHz and a 2x NVIDIA GeForce RTX 4090 GPU.
[0167] Mean absolute error is the measurement indicator, where m is the total amount of data used for training or testing. The training results of the multi-task modeling network are as follows Figure 4 As shown in the results, the mean absolute error decreases with increasing training times, and the multi-task modeling network achieves good results in both the training and validation sets. Two typical spatial poses from the test set, the "U" pose and the "S" pose, are selected to analyze the proposed modeling method.
[0168] The "U"-shaped posture requires the upper half of the flexible arm to remain unchanged, while the lower half is bent upward; the "S"-shaped posture requires the upper and lower halves of the flexible arm to bend in opposite directions. Compared with the refined state, these two postures have great movement changes and have extremely high requirements for the modeling method. The modeling method proposed in this invention has the following modeling effects on the "U"-shaped posture and the "S"-shaped posture methods: Figures 5 and 6 As shown in FIG, it can be seen from the modeling results that the modeling method proposed in the present invention can achieve accurate modeling of the above two postures.
[0169] like Figure 7 As shown, this embodiment selects seven typical neural network modeling methods, including artificial neural network (ANN), recurrent neural network (RNN), long short-term memory neural network (LSTM), Transformer, Mamba, liquid neural network (LNN), and multi-gate mixture-of-experts (MMoE), and compares them with the modeling method provided by the present invention ( Figure 7Ours is shown in the figure for comparison. ANN, RNN, and LSTM are relatively basic neural network models. Transformer and Mamba are neural network models that further incorporate attention and selection mechanisms, enabling focused selective learning of information. LNN is a newer neural network model that can dynamically adjust the connections between neurons, offering greater advantages for solving complex nonlinear problems. MMoE is a commonly used multi-task learning network used to model and process coupled information.
[0170] In this embodiment, R square Mean absolute error Root mean square error and mean absolute percentage error The four modeling evaluation indicators are compared. From the comparison results, it can be seen that the modeling method provided by the present invention is superior to other modeling methods in every aspect, which further illustrates the superiority of the modeling method provided by the present invention.
[0171] The present invention also provides a device comprising a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above-mentioned method for modeling the spatial motion posture of the rope-driven flexible arm, and the processor is configured to execute the program stored in the memory.
[0172] The present invention also provides a storage medium storing a computer program. When the computer program is run by a processor, the steps of the above-mentioned method for modeling the spatial motion posture of the rope-driven flexible arm are executed.
[0173] The proposed method for modeling the spatial motion posture of a rope-driven flexible arm is a purely data-based modeling approach. Its adaptive spatial motion posture analysis network addresses issues such as high-dimensional nonlinearity and behavioral coupling in the flexible arm's spatial motion posture. Control information embedding and encoding expands and extracts hidden information from a single servo. An adaptive parameter balancing mechanism addresses the negative transfer issue in multi-task learning networks, increasing the network's training capabilities. Requiring only data collection and no hypothesis analysis, the proposed method is applicable to any system with the same problem and exhibits strong generalization capabilities.
[0174] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for modeling the spatial motion posture of a rope-driven flexible arm, characterized in that: The following steps are involved: S1, collects the spatial motion posture information of the flexible arm and the control signal of the servo steering gear to construct the original data set; S2, parsing the spatial motion posture information based on an adaptive spatial motion posture parsing network; the adaptive spatial motion posture parsing network includes a hybrid multi-dimensional scaling module and a spatiotemporal adaptive attention mechanism module; S3, embedding the control signal based on the encoding module to extract the hidden features in the control signal; S4, constructs a new dataset based on the processed spatial motion posture information and control signals, and builds a multi-task modeling network with an adaptive gradient balancing mechanism.
2. The method for modeling the spatial motion posture of a rope-driven flexible arm according to claim 1, characterized in that: Said S1 comprises: S11, collect the control signals of the servo motor at different times t in, represents the control signal of the i1th servo actuator at time t, where N is the total number of servo actuators; S12, collect the spatial motion posture information of the flexible arm through the motion capture system, and record the spatial motion posture information of the flexible arm at different times t in, represents the position information of the kth point on the i2th metal disk at time t, and P represents the total number of metal disks.
3. The method for modeling the spatial motion posture of a rope-driven flexible arm according to claim 1, characterized in that: The S2 includes: S21, decoupling and reducing the dimensionality of the spatial motion posture information using a hybrid multidimensional scaling transform; the hybrid multidimensional scaling transform module includes a metric multidimensional scaling transform and a non-metric multidimensional scaling transform; S22, uses the spatiotemporal adaptive attention mechanism module to capture the spatial motion characteristics of spatial motion posture information.
4. The method for modeling the spatial motion posture of a rope-driven flexible arm according to claim 3, characterized in that: The S21 includes: S211, mapping the high-dimensional data to a low-dimensional space based on the metric multidimensional scaling transformation, and performing dimensionality reduction mapping on the relative position information in the original position, specifically: The location information of the sample points after dimensionality reduction is recorded as The metric multidimensional scaling transformation is transformed into an optimization problem to solve the following formula: Among them, D position Represents the Euclidean distance between the original points, D M-MDS Represents the Euclidean distance between points after dimensionality reduction; S212, based on non-metric multidimensional scaling transformation, minimizes the difference between the mapped point order and the original point order, and approaches the low-dimensional representation of the original structure, specifically: Note the different points between the flexible arms The distance between different points is expressed as M1 represents the total number of points used to collect the flexible arm posture information; Identify a set of dissimilarities The non-metric multidimensional scaling transformation is transformed into a minimization problem to solve the following formula: in, represents the distance between the i-th and j-th dissimilar points in the dimensionality reduction space, Represents the distance between the kth dissimilar point and the lth dissimilar point after dimensionality reduction, represents the distance between the i-th and n-th dissimilar points in the dimensionality reduction space, represents the distance between the jth and nth dissimilar points in the dimensionality reduction space, represents the distance between the kth and nth dissimilar points in the dimensionality reduction space, f(·) represents the loss function, k={1,2,....,M1}, l={1,2,....,M1}, k≠l≠j, represents the position of the i-th different point at time t, represents the position of the jth different point at time t, represents the position of the kth different point at time t; S213, fusing the metric multidimensional scaling transformation result and the non-metric multidimensional scaling transformation result, using the following logic representation: in, It represents the spatial motion posture information of the flexible arm at different times t after dimensionality reduction mapping, and ⊕ represents the exclusive-OR operation.
5. The method for modeling the spatial motion posture of a rope-driven flexible arm according to claim 4, characterized in that: The S22 includes: S221, the information after dimensionality reduction mapping Divide into S×S non-overlapping areas, and use the adaptive adjustment factor ε to perform differential operations on different sliding window areas to obtain S222, obtain the query matrix corresponding to the u-th region at time t through linear mapping Bond Matrix and the value matrix S223, considering the influence relationship between different regions, find the influence degree corresponding to different key values through the traction matrix, and find the region that should participate in each given region; specifically, calculate the query vector and key vector of each region, use cosine similarity to calculate the correlation traction matrix, and obtain the traction matrix S223, using the traction matrix between regions Calculate the degree of traction between different areas; specifically, With the traction matrix Combine to get in, represents the correlation bond matrix, represents the correlation value matrix, represents the Ronecker product; S224, calculate the attention coefficient of each area through the traction matrix Remove the influence of redundant and invalid information on the area of attention and improve the accuracy of attention; S225, based on a multi-channel convolutional network, fuses multiple convolution kernels to extract directional local features. Specifically, the square matrix convolution kernel is fused with the row and column vector convolution kernels, using the following logic representation: in, Represents the result after square matrix convolution kernel processing, Represents the result after column convolution kernel processing, Represents the result after row convolution kernel processing, It represents the result after the dimensionality reduction processing and the attention concentration processing, Indicates k s ×k s is the size of the convolution kernel, and the depth convolution is performed. The number of convolution channels is g. Indicates 1×k c is the size of the convolution kernel, and the depth convolution is performed. The number of convolution channels is g. Indicates k r ×1 is the size of the convolution kernel, depthwise convolution is performed, and the number of convolution channels is g; S226, fuses the results of the multi-channel convolutional network, uses regularization and multi-layer perceptron to describe one-dimensional feature information, and outputs the parsing results of the adaptive spatial motion posture parsing network.
6. The method for modeling the spatial motion posture of a rope-driven flexible arm according to claim 1, characterized in that: The S3 is specifically: Through the linear mapping layer, the control input information is expanded and the nonlinear characteristics contained in the control input information are extracted, which is expressed as The data is encoded and a mask mechanism is added to dynamically shield the redundant information that exists when the information is expanded. The following logic is used to represent it: in, is the result after feature encoding, is the result after mask processing, Indicates the actual maximum control output of the servo motor, M represents the randomly initialized mask unit, M N,T Indicates the mask unit corresponding to the Nth control information unit at time T, M mask Represents the mask matrix.
7. The method for modeling the spatial motion posture of a rope-driven flexible arm according to claim 1, characterized in that: The S4 includes: S41, combines the multi-expert network, weight regulation network and output mapping network to build a multi-task modeling network; S42, a gradient paradigm for parameter updates of shared parameters defined based on multi-task modeling losses and dynamic weights that fluctuate over time; S43, initialize the target gradient, use cosine similarity to calculate whether the target gradient and the original gradient between any tasks conflict, if a conflict occurs, the target gradient will be updated by vector projection according to the conflict relationship between the conflicting gradients, and the following logic is used to represent the update of the target gradient: in, represents the i3th network parameter update gradient, represents the j3th network parameter gradient, Indicates the initialization target gradient, take 8. The method for modeling the spatial motion posture of a rope-driven flexible arm according to claim 1, characterized in that: 60% of all the data in the new dataset was used as a training set, 20% as a validation set, and 20% as a test set. A multi-task modeling network was constructed based on the torch framework of Python 3.
8.
9. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the spatial motion posture modeling method of the rope-driven flexible arm according to any one of claims 1 to 8, and the processor is configured to execute the program stored in the memory.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for modeling the spatial motion posture of a rope-driven flexible arm according to any one of claims 1 to 8 are executed.
Citation Information
Patent Citations
Control method of rope-driven flexible arm
CN111421529A
Method for controlling tail end position of rope-driven flexible arm
CN118493364A