Twin modeling method, device and equipment for modular object production line
By using a twin modeling method for modular object production lines and training virtual contours with a deep Q-network reinforcement learning model, the problem of the narrow application scope of modular object digital twin systems is solved, and efficient modular object testing and production are achieved.
Patent Information
- Application Number
- CN202510055778.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-01-14
AI Technical Summary
In existing technologies, digital twin systems for modular objects have a narrow range of applications and are difficult to support structural innovation and flexible modular product design.
By adopting a twin modeling method for modular object production lines, virtual contours are trained using a model-free reinforcement learning model of deep Q-networks by acquiring the component and environmental parameters, and combined with dynamics and kinematics simulations, high-precision virtual reproduction of modular objects is achieved.
It supports the creation of custom modular objects for models and modification of entities, which improves testing efficiency, reduces costs, and achieves high-precision simulation of the real world.
Smart Images

Figure CN119882644B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer simulation, and in particular to a twin modeling method, device and equipment for a modular object production line. BACKGROUND
[0002] At present, a large number of physical manufacturing industries tend to adopt the design and manufacturing process of modular products based on the consideration of product update iteration, cost control, corresponding production line modification, etc. On this basis, in order to further facilitate design, modification and even testing, the corresponding digital twin system, i.e. the digitalization of physical modular products and the reproduction of them in a virtual world / system at a scale of 1:1, is gradually popularized and applied. At present, the digital twin system of modular objects (such as robots, vehicle chassis, etc.) usually collects data from a certain ready-made object or fixed production line and simulates the reproduction. As can be seen, the application of digital twin is generally limited to fixed objects or fixed fields, and the application range is relatively narrow, which is not conducive to structural innovation. SUMMARY
[0003] In order to overcome the shortcomings of the prior art, according to a first aspect of the present application, a twin modeling method for a modular object production line is provided, which comprises the following steps:
[0004] - obtaining characteristic parameters of each component member of the modular object as a first input, and obtaining environmental parameters of a target working environment of the modular object as a second input;
[0005] - creating an initial virtual contour of the modular object based on at least part of the first input;
[0006] - importing the first input, the second input and the initial virtual contour into a simulation system;
[0007] - training and simultaneously automatically adjusting the first input, the second input and the initial virtual contour in the simulation system by using a model-free reinforcement learning model;
[0008] - calculating at least one of a first output associated with the first input, a second output associated with the second input and a finished product virtual contour associated with the initial virtual contour.
[0009] In one embodiment, the model-free reinforcement learning model is in the form of a deep Q network, wherein the Q function in the deep Q network is:
[0010]
[0011] wherein a represents a learning rate, γ represents a discount factor of a Markov reward process, s and S ′ represent the current and next environment states respectively, and a and a′ respectively, r represents the reward obtained by the current state.
[0012] In another embodiment, the Q-function is updated by:
[0013]
[0014] where ω represents the loss function, and N represents the total number of states / executable operations.
[0015] In yet another embodiment, the twin modeling method of the modular object production line further comprises:
[0016] - selecting any executable operation as an exploration operation with a probability ∈;
[0017] - selecting an executable operation with the largest expected reward estimate value in past experience as an exploitation operation with a probability 1-∈;
[0018] - balancing the exploration operation and the exploitation operation to avoid local optimal operation a t :
[0019]
[0020] wherein, represents the set of all operations.
[0021] In yet another embodiment, the probability ∈ is set to decay over time:
[0022]
[0023] wherein t represents time, and t>0.
[0024] In yet another embodiment, the first input comprises at least one of a geometric property parameter, a material property parameter, a mass property parameter, a connection mode, a constraint condition, an actuator type, an output property parameter, and a dynamic property parameter; and the second input comprises at least one or more of a gravity environment parameter, a friction environment parameter, a collision environment parameter, and a vibration environment parameter.
[0025] In yet another embodiment, the simulation system has a dynamics calculation and a kinematics calculation, and the dynamics calculation and the kinematics calculation are based on an ODE physics engine.
[0026] According to a second aspect of the present application, there is provided a twin modeling device of a modular object production line, comprising:
[0027] The input module is configured to obtain characteristic parameters of each component of the modular object as a first input, and obtain environmental parameters of a target working environment of the modular object as a second input.
[0028] The creation module is configured to create an initial virtual profile of the modular object based on at least part of the first input.
[0029] The import module is configured to import the first input, the second input, and the initial virtual profile into the simulation system.
[0030] The training module is configured to train the first input, the second input, and the initial virtual profile in the simulation system by using a model-free reinforcement learning model and simultaneously automatically adjusting them.
[0031] The calculation module is configured to calculate at least one of a first output associated with the first input, a second output associated with the second input, and a finished virtual profile associated with the initial virtual profile.
[0032] According to a third aspect of the present application, an electronic device is provided, which comprises:
[0033] a memory storing non-transitory computer-readable instructions;
[0034] at least one processor configured to, when executing the non-transitory computer-readable instructions, perform at least one step of the twin modeling method of the modular object production line according to the first aspect of the present application.
[0035] The present application supports self-creation of a model of a modular object, and simultaneously supports modification of an existing, physical modular object, and realizes a simulation process and result sufficient to resemble the real world. With the present application, manufacturers in various industrial fields can use virtual components to complete structural construction and high-precision testing of a modular object in advance, and then use hardware components to complete batch production of a physical modular object, greatly improving the testing efficiency of each version of the scheme and reducing the cost of testing. BRIEF DESCRIPTION OF DRAWINGS
[0036] The specific embodiments of the present application will be further described in detail below with reference to the accompanying drawings, in which:
[0037] Figure 1 A flowchart of the twin modeling method of the modular object production line according to an embodiment of the present application is shown.
[0038] Figure 2 A flowchart of solving by using a deep Q network according to an embodiment of the present application is shown.
[0039] Figure 3A flow chart showing the twin modeling of a modular vehicle chassis according to another embodiment of the present application is shown.
[0040] Figure 4 A schematic diagram showing a twin modeling device of a modular object production line according to yet another embodiment of the present application is shown. DETAILED DESCRIPTION
[0041] The present application will be described in detail below with reference to the accompanying drawings and specific embodiments. The illustrative embodiments and descriptions thereof are used to explain the present application, but are not intended to limit the present application.
[0042] Embodiment One
[0043] In this embodiment, the modular object is a modular robot. The hardware components and virtual components of the modular robot are respectively composed of various robot modules, which include structural members, actuators, control devices, sensors, etc. These modules use a common mating interface and can be quickly assembled into various common types of robots, such as robotic arms, mobile robots, multi-legged robots, etc., thereby constructing a robot working environment.
[0044] The hardware components include various types of physical hardware that are applied in real scenarios, and the virtual components are composed of various digital models that are reproduced in virtual simulation systems. The hardware components and virtual components have the same physical and electrical properties and together form a digital twin. Users can expand the hardware / software component system with the help of a reinforcement learning system. The reinforcement learning system will collect characteristic parameters of the new hardware components during operation, analyze the parameters, and correct the parameter calibration of the digital model corresponding to the new components, thereby establishing a new digital twin.
[0045] The hardware components include various common robot modules and other auxiliary modules, which can be freely combined to the greatest extent. At the level of module types, the robot modules and other auxiliary modules can include structural modules, actuator modules, control modules, and sensor modules. The virtual components are digital models of the hardware components, and the module types of the virtual components are the same as those of the hardware components, which are composed of three-dimensional models and simulation parameters. The three-dimensional models are modeled by the modules in the hardware components through 1:1 modeling, defining the appearance and collision boundary of the modules, and the simulation parameters determine the physical and electrical properties of the modules in the simulation system, including mass, moment of inertia, friction coefficient, collision coefficient, damping, stiffness, etc.
[0046] The virtual simulation system is mainly used for simulating the virtual robot built by the user through the virtual components and performing motion simulation. The dynamics and kinematics calculation of the simulation system is based on the open source physics engine (ODE). According to the parameters of each module in the virtual component, the physical processes such as force, collision and constraint are calculated with high precision and in real time. At the same time, the ODE can also simulate the motor output power and various sensor data acquisition processes. The simulation system can adopt the BS (Browser / Server Architecture) architecture, in which the server side is used for simulation calculation, and the browser side is used for presenting the graphical picture of simulation and real-time presentation of simulation data. The user can write a control program in the virtual simulation system to control the robot in the simulation system, and the program written can also be copied to the hardware module for direct running.
[0047] The reinforcement learning system is a model-free reinforcement learning model, which collects the running characteristics of the hardware model and inputs them into the model-free reinforcement learning model, and based on the intelligent adjustment parameter function of the deep reinforcement learning model, the hardware components and virtual components are constructed into digital twins, so that the virtual components have the same physical and electrical properties as the hardware components, that is, under the same system input, the system output is always consistent.
[0048] The process of constructing virtual components by using the reinforcement learning system is as follows:
[0049] The user completes the 1:1 three-dimensional model building of the virtual component by measuring and disassembling the hardware components. At this time, the physical and electrical parameters of the virtual component are to be determined.
[0050] The sensor components are used to collect various data during the operation of the robot and send them to the control system or storage system, including collection frequency, measurement accuracy, noise, data structure, communication form, etc. At the same time, the characteristic parameters of the actuator are read, including the motor and associated transmission devices, especially their power, speed, force / torque, time constant, etc.
[0051] The first input includes: geometric characteristic parameters such as the length, width, height and shape of the structure module; material characteristic parameters such as the purpose, elastic modulus and Poisson's ratio of the structure module; mass characteristic parameters such as the total mass, center of mass and inertia tensor of the structure module; connection mode and constraint condition refers to the connection mode and corresponding constraint condition between the structure modules; actuator type includes rotary motor or linear motor, etc.; output characteristic parameters such as the output force, torque, speed and frequency of the actuator; and dynamic characteristic parameters such as the response time or bandwidth of the actuator, etc.
[0052] The second input includes: gravity environment parameters, i.e. the magnitude and direction of the gravitational acceleration; friction environment parameters, i.e. the friction coefficient between the contact surfaces of each module; collision environment parameters, i.e. the elastic and inelastic characteristics of the collision; vibration environment parameters, i.e. the natural vibration frequency and damping characteristics of the system; and natural environment parameters, i.e. the airflow, humidity, air temperature and other natural environment parameters of the environment;
[0053] - using a Deep Q-Network (DQN) model-free reinforcement learning model to train the above parameters and complete the construction of the digital twin model;
[0054] Among them, the state s is used to describe the current situation of the environment; the operation a represents the operation that the robot can execute; the reward r represents the value of the environment feedback after executing the action, which is a reward function, used to guide learning; and the policy p represents the rule for determining that the robot takes a specific operation a in a certain state s; Q represents the expected cumulative reward obtained after taking operation a in state s. On this basis, in order to learn the optimal policy by continuously updating the Q value, the Q function of the reinforcement learning based on the time difference algorithm is:
[0055]
[0056] Among them, α represents the learning rate, γ represents the discount factor of the Markov reward process, s and S ′ respectively represent the current and next environment states, a and a ′ respectively represent the current and next executable operations, and r represents the reward obtained in the current state. Preferably, the value of γ is in the range [0, 1), and particularly preferably, its value is 0.95; the value of the learning rate α is 0.001;
[0057] Then, a deep neural network is used to fit the value of the above Q function. Common neural network structures include multi-layer perceptron (MLP) and convolutional neural network (CNN), and the parameters of these neural networks are usually optimized through training. Moreover, the output size of these neural networks is comparable to the dimension of the operation space.
[0058] DQN can update the model by minimizing the mean square error loss of the Q function:
[0059]
[0060] Among them, ω represents the loss function, and N represents the total number of states / execution operations.
[0061] In addition, DQN introduces an experience replay mechanism, which stores past experiences in a buffer, for example, the state transition (s, a, r, S ′) into a specific memory bank and randomly sample during training, thus improving the utilization of data, the stability of training, and reducing the correlation between samples. Specifically, an experience replay mechanism is introduced, which can satisfy the independence assumption on one hand. The data obtained by interactive sampling in Markov Decision Process (MDP) does not satisfy the independence assumption because the state s at the current time and the state s ′ The non-independent and identically distributed data has a great influence on the training of neural networks, which can make the neural network fit to the nearest training data. Experience replay breaks / reduces the correlation between samples to satisfy the independence assumption. On the other hand, each sample can be reused, thus suitable for gradient learning of deep neural networks.
[0062] In order to further stabilize the training process, DQN uses two neural networks with the same structure but different parameters: one is used to predict the Q value, called the main network, and the other is used to calculate the target Q value, called the target network. The parameters of the target network are updated regularly, which helps to reduce the instability in the training process.
[0063] In the training process, in order to save training time as much as possible while obtaining the optimal solution of the training data, a greedy algorithm is introduced. Therefore, the twin modeling method of the modular robot production line of the embodiment further comprises:
[0064] - selecting any executable operation as an exploration operation with a probability ∈;
[0065] - selecting an executable operation with the largest expected reward estimate value in the past experience as an exploitation operation with a probability of 1-∈;
[0066] - balancing the exploration operation and the exploitation operation to avoid local optimal operation a t :
[0067]
[0068] wherein, represents the set of all operations.
[0069] In some embodiments, the probability ∈ is set to decay over time t:
[0070] ∈_t=1 / t
[0071] wherein t represents time and t>0.
[0072] Referring to Figures 1-2 , the twin modeling method of the modular object (in particular, robot) production line provided by the embodiment comprises:
[0073] - obtaining characteristic parameters of each component of the modular robot as the first input, and obtaining environmental parameters of the target working environment of the modular robot as the second input;
[0074] - creating an initial virtual profile of the modular robot based on at least part of the first input, in particular, geometric characteristic parameters, material characteristic parameters, mass characteristic parameters, connection modes and constraint conditions, etc.;
[0075] - importing the first input, the second input and the initial virtual profile into a simulation system, such as a BS architecture based on ODE;
[0076] - training and automatically adjusting the first input, the second input and the initial virtual profile in the simulation system by using a DQN-free model reinforcement learning model;
[0077] - calculating at least one of the first output associated with the first input, the second output associated with the second input and the finished virtual profile associated with the initial virtual profile.
[0078] It can be imagined that the first output, the second output and the finished virtual profile are the desired training data. These obtained training data can be used on the physical modular product in turn.
[0079] Embodiment two
[0080] In this embodiment, the modular object is a modular vehicle chassis. On this basis, the first input that can be obtained on the physical hardware includes the shape, weight and moment of inertia of each component of the modular vehicle chassis, the output speed and torque of the actuator, and the second input that can be obtained includes the gravity environmental parameters. These obtainable parameters are deterministic parameters. In contrast, the training data expected to be obtained by the twin modeling method of this embodiment include the damping and dynamic and static friction coefficients of the shaft, the dynamic and static friction coefficients and the collision coefficients of the wheel and the ground. These expected parameters are uncertain parameters.
[0081] With reference to Figure 3 , the twin modeling method of the modular object (in particular, the vehicle chassis) production line provided by the embodiment comprises:
[0082] - installing an encoder for the physical vehicle chassis, controlling the speed and output torque of the motor, and collecting various data of the vehicle chassis during operation, such as including the speed, acceleration, position of the chassis, the speed and torque of the motor;
[0083] - using a reinforcement learning system to solve the relationship between the state s and the action a of the digital / virtual twin model in the simulation system. This relationship represents the optimal control strategy of the digital / virtual twin model, so that the digital / virtual twin model is consistent with the motion performance of the entity vehicle chassis. The specific solving process can be described as:
[0084] S1, add and sample data in the experience replay pool; the data part needs to sample valuable data, for example, set a plurality of different (first / second) input values under normal working conditions of the vehicle chassis, and obtain different states s, such as data under different conditions such as straight running, backing up, turning, etc.
[0085] S2, define the number of layers N of the Q network, set according to the number of unknown parameters Num, the number of layers N = lnNum, set 32 neurons in each layer, and use the ReLU activation function, that is, f(x) = max(0, x), so that when the input value x is greater than 0, the ReLU function remains x; when the input value x is less than or equal to 0, the ReLU activation function outputs 0;
[0086] S3, initialize parameters, Q(s, a), learning rate α = 0.001, discount factor γ = 0.95; the final mean square error threshold is [0, 1]; the number of data groups buffer_size = 1000, the number of data groups batch_size is 64, and the probability of the greedy algorithm ∈ = 1 / t; the target network is updated every 10 times;
[0087] S4, select operation: select operation according to the current state s using the greedy algorithm strategy;
[0088] S5, execute the action in the simulation environment and observe the reward: take the operation a, interact with the environment, and observe the next state
[0089] S ′ and immediate reward r;
[0090] S6, store experience: store (s, a, r, S ′ ) in the experience replay pool;
[0091] S7, randomly sample a batch of data from the experience replay pool;
[0092] S8, calculate the target Q value: calculate the target Q value using the target network, that is,
[0093] S9, update the main network: update the model parameters according to the loss function L;
[0094] S10, update the target network: update the parameters of the target network regularly, and the update period is set to every 10 steps here;
[0095] S11, repeating steps S2-S8 until a termination condition is reached, i.e. a mean square error threshold is [0, 1].
[0096] Embodiment Three
[0097] With reference to Figure 4 which shows a schematic diagram of a twin modeling device of a modular object production line according to the present application. In the present embodiment, a twin modeling device of a modular object production line is provided, which comprises:
[0098] an input module 100 configured to obtain characteristic parameters of each component member of the modular object as a first input, and to obtain environmental parameters of a target working environment of the modular object as a second input;
[0099] a creation module 200 configured to create an initial virtual contour of the modular object based on at least a part of the first input;
[0100] an import module 300 configured to import the first input, the second input, and the initial virtual contour into a simulation system;
[0101] a training module 400 configured to train and simultaneously automatically adjust the first input, the second input, and the initial virtual contour in the simulation system by using a model-free reinforcement learning model;
[0102] a calculation module 500 configured to calculate at least one of a first output associated with the first input, a second output associated with the second input, and a finished product virtual contour associated with the initial virtual contour.
[0103] Embodiment Four
[0104] The present application also provides an electronic device comprising a memory storing non-transitory computer readable instructions; and at least one processor configured to perform at least one step of the twin modeling method of a modular object production line according to the first aspect of the present application when executing the non-transitory computer readable instructions.
[0105] The above describes the technical solutions provided by the embodiments of the present application in detail. The principles and implementation manners of the embodiments of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the principles of the embodiments of the present application; meanwhile, for those skilled in the art, the specific implementation manners and application ranges of the embodiments of the present application will be changed. In view of the above, the content of the present description should not be understood as a limitation of the present application.
Claims
1. A twin modeling method of a modular object production line, characterized by, comprises the following steps: - obtaining characteristic parameters of individual components of a modular object as a first input and obtaining environmental parameters of a target working environment of the modular object as a second input; - creating an initial virtual profile of the modular object based on at least a portion of the first input; - importing the first input, the second input, and the initial virtual profile into a simulation system; - training and simultaneously automatically adjusting the first input, the second input, and the initial virtual profile in the simulation system using a model-free reinforcement learning model; - calculating at least one of a first output associated with the first input, a second output associated with the second input, and a finished virtual profile associated with the initial virtual profile; wherein the model-free reinforcement learning model is in the form of a deep Q network, wherein a Q function in the deep Q network is: ; wherein, denotes a learning rate, denotes a discount factor of the Markov reward process, and denote the current and next environment state, respectively, and denote the current and next executable operation, respectively, denotes the reward obtained for the current state; wherein the Q function is updated by: ; wherein ω represents a loss function and N represents a total number of states / execution operations; further comprising: - with a probability selecting any of the executable operations as the exploration operation; - with a probability selecting, as the utility operation, the executable operation having the greatest expected reward estimate value among past experiences; - balancing the exploration operation and the exploitation operation to avoid local optimal operation : ; wherein represents the set of all operations.
2. The twin modeling method of a modular object production line according to claim 1, characterized in that, the probability e is set to decay over time t: ; wherein t represents time and has a value ranging from t .
3. The twin modeling method of a modular object production line according to claim 1, wherein, the first input comprises at least one of geometric characteristic parameters, material characteristic parameters, mass characteristic parameters, connection modes, constraint conditions, actuator types, output characteristic parameters, and dynamic characteristic parameters; and the second input comprises at least one of gravity environmental parameters, friction environmental parameters, collision environmental parameters, and vibration environmental parameters.
4. The twin modeling method of a modular object production line according to claim 3, wherein, the simulation system has dynamics calculation and kinematics calculation based on an ODE physical engine.
5. The twin modeling method of a modular object production line according to claim 4, characterized in that, the simulation system is compatible with a ROS2 operating system.
6. A twin modeling device of a modular object production line, characterized by, comprises: an input module configured to obtain characteristic parameters of individual components of a modular object as a first input and obtain environmental parameters of a target working environment of the modular object as a second input; a creation module configured to create an initial virtual profile of the modular object based on at least a portion of the first input; an import module configured to import the first input, the second input, and the initial virtual profile into a simulation system; a training module configured to train and simultaneously automatically adjust the first input, the second input, and the initial virtual profile in the simulation system using a model-free reinforcement learning model; a calculation module configured to calculate at least one of a first output associated with the first input, a second output associated with the second input, and a finished virtual profile associated with the initial virtual profile; wherein the model-free reinforcement learning model is in the form of a deep Q network, wherein a Q function in the deep Q network is: ; wherein, denotes a learning rate, denotes a discount factor of the Markov reward process, and denote the current and next environment state, respectively, and denote the current and next executable operation, respectively, denotes a reward obtained for the current state; wherein the Q function is updated by: ; wherein ω represents a loss function and N represents a total number of states / execution operations; further comprising: - with a probability selecting any of the executable operations as the exploration operation; - with a probability selecting, as the utility operation, the executable operation having the greatest expected reward estimate value among past experiences; - balancing the exploration operation and the exploitation operation to avoid local optimal operation : ; wherein represents the set of all operations.
7. An electronic device, comprising: comprises: a memory storing non-transitory computer-readable instructions; at least one processor configured to, when executing the non-transitory computer-readable instructions, perform at least one step of the twin modeling method of the modular object production line according to any one of claims 1-6.
Citation Information
Patent Citations
Digital twin object processing system and method
CN118747479A