An FSM simulation method and system based on a virtual factory robot

By using the Q-Learning algorithm in a virtual factory to optimize the robot's decision-making process and dynamically adjust the discount factor, the problem of inefficient decision-making in complex environments is solved, and efficient task completion and intelligent production management are achieved.

CN119781315BActive Publication Date: 2025-07-25HANGZHOU TAOPU TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510028092.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-07-25
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

In the prior art, industrial robots have difficulty choosing the optimal decision in complex and dynamically changing manufacturing plant environments, resulting in inefficient decision making.

Method used

The Q-Learning algorithm is used to optimize the decision-making process of the robot, and the exponential weighted moving average and probability density analysis of the historical reward values of future tasks is performed dynamically to adjust the discount factor, and combine the robot's work performance under different discount factors, and selectively use the updated discount factor.

Benefits of technology

It improves the task completion efficiency of robots in complex and dynamic factory environments, and realizes intelligent production management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119781315B_ABST
    Figure CN119781315B_ABST
Patent Text Reader

Abstract

The present invention discloses an FSM simulation method and system based on a virtual factory robot, specifically relating to the technical field of FSM simulation. It includes using the Q-Learning algorithm in FSM simulation to optimize the decision-making process of the robot, extracting the current task and future tasks of the robot during the decision-making process of the robot, analyzing the historical reward values of different future tasks of the robot, performing exponentially weighted moving average analysis and probability density analysis on the historical reward values of future tasks, determining the uncertainty information of future tasks, obtaining the dependency information of future tasks through comparative analysis, comprehensively analyzing the uncertainty information and dependency information of future tasks, dynamically adjusting the discount factor of the Q-Learning algorithm, and selectively using the updated discount factor. Through FSM simulation, the present invention helps the robot complete tasks more efficiently in a complex and dynamic factory environment, thereby realizing intelligent production management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of FSM simulation, and more specifically, to an FSM simulation method and system based on a virtual factory robot. Background Art

[0002] With the continuous development of automation and artificial intelligence technologies, the decision-making ability of robots in complex environments has become increasingly important. As an adaptive learning technology, reinforcement learning has gradually received attention. Among them, the Q-Learning algorithm has become a popular choice for robot decision-making optimization due to its simple and effective characteristics.

[0003] In an existing virtual simulation training system based on an industrial robot (CN116229792B), by using the methods of data integration analysis and quantitative setting of workstations, the clear setting of the number of industrial robot workstations required by a manufacturing factory is realized, and then the simulation and planning layout scheme of the industrial robot is clarified. However, as the task complexity of the industrial robot increases, it is difficult to adapt to a dynamically changing environment, resulting in the industrial robot not being easy to select the optimal decision.

[0004] To solve the above defects, a technical solution is provided now. Summary of the Invention

[0005] In order to overcome the above defects of the prior art, an embodiment of the present invention provides an FSM simulation method and system based on a virtual factory robot to solve the problems raised in the above background art.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] An FSM simulation method based on a virtual factory robot, characterized in that it specifically includes the following steps:

[0008] S1: Use the Q-Learning algorithm in the FSM simulation to optimize the decision-making process of the robot, and extract the current task and future tasks of the robot during the decision-making process of the robot.

[0009] S2: Analyze the historical reward values of different future tasks of the robot, perform exponentially weighted moving average analysis and probability density analysis on the historical reward values of the future tasks to determine the uncertainty information of the future tasks, and obtain the dependency information of the future tasks through comparative analysis.

[0010] S3: Comprehensively analyze the uncertainty information and dependency information of the future tasks, dynamically adjust the discount factor of the Q-Learning algorithm, and determine the updated discount factor and the basic discount factor.

[0011] S4: Analyze the difference information and efficiency information by combining the working performance of the robot under the updated discount factor and the basic discount factor, and selectively use the updated discount factor.

[0012] In a preferred embodiment, the Q-Learning algorithm is used in the FSM simulation to optimize the decision-making process of the robot, including:

[0013] Use the Q-learning algorithm for the robot in the virtual factory to help the robot autonomously learn task priorities according to real-time environment and task changes, and optimize its state transition decisions, including:

[0014] Determine the state information of the robot in the virtual factory, including the current position, target position, task status, surrounding environment information, and resource status of the robot. The state information will be used as the input of the Q-learning algorithm to guide the robot to decide the next action to be executed;

[0015] Determine the action information of the robot in the virtual factory, including the moving direction, task selection, and path planning of the robot. The action information will be used as the output of the Q-learning algorithm to indicate the actions that the robot should take in the current state.

[0016] In a preferred embodiment, determine the uncertainty information of future tasks, including:

[0017] Represent the uncertain information by the reward fluctuation amplitude coefficient and the probability distribution coercion coefficient;

[0018] The acquisition logic of the reward fluctuation amplitude coefficient is as follows: Extract the future tasks of the robot, collect the historical data of the robot in the future tasks, and use the exponentially weighted moving average to smooth the reward volatility of the future tasks. The expression of the exponentially weighted moving average is: ; where represents the smoothed value at the current moment, represents the smoothing factor, represents the smoothed value at the previous moment, represents the immediate reward value of the future task;

[0019] According to the historical data of the future tasks, obtain the smoothed value of each task of the future tasks, calculate the difference between the smoothed value of each task of the future tasks and the actual reward value, and mark the difference between the smoothed value of each task of the future tasks and the actual reward value as: where is the smoothed value of each task of the future tasks, is the actual reward value of each task of the future tasks, n = 1, 2, 3,..., N, N is a positive integer, and n is the number of the historical data of the future tasks;​

[0020] Calculate the reward fluctuation amplitude coefficient, and the calculation formula is: ; where Reward fluctuation amplitude coefficient.

[0021] In a preferred embodiment, the probability distribution coercion coefficient includes:

[0022] The acquisition logic of the probability distribution coercion coefficient is: According to the historical data of the collection robot in future tasks, use the Gaussian function to represent the probability density function of the future task reward value, and the expression is: ; where Is the standard deviation of the future task reward value, Is the average value of the future task reward value, , ;

[0023] Calculate the probability distribution coercion coefficient, and the calculation formula is: ; where Is the minimum value of the reward value of the future task in the historical data, Is the maximum value of the reward value of the future task in the historical data.

[0024] In a preferred embodiment, the dependency information of the future task includes:

[0025] Express the dependency information of the future task through the priority influence coefficient;

[0026] The acquisition logic of the priority influence coefficient is: According to the historical data of the robot's future tasks, determine the average reward value after the robot selects a future task after completing the current task, and mark the average reward value after the robot selects a future task after completing the current task as: , where i = 1, 2, 3,..., I, I is a positive integer, and i is the number of the future task;

[0027] Obtain the maximum and minimum values of the reward value after the robot selects a future task after completing the current task, and mark the maximum and minimum values of the reward value after the robot selects a future task after completing the current task as: And ; Calculate the priority influence coefficient, and the calculation formula is: ; where Is the priority influence coefficient.

[0028] In a preferred embodiment, dynamically adjust the discount factor of the Q-Learning algorithm, including:

[0029] Through the comprehensive analysis of the uncertainty information and dependency information of tasks by a robot, the reward fluctuation amplitude coefficient, the probability distribution coercion coefficient, and the priority influence coefficient are weighted and calculated to construct an adjustment evaluation model and generate an adjustment evaluation coefficient. The expression is as follows: ; where, is the adjustment evaluation coefficient, , , are the proportionality coefficients of the reward fluctuation amplitude coefficient, the probability distribution coercion coefficient, and the priority influence coefficient, , , are all greater than 0 respectively;

[0030] Set the adjustment evaluation coefficient threshold, and compare the adjustment evaluation coefficient with the adjustment evaluation coefficient threshold. If the adjustment evaluation coefficient is greater than the adjustment evaluation coefficient threshold, then adjust the discount factor with the adjustment evaluation coefficient. If the adjustment evaluation coefficient is less than the adjustment evaluation coefficient threshold, then do not adjust the discount factor with the adjustment evaluation coefficient;

[0031] Set the basic discount factor, adjust the discount factor with the adjustment evaluation coefficient, determine the updated discount factor according to the current working environment of the robot, and mark the updated discount factor as: , where, , where, is the correction coefficient.

[0032] In a preferred embodiment, the difference information includes:

[0033] The difference information is represented by the Q-value difference coefficient;

[0034] The acquisition logic of the Q-value difference coefficient is as follows: For the same robot to conduct multiple comparative experiments on the task queue of the same order, respectively conduct experiments using the Q-Learning algorithm with the updated discount factor and the Q-Learning algorithm with the basic discount factor, respectively obtain the cumulative sum of Q-values after using the Q-Learning algorithm with the updated discount factor and the Q-Learning algorithm with the basic discount factor, and mark the average cumulative sum of Q-values of the Q-Learning algorithm with the updated discount factor as: , and mark the average cumulative sum of Q-values of the Q-Learning algorithm with the basic discount factor as: ;

[0035] Calculate the Q-value difference coefficient, and the calculation formula is: ; where, is the Q-value difference coefficient.

[0036] In a preferred embodiment, the efficiency information includes:

[0037] The efficiency information is represented by a time deviation coefficient;

[0038] The acquisition logic of the time deviation coefficient is as follows: Multiple comparative experiments are carried out on the task queue of the same order by the same robot. Experiments are carried out respectively by using the Q-Learning algorithm with an updated discount factor and the Q-Learning algorithm with a basic discount factor. Compare the time used after completing the same number of tasks. Mark the average time used by the Q-Learning algorithm with the updated discount factor as: Mark the average time used by the Q-Learning algorithm with the basic discount factor as: ;

[0039] Calculate the time deviation coefficient. The calculation formula is: ; where is the time deviation coefficient.

[0040] In a preferred embodiment, selectively using the updated discount factor includes:

[0041] Perform a weighted calculation on the Q-value difference coefficient and the time deviation coefficient to construct an updated evaluation model and generate an updated evaluation coefficient. The expression is: ; where is the updated evaluation coefficient, and are the proportionality coefficients of the Q-value difference coefficient and the time deviation coefficient respectively, and are both greater than 0;

[0042] Set the updated evaluation coefficient threshold, compare the updated evaluation coefficient with the updated evaluation coefficient threshold. If the updated evaluation coefficient is greater than the updated evaluation coefficient threshold, continue to use the updated discount factor. If the updated evaluation coefficient is less than the updated evaluation coefficient threshold, use the basic discount factor.

[0043] In a preferred embodiment, an FSM simulation system based on a virtual factory robot includes a Q-Learning algorithm module, a data analysis module, a discount factor dynamic adjustment module, and a selection module, and the modules are connected by signals;

[0044] The Q-Learning algorithm module is used to optimize the decision-making process of the robot using the Q-Learning algorithm and learn the best strategy for taking different actions in different states;

[0045] The data analysis module is used to determine the uncertainty information and dependency information of future tasks, comprehensively analyze the uncertainty information and dependency information of future tasks, and construct an adjustment evaluation model;

[0046] A discount factor dynamic adjustment module, which is used to dynamically adjust the discount factor of the Q-Learning algorithm and analyze the working performance of the robot under the updated discount factor and the basic discount factor;

[0047] A selection module, which is used to determine the Q-value difference and efficiency deviation existing in the robot when performing the same work according to the working performance of the robot under the updated discount factor and the basic discount factor.

[0048] The technical effects and advantages of the present invention:

[0049] In the present invention, the FSM simulation of the robot is carried out by using the Q-Learning algorithm in a virtual factory. By dynamically adjusting the discount factor and the reward mechanism, it can better adapt to the uncertainties and dependencies in the factory environment. And by dynamically adjusting the discount factor and the comprehensive evaluation mechanism, and selectively using the dynamically adjusted discount factor, it helps the robot to complete tasks more efficiently in a complex and dynamic factory environment, realizing intelligent production management. Description of the Drawings

[0050] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the accompanying drawings;

[0051] Figure 1 It is a schematic flow chart of a method for FSM simulation of a robot based on a virtual factory of the present invention;

[0052] Figure 2 It is a schematic block diagram of the structure of a system for FSM simulation of a robot based on a virtual factory of the present invention. Specific Embodiments

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. Embodiment 1

[0054] As Figure 1 A schematic flow chart of a method for FSM simulation of a robot based on a virtual factory is given, which specifically includes the following steps:

[0055] S1: Use the Q-Learning algorithm in the FSM simulation to optimize the decision-making process of the robot, and extract the current task and future tasks of the robot during the decision-making process of the robot;

[0056] S2: Analyze the historical reward values of different future tasks of the robot, conduct exponentially weighted moving average analysis and probability density analysis on the historical reward values of future tasks, determine the uncertainty information of future tasks, and obtain the dependency information of future tasks through comparative analysis;

[0057] S3: Conduct a comprehensive analysis of the uncertainty information and dependency information of future tasks, dynamically adjust the discount factor of the Q-Learning algorithm, and determine the updated discount factor and the base discount factor;

[0058] S4: Combine the working performance of the robot under the updated discount factor and the base discount factor, analyze the difference information and efficiency information, and selectively use the updated discount factor.

[0059] Based on the real-time situation of the factory, conduct virtual construction, and use the robots in the model library to simulate actual operations. Through the simulation environment, vividly reproduce the real-time state of factory production, and perform various tasks through the robot model, including:

[0060] Understand the operation of various factory equipment through virtual simulation, schedule robots for task allocation in real time. Before introducing new equipment or optimizing the production line, test the feasibility and potential problems of the simulation plan through simulation, use the simulation platform to simulate robot behavior, analyze the task execution efficiency of the robot, and adjust its work process;

[0061] By using modeling tools, the simulation system can automatically generate a virtual environment corresponding to the factory layout, equipment configuration, process flow, etc. according to the real-time data of the factory;

[0062] The robots in the model library represent the robots used in actual production. The simulation system extracts the robots that meet the task requirements from the model library and conducts simulation operations for actual factory tasks.

[0063] For virtual scene restoration, first, it is necessary to restore the virtual scene of the real production line in the FSM core factory software, which involves two parts of work: modeling and construction, including:

[0064] Modeling is carried out through the FSM model editor plugin of the core factory. The plugin can realize the modeling and configuration of common models such as standard industrial robots, non-standard servo mechanisms, grippers, suction cups, conveyor belts, sensors, stackers, and shelves;

[0065] Users can import the modeled models into the FSM simulation engine of the core factory, and then realize the construction and restoration of the entire line in the virtual scene through the construction function of the software.

[0066] Based on the established virtual factory, the Q-learning algorithm is used for the robots in the virtual factory to autonomously learn task priorities according to real-time environment and task changes, and optimize their state transition decisions, including:

[0067] Determine the state information of the robot in the virtual factory, including the robot's current position, target position, task status, surrounding environment information, and resource status, etc. The state information will be used as the input of the Q-learning algorithm to guide the robot to decide the next action to be executed;

[0068] Determine the action information of the robot in the virtual factory, including the robot's moving direction, task selection, and path planning, etc. The action information will be used as the output of the Q-learning algorithm to indicate the actions that the robot should take in the current state;

[0069] Define a reward function to guide the robot to obtain feedback through actions. When defining the reward, the goals and requirements in the actual scenario need to be considered.

[0070] Among them, the expression for updating the Q value by the Q-learning algorithm is: ; among them, the expression means that in state take action After that, according to the next state and the obtained reward , update the current Q value, where represents the current state, means the action executed in state , means the immediate reward obtained after executing the action , means the next state reached after executing the action , represents the discount factor, represents the learning rate, means the optimal action value (i.e., the maximum Q value) in the next state , that is, it represents the expected reward brought by choosing the best action in the next state.

[0071] Dynamically adjust the discount factor in the virtual factory to help the robot better adapt to the changes in the environment and the complexity of the tasks. Therefore, collect the uncertainty information and dependency information of the tasks in the robot's future tasks, represent the uncertainty information through the reward fluctuation amplitude coefficient and the probability distribution coercion coefficient, and represent the dependency information through the priority influence coefficient.

[0072] It should be noted that the discount factor is usually a fixed parameter in the reinforcement learning algorithm. When the robot faces multi-objective tasks, especially when the priorities of multi-objective tasks are the same, the discount factor should be adjusted dynamically in real time according to factors such as the robot's tasks, environment, and priorities.

[0073] In a factory environment, the volatility of task rewards is affected by various factors, including the production environment, task complexity, human factors, and randomness. There is randomness or uncertainty in the robot's reward mechanism itself, including:

[0074] During the process of the robot's work, the task execution environment may be affected by external factors, resulting in the nature and priority of the task may change at any time. For example, suddenly increasing the production volume of a certain product will cause the production tasks of other products to be postponed, thus affecting its reward;

[0075] During the production process, the quality of the product may fluctuate due to various factors, and these fluctuations will directly affect the reward of the production task. For example, improper equipment calibration may lead to a decrease in product consistency and affect the reward.

[0076] The acquisition logic of the reward fluctuation amplitude coefficient is as follows: extract the future tasks of the robot, collect the historical data of the robot in the future tasks, and use the exponentially weighted moving average to smooth the reward volatility of the future tasks. The expression of the exponentially weighted moving average is: ; where represents the smoothed value at the current moment, represents the smoothing factor, represents the smoothed value at the previous moment, represents the immediate reward value of the future task;

[0077] It should be noted that using the exponentially weighted moving average (EWMA) can be adjusted dynamically according to the latest data, especially in the case of fluctuating, trending, or noisy task rewards, it can more accurately quantify the uncertainty of the task. EWMA assigns different weights to historical reward values, with larger weights for newer reward values and gradually decreasing weights for older reward values.

[0078] According to the historical data of the future tasks, obtain the smoothed value of each future task, calculate the difference between the smoothed value of each future task and the actual reward value, and mark the difference between the smoothed value of each future task and the actual reward value as: , where , is the smoothed value of each future task, is the actual reward value of each future task, n = 1, 2, 3,..., N, N is a positive integer, and n is the number of the historical data of the future tasks;

[0079] Calculate the reward fluctuation amplitude coefficient, and the calculation formula is: ; where is the reward fluctuation amplitude coefficient.

[0080] As can be seen from the formula, the larger the reward fluctuation amplitude coefficient, the higher the uncertainty of future rewards. The discount factor can be appropriately reduced so that when the robot selects tasks, it is more inclined to choose the current stable tasks rather than the tasks that may have large fluctuations in the future.

[0081] For the historical data obtained by simulating the robot using the FSM, it may be incomplete or noisy. It becomes unreliable to directly use a single historical reward value to evaluate future rewards. By constructing a probability density function, the potential variation range and distribution characteristics of rewards can be captured more comprehensively, and the entropy of continuous random variables can be used to quantify the uncertainty of future task rewards.

[0082] The acquisition logic of the probability distribution stress coefficient is as follows: According to the historical data of the robot in future tasks, use the Gaussian function to represent the probability density function of future task reward values, and the expression is: ; where is the standard deviation of future task reward values, is the average value of future task reward values, , ;

[0083] Calculate the probability distribution stress coefficient, and the calculation formula is: ; where is the minimum value of the reward value of the future task in the historical data, is the maximum value of the reward value of the future task in the historical data.

[0084] As can be seen from the formula, the larger the probability distribution stress coefficient, the more dispersed the distribution of future task reward values, and the higher the resulting uncertainty. When the robot selects tasks, it is more inclined to choose the current stable tasks rather than the tasks that may have greater uncertainty in the future.

[0085] Considering the influence of the dependency relationship on the discount factor is mainly to help the robot more reasonably balance the benefits of current tasks and future tasks. When the future task has a strong dependency relationship with the current task, if the potential benefit of the future task is high, the robot may need a higher discount factor to value the return of the future task, which can prevent the robot from investing too many resources in the current task and missing high-benefit tasks in the future.

[0086] The acquisition logic of the priority influence coefficient is as follows: Based on the historical data of the robot's future tasks, determine the average reward value after the robot selects a future task after completing the current task, and mark the average reward value after the robot selects a future task after completing the current task as: , where i = 1, 2, 3,..., I, I is a positive integer, and i is the number of the future task;

[0087] Obtain the maximum and minimum values of the reward values after the robot selects a future task after completing the current task, and mark the maximum and minimum values of the reward values after the robot selects a future task after completing the current task as: and ;

[0088] Calculate the priority influence coefficient, and the calculation formula is: ; where, is the priority influence coefficient.

[0089] It can be seen from the formula that the larger the priority influence coefficient, the higher the average reward value of this future task compared to other tasks. The importance of high-reward tasks can be emphasized by increasing the discount factor.

[0090] Through comprehensive analysis of the uncertainty information and dependency information of tasks in the robot's tasks, perform weighted calculations on the reward fluctuation amplitude coefficient, probability distribution coercion coefficient, and priority influence coefficient, construct an adjustment evaluation model, and generate an adjustment evaluation coefficient. The expression is: ; where, is the adjustment evaluation coefficient, , , are the proportionality coefficients of the reward fluctuation amplitude coefficient, probability distribution coercion coefficient, and priority influence coefficient, , , are all greater than 0.

[0091] Set the adjustment evaluation coefficient threshold, compare the adjustment evaluation coefficient with the adjustment evaluation coefficient threshold. If the adjustment evaluation coefficient is greater than the adjustment evaluation coefficient threshold, it means that the importance of the future task is relatively high, and the adjustment evaluation coefficient is used to adjust the discount factor. The Q value of the future task can be enhanced by increasing the discount factor. If the adjustment evaluation coefficient is less than the adjustment evaluation coefficient threshold, it means that the importance of the future task is relatively low, and the adjustment evaluation coefficient is not used to adjust the discount factor. The Q value of the future task can be weakened by reducing the discount factor.

[0092] Set the basic discount factor, adjust the discount factor with the adjustment evaluation coefficient, determine the updated discount factor according to the current working environment of the robot, and mark the updated discount factor as: , where, , where is the correction coefficient.

[0093] Simulate according to the FSM after using the updated discount factor, analyze the working performance of the robot under the updated discount factor and the basic discount factor. By setting the task queue in the same order, conduct a comparative experiment, collect difference information and efficiency information, represent the difference information by the Q-value difference coefficient, and represent the efficiency information by the time deviation coefficient.

[0094] When the same robot conducts a comparative experiment on the task queue in the same order, conduct experiments respectively by using the Q-Learning algorithm with the updated discount factor and the Q-Learning algorithm with the basic discount factor, and obtain the cumulative sum of Q-values after using the Q-Learning algorithm with the updated discount factor and the Q-Learning algorithm with the basic discount factor respectively, and compare the final Q-values, including:

[0095] Ensure that the two experiments (using the updated discount factor and not using the updated discount factor) are consistent in terms of environment, tasks, initial conditions, learning rate, exploration strategy (such as ε-greedy), etc., to exclude the interference of external factors on the results;

[0096] Conduct enough experiments to ensure the statistical significance of the results. The results of a single experiment may be affected by accidental factors, and increasing the number of experiments helps to obtain more reliable conclusions;

[0097] During the experiment, monitor the learning curve of Q-values (such as the change trend of Q-values, cumulative rewards, convergence speed, etc.) to detect abnormal situations in the learning process in a timely manner.

[0098] The acquisition logic of the Q-value difference coefficient is as follows: When the same robot conducts multiple comparative experiments on the task queue in the same order, conduct experiments respectively by using the Q-Learning algorithm with the updated discount factor and the Q-Learning algorithm with the basic discount factor, and obtain the cumulative sum of Q-values after using the Q-Learning algorithm with the updated discount factor and the Q-Learning algorithm with the basic discount factor respectively, and mark the average cumulative sum of Q-values of the Q-Learning algorithm with the updated discount factor as: Mark the average cumulative sum of Q-values of the Q-Learning algorithm with the basic discount factor as: ;

[0099] Calculate the Q-value difference coefficient, and the calculation formula is: ; where is the Q-value difference coefficient.

[0100] As can be seen from the formula, the greater the Q-value difference coefficient, the better the performance of the Q-Learning algorithm using the updated discount factor, indicating that a higher cumulative Q-value means the robot obtains more total rewards when performing tasks.

[0101] The acquisition logic of the time deviation coefficient is as follows: For the same robot, conduct multiple comparative experiments on the same task queue in the same order. Conduct experiments using the Q-Learning algorithm with the updated discount factor and the Q-Learning algorithm with the basic discount factor respectively. Compare the time used after completing the same number of tasks. Mark the average time used by the Q-Learning algorithm with the updated discount factor as: Mark the average time used by the Q-Learning algorithm with the basic discount factor as: ;

[0102] Calculate the time deviation coefficient. The calculation formula is: ; where is the time deviation coefficient.

[0103] As can be seen from the formula, the smaller the time deviation coefficient, the better the performance of the Q-Learning algorithm using the updated discount factor, indicating that the robot may complete tasks faster when using the updated discount factor.

[0104] Perform weighted calculation on the Q-value difference coefficient and the time deviation coefficient to construct an update evaluation model and generate an update evaluation coefficient. The expression is: ; where is the update evaluation coefficient, and are the proportionality coefficients of the Q-value difference coefficient and the time deviation coefficient respectively, and are both greater than 0.

[0105] Set the update evaluation coefficient threshold, and compare the update evaluation coefficient with the update evaluation coefficient threshold. If the update evaluation coefficient is greater than the update evaluation coefficient threshold, continue to use the updated discount factor, indicating that using the updated discount factor can improve the working efficiency of the robot. If the update evaluation coefficient is less than the update evaluation coefficient threshold, use the basic discount factor.

[0106] In the present invention, the Q-Learning algorithm is used to perform FSM simulation on the robot in a virtual factory. By dynamically adjusting the discount factor and the reward mechanism, it better adapts to the uncertainties and dependencies in the factory environment. And by dynamically adjusting the discount factor and the comprehensive evaluation mechanism, selectively using the dynamically adjusted discount factor helps the robot complete tasks more efficiently in a complex and dynamic factory environment, realizing intelligent production management. Embodiment 2

[0107] As shown in Figure 2 Figure 1 shows a schematic block diagram of the structure of an FSM simulation system based on a virtual factory robot, which specifically includes a Q-Learning algorithm module, a data analysis module, a discount factor dynamic adjustment module, and a selection module, with signal connections between the modules;

[0108] The Q-Learning algorithm module is used to optimize the decision-making process of the robot using the Q-Learning algorithm and learn the best strategies for taking different actions in different states;

[0109] The data analysis module is used to determine the uncertainty information and dependency information of future tasks, comprehensively analyze the uncertainty information and dependency information of future tasks, and construct an adjustment evaluation model;

[0110] The discount factor dynamic adjustment module is used to dynamically adjust the discount factor of the Q-Learning algorithm and analyze the working performance of the robot using the updated discount factor and the basic discount factor;

[0111] The selection module is used to determine the Q-value difference and efficiency deviation existing in the robot when performing the same work according to the working performance of the robot using the updated discount factor and the basic discount factor.

[0112] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula that is closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0113] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0114] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not imply the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0115] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0116] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0117] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0118] If the described functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, and other various media that can store program codes.

[0119] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. An FSM simulation method based on a virtual factory robot, characterized in that Specifically, it includes the following steps: S1: Use the Q-Learning algorithm in FSM simulation to optimize the decision-making process of the robot, and extract the current task and future tasks of the robot during the decision-making process of the robot; S2: Analyze the historical reward values of different future tasks of the robot, conduct exponentially weighted moving average analysis and probability density analysis on the historical reward values of future tasks, determine the uncertainty information of future tasks, and obtain the dependency information of future tasks through comparative analysis; S3: Conduct a comprehensive analysis of the uncertainty information and dependency information of future tasks, dynamically adjust the discount factor of the Q-Learning algorithm, and determine the updated discount factor and the basic discount factor; S4: Combine the working performance of the robot under the updated discount factor and the basic discount factor, analyze the difference information and efficiency information, and selectively use the updated discount factor; Among them, the uncertainty information is represented by the reward fluctuation amplitude coefficient and the probability distribution coercion coefficient; The acquisition logic of the reward fluctuation amplitude coefficient is as follows: extract the future tasks of the robot, collect the historical data of the robot in the future tasks, and use the exponentially weighted moving average to smooth the reward volatility of the future tasks. The expression of the exponentially weighted moving average is: ; where represents the smoothed value at the current moment, represents the smoothing factor, represents the smoothed value at the previous moment, represents the immediate reward value of the future task; Based on the historical data of future tasks, obtain the smoothed value for each task of the future task, calculate the difference between the smoothed value for each task of the future task and the actual reward value, and mark the difference between the smoothed value for each task of the future task and the actual reward value as: , where , is the smoothed value for each task of the future task, is the actual reward value for each task of the future task, n = 1, 2, 3, ……, N, N is a positive integer, and n is the number of the historical data of the future task; Calculate the reward fluctuation amplitude coefficient, and the calculation formula is: ; where is the reward fluctuation amplitude coefficient; The acquisition logic of the probability distribution coercion coefficient is: According to the historical data of the robot collected during future tasks, use the Gaussian function to represent the probability density function of the future task reward value, and the expression is: ; where is the standard deviation of future task reward values, is the average value of future task reward values, , ; Calculate the probability distribution stress coefficient, and the calculation formula is: ; where is the minimum value of the reward value of the future task in the historical data, is the maximum value of the reward value of the future task in the historical data; The dependency information of future tasks is represented by the priority influence coefficient; The acquisition logic of the priority influence coefficient is as follows: Based on the historical data of the robot's future tasks, determine the average reward value after the robot selects a future task after completing the current task, and mark the average reward value after the robot selects a future task after completing the current task as: , where i = 1, 2, 3,..., I, I is a positive integer, and i is the number of the future task; Obtain the maximum and minimum values of the reward value after the robot selects a future task after completing the current task, and mark the maximum and minimum values of the reward value after the robot selects a future task after completing the current task as: and ; Calculate the priority influence coefficient, and the calculation formula is: where is the priority influence coefficient.

2. The FSM simulation method based on a virtual factory robot according to claim 1, characterized in that, Using the Q-Learning algorithm in FSM simulation to optimize the decision-making process of the robot includes: Using the Q-learning algorithm for the robot in the virtual factory to help the robot autonomously learn task priorities according to the real-time environment and task changes, and optimize its state transition decision-making, including: Determine the state information of the robot in the virtual factory, including the current position, target position, task status, surrounding environment information, and resource status of the robot. The state information will be used as the input of the Q-learning algorithm to guide the robot to decide the next action to be executed; Determine the action information of the robot in the virtual factory, including the moving direction, task selection, and path planning of the robot. The action information will be used as the output of the Q-learning algorithm to indicate the actions that the robot should take in the current state.

3. A FSM simulation method based on a virtual factory robot according to claim 2, characterized in that, Dynamically adjusting the discount factor of the Q-Learning algorithm includes: Through the comprehensive analysis of the uncertainty information and dependency information of tasks in the future tasks of the robot, the reward fluctuation amplitude coefficient, the probability distribution coercion coefficient, and the priority influence coefficient are weighted and calculated to construct an adjustment evaluation model and generate an adjustment evaluation coefficient. The expression is as follows: ; where is the adjustment evaluation coefficient, , , are the proportionality coefficients of the reward fluctuation amplitude coefficient, the probability distribution coercion coefficient, and the priority influence coefficient, , , are all greater than 0 respectively; Set the adjustment evaluation coefficient threshold, compare the adjustment evaluation coefficient with the adjustment evaluation coefficient threshold. If the adjustment evaluation coefficient is greater than the adjustment evaluation coefficient threshold, then adjust the discount factor with the adjustment evaluation coefficient. If the adjustment evaluation coefficient is less than the adjustment evaluation coefficient threshold, then do not adjust the discount factor with the adjustment evaluation coefficient; Set a basic discount factor, adjust the discount factor by adjusting the evaluation coefficient, determine the updated discount factor according to the current working environment of the robot, and mark the updated discount factor as: , where , where is the correction coefficient.

4. A FSM simulation method based on a virtual factory robot according to claim 3, characterized in that, The difference information includes: The difference information is represented by the Q-value difference coefficient; The acquisition logic of the Q-value difference coefficient is as follows: Conduct multiple comparative experiments on the same task queue in the same order by the same robot, and conduct experiments using the Q-Learning algorithm with the updated discount factor and the Q-Learning algorithm with the basic discount factor respectively, and obtain the cumulative sum of Q-values after using the Q-Learning algorithm with the updated discount factor and the Q-Learning algorithm with the basic discount factor respectively. Mark the average cumulative sum of Q-values of the Q-Learning algorithm with the updated discount factor as: , and mark the average cumulative sum of Q-values of the Q-Learning algorithm with the basic discount factor as: ; Calculate the Q-value difference coefficient, and the calculation formula is: ; where is the Q-value difference coefficient.

5. A FSM simulation method based on a virtual factory robot according to claim 4, characterized in that, The efficiency information includes: The efficiency information is represented by the time deviation coefficient; The acquisition logic of the time deviation coefficient is as follows: Multiple comparative experiments are carried out on the same task queue in the same order by the same robot. Experiments are respectively carried out by using the Q-Learning algorithm with the updated discount factor and the Q-Learning algorithm with the basic discount factor. The time used after completing the same number of tasks is compared, and the average time used by the Q-Learning algorithm with the updated discount factor is marked as: , and the average time used by the Q-Learning algorithm with the basic discount factor is marked as: ; Calculate the time deviation coefficient, and the calculation formula is: ; where is the time deviation coefficient.

6. A FSM simulation method based on a virtual factory robot according to claim 5, characterized in that, Selectively using the updated discount factor includes: The Q-value difference coefficient and the time deviation coefficient are weighted and calculated to construct an updated evaluation model and generate an updated evaluation coefficient. The expression is as follows: ; where is the updated evaluation coefficient, and are the proportionality coefficients of the Q-value difference coefficient and the time deviation coefficient respectively, and are both greater than 0; Set the updated evaluation coefficient threshold, compare the updated evaluation coefficient with the updated evaluation coefficient threshold. If the updated evaluation coefficient is greater than the updated evaluation coefficient threshold, then continue to use the updated discount factor. If the updated evaluation coefficient is less than the updated evaluation coefficient threshold, then use the basic discount factor.

7. An FSM simulation system based on a virtual factory robot, which is used to implement an FSM simulation method based on a virtual factory robot according to any one of claims 1-6, characterized in that, It includes a Q-Learning algorithm module, a data analysis module, a discount factor dynamic adjustment module, and a selection module, with signal connections among the modules; The Q-Learning algorithm module is used to optimize the decision-making process of the robot using the Q-Learning algorithm and learn the best strategies for taking different actions in different states; The data analysis module is used to determine the uncertainty information and dependency information of future tasks, comprehensively analyze the uncertainty information and dependency information of future tasks, and construct an adjustment evaluation model; The discount factor dynamic adjustment module is used to dynamically adjust the discount factor of the Q-Learning algorithm and analyze the working performance of the robot under the updated discount factor and the basic discount factor; The selection module is used to determine the Q-value difference and efficiency deviation existing in the robot when performing the same work according to the working performance of the robot under the updated discount factor and the basic discount factor.

Citation Information

Patent Citations

  • A virtual simulation training system based on industrial robots

    CN116229792B

  • Planning method for behavior system structure of intelligent underwater robot based on deep Q learning

    CN108873687A

  • Production line mobile robot aggregation type recovery warehousing simulation method and system

    CN113110101A