A World Model Perturbation Method, Device, Equipment and Storage Medium
By constructing a target perturbation model, combining the target world model and the random residual model set for state perturbation prediction, the problem of unstable data performance of the world model outside the distribution is solved, and the robustness of the reinforcement learning strategy and the effect in the real environment is improved.
Patent Information
- Application Number
- CN202411220476.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-02
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-09-02
AI Technical Summary
The performance of existing world models on out-of-distribution data is unstable, resulting in inaccurate decision-making of reinforcement learning strategies in real environments.
By constructing a target perturbation model, combining the target world model and a random residual model set, state perturbation prediction is performed to increase the coverage of the world model simulation data.
It improves the performance of world models on out-of-distribution data, enhances the robustness of reinforcement learning strategies, and ensures effectiveness in real physical environments.
Smart Images

Figure CN119129695B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method, device, equipment and storage medium for perturbing a world model. Background Art
[0002] A world model is a method used in the fields of artificial intelligence and machine learning to describe and understand the environment. Based on the current state and past observations, the world model can predict future states. By simulating possible future states, the world model helps agents (such as robots or virtual agents) make better decisions, enabling them to try different actions in the internal model and select those that can maximize rewards or achieve goals.
[0003] Model-based reinforcement learning is a specific application of the world model, which accelerates the process of agent reinforcement learning by utilizing the world model. In this method, the agent not only obtains experience by interacting with the actual environment but also simulates future states and results through the world model, thereby conducting a large number of trials and errors in such a virtual world model.
[0004] Due to limited training data and the inability of training data to cover all possible situations, the world model often can only learn accurately in a part of the region. For data outside the training scope, the world model often has to "guess" based on its generalization ability. The "guessed" data is what we call "out-of-distribution data". The performance of the world model on out-of-distribution data is not guaranteed and is very likely to be incorrect. Therefore, the policy obtained through reinforcement learning in the world model also cannot make accurate decisions on out-of-distribution data. Summary of the Invention
[0005] The present invention provides a method, device, equipment and storage medium for perturbing a world model. By perturbing the output of the world model, the coverage rate of the simulated data of the world model is increased, so that the reinforcement learning policy trained in the world model can be robust to out-of-distribution data.
[0006] According to one aspect of the present invention, a method for perturbing a world model is provided. The method includes:
[0007] Obtain the task scenario state and task execution action of the currently executing task;
[0008] Input the task scenario state and the task execution action into a pre-trained target perturbation model for state perturbation prediction, where the target perturbation model is composed of a target world model with the same model dimension and a set of random residual models, and the set of random residual models includes at least one neural network model;
[0009] Predict the target task scenario state of the current executing task at the next moment according to the output of the target world model.
[0010] According to another aspect of the present invention, there is provided a world model perturbation device. The device includes:
[0011] A task information acquisition module, configured to acquire the task scenario state and task execution actions of the current executing task;
[0012] A task state perturbation module, configured to input the task scenario state and the task execution actions into a pre-trained target perturbation model for state perturbation prediction, where the target perturbation model is composed of a target world model and a set of random residual models with the same model dimension, and the set of random residual models includes at least one neural network model;
[0013] A scenario state prediction module, configured to predict the target task scenario state of the current executing task at the next moment according to the output of the target world model.
[0014] According to another aspect of the present invention, there is provided an electronic device, the electronic device includes:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the world model perturbation method according to any embodiment of the present invention.
[0018] According to another aspect of the present invention, there is provided a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the world model perturbation method according to any embodiment of the present invention when executed.
[0019] The technical solution of the embodiment of the present invention obtains the task scenario state and task execution actions of the currently executing task. The task scenario state and the task execution actions are input into a pre-trained target perturbation model for state perturbation prediction, where the target perturbation model is composed of a target world model and a set of random residual models with the same model dimension, and the set of random residual models includes at least one neural network model. According to the output of the target world model, the target task scenario state of the currently executing task at the next moment is predicted. By performing parameter perturbation in the parameter space of the world model, the coverage of the simulated data of the world model is improved, so that the agent can see a large amount of data during the reinforcement learning training in the world model, improve its effect on out-of-distribution data, and finally ensure the effect in the real physical environment.
[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0022] Figure 1 is a flowchart of a world model perturbation method provided in Embodiment 1 of the present invention;
[0023] Figure 2 is a flowchart of a world model perturbation method provided in Embodiment 2 of the present invention;
[0024] Figure 3 is a schematic diagram of the principle of the target perturbation model provided in Embodiment 2 of the present invention;
[0025] Figure 4 is a structural diagram of a world model perturbation device provided in Embodiment 3 of the present invention;
[0026] Figure 5 is a schematic diagram of the structure of an electronic device for implementing the world model perturbation method of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0029] Embodiment 1
[0030] Figure 1 A flowchart of a world model perturbation method is provided for Embodiment 1 of the present invention. This embodiment is applicable to the situation of improving the coverage of world model simulation data. This method can be executed by a world model perturbation device, which can be implemented in the form of hardware and / or software, and the world model perturbation device can be configured in an electronic device. As Figure 1 shown, the method includes:
[0031] S101. Obtain the task scenario state and task execution actions of the currently executing task.
[0032] Among them, the currently executing task may refer to any execution task that requires state prediction. The task scenario state refers to the specific situation of the scenario of the currently executing task. The task scenario state may include various types of information, specifically depending on the application field. For example, in a game environment, the task scenario state may include the position, speed, health value, item holding situation, etc. of the character; in a robot control task, the state may include the joint angles, speed, acceleration, and sensor readings of the robot.
[0033] The task execution action refers to the specific action taken by the currently executing task, that is, the behavior selected and executed by the agent according to its strategy. The task execution action can be continuous (such as controlling the angle change of a robot arm) or discrete (such as selecting a moving direction in a game).
[0034] Exemplarily, when the current execution task is a maze navigation task, the task scenario state can be used to describe the position of the agent in the maze and other possible relevant information. For example, in a maze composed of grids, each grid can be represented by a pair of coordinates (x, y). Therefore, the task scenario state can be simply represented as the coordinates (x, y) of the agent's current position, the visible part of the maze map, the direction of the agent, whether it is carrying a collected item, and the remaining task time, etc. The task execution action can be an action that the agent can take. For example, in the maze navigation task, the task execution action is usually a moving direction (such as moving up, moving down, moving left, and moving right, etc.). In some cases, the task execution action may also include staying still, collecting items, and placing items, etc.
[0035] S102. Input the task scenario state and the task execution action into a pre-trained target perturbation model for state perturbation prediction.
[0036] It should be noted that the target perturbation model can be composed of a target world model and a set of random residual models. The target world model and the random residual models should be models with the same dimension. For example, both the target world model and the random residual models have a 3-layer fully connected neural network.
[0037] Among them, the set of random residual models includes at least one neural network model. The neural network models in the set of random residual models are usually simple and randomly initialized neural networks, which can be used to generate perturbations without being trained. The neural network models in the set of random residual models can include, but are not limited to, fully connected networks, convolutional neural networks, recurrent neural networks, variational autoencoders, and generative adversarial networks, etc.
[0038] The target world model can be obtained by collecting a batch of real environment data of state-action-next moment state, and optimizing the world model using the MSE loss function and the Adam optimizer.
[0039] Specifically, use the task scenario state and the task execution action as prediction conditions and input them into a pre-trained target perturbation model, so that the target perturbation model performs state perturbation prediction.
[0040] Exemplarily, the process of state perturbation prediction in the target perturbation model is as follows:
[0041] Based on the task scenario state, the task execution action, and the target world model, obtain the predicted task scenario state; based on the task scenario state, the task execution action, and the set of stochastic residual models, determine the incremental task scenario state; combine the predicted task scenario state and the incremental task scenario state for processing to enable the target perturbation model to perform state perturbation prediction.
[0042] Among them, the predicted task scenario state refers to the task scenario state predicted by the target world model for the next moment. The incremental task scenario state can refer to the perturbation amount provided by the set of stochastic residual models.
[0043] That is to say, input the task scenario state and the task execution action into the target world model for scenario state prediction, and the predicted task scenario state can be obtained. Input the task scenario state and the task execution action into the set of stochastic residual models for scenario state prediction, and the incremental task scenario state can be obtained. Combining the predicted task scenario state and the incremental task scenario state can obtain the final state to enable the target perturbation model to perform state perturbation prediction.
[0044] It should be noted that the combination processing of the predicted task scenario state and the incremental task scenario state can be the addition processing of the predicted task scenario state and the incremental task scenario state, or the multiplication operation of the predicted task scenario state and the incremental task scenario state with their respective preset coefficients, and then the addition processing of the multiplication results.
[0045] Continuing with the example where the current task being executed is a maze navigation task, the starting state of the task scenario state may be (1,1), which means the agent starts at the lower left corner of the maze. If the agent chooses to move right, the task execution action is to move right. The target world model predicts the next state as (2,1) based on the current state (1,1) and the action "move right". The neural network model in the set of stochastic residual models may give a small offset, such as (-1,0), which means the agent may deviate slightly from the expected path. The output result of the target perturbation model is (2,1)+(-1,0)=(1,1). Then the state at the next moment is updated to (1,1).
[0046] It should be noted that the example here is idealized, and in actual applications, the target perturbation model may give more complex outputs. In addition, the specific representations of the task scenario state and the task execution action depend on the design of the maze environment and the complexity of the task. The example is only for detailed illustration and not a limitation of the present invention.
[0047] Exemplarily, based on the task scenario state, the task execution action, and the set of stochastic residual models, determining the incremental task scenario state includes: screening a target neural network model from the set of stochastic residual models based on a preset screening method; inputting the task scenario state and the task execution action into the target neural network model for incremental calculation to obtain the incremental task scenario state.
[0048] The preset screening method can be set according to actual needs. For example, the screening method can be set according to actual experience, and it can be the screening method for the same type of task. The present invention does not limit this. Preferably, the present invention adopts a random screening method, and the screening quantity is single. That is, randomly screen a neural network model from the set of stochastic residual models, and determine this neural network model as the target neural network model. Input the task scenario state and the task execution action into the target neural network model for incremental calculation, and obtain the incremental task scenario state according to the output of the target neural network model, thereby improving the randomness of the increment. Furthermore, all possible situations are simulated through the target perturbation model, hoping to cover the real situation, so that the policy can also see the real situation when training in the world model, and improve the robustness of the policy on out-of-distribution data.
[0049] S103. Predict the target task scenario state of the currently executing task at the next moment according to the output of the target perturbation model.
[0050] Among them, the target task scenario state may refer to the perturbation scenario state obtained by the target perturbation model for perturbation prediction.
[0051] Specifically, receive the output of the target perturbation model, and determine the output result as the target task scenario state of the currently executing task at the next moment. After obtaining the target task scenario state, define appropriate evaluation metrics to measure the performance of the agent policy, including the performance on in-distribution data and out-of-distribution data. At the same time, the performance of the agent policy under different perturbation intensities and the generalization ability on different tasks can also be tested.
[0052] Exemplarily, the expression for predicting the target task scenario state of the currently executing task at the next moment is as follows:
[0053] s′ = f(s,a;θ) + clip(g(s,a;φ i ), -c, c);
[0054] i ∼ uniform({1,2,...,N});
[0055] Among them, f(s,a;θ) refers to the target world model, g(s,a;φ iwhere represents the set of random residual models, N represents the number of neural network models in the set of random residual models, i represents the i-th neural network model in the set of random residual models, s represents the task scenario state, a represents the task execution action, θ represents the model parameters of the target world model, and φ i represents the model parameters of the set of random residual models, c represents the range of randomized perturbations, and the clip() function is used to limit the output range of the neural network model.
[0056] uniform means randomly selecting an element uniformly from the set within the parentheses, that is, randomly screening a neural network model from the set of random residual models. clip() means clipping the value of the first item in the parentheses so that the value is within the range of the second and third items in the parentheses. The clip() function is specifically used to limit the output range of the residual model to ensure that the perturbation is not too large. The value of c determines the amplitude of the perturbation and needs to be adjusted according to the actual application scenario to achieve the best perturbation effect without destroying the stability of the model, ensuring that the perturbation does not cause the state prediction to deviate too much from the normal range, thereby helping to maintain the stability of the model.
[0057] The technical solution of the embodiment of the present invention obtains the task scenario state and the task execution action of the currently executing task. Input the task scenario state and the task execution action into a pre-trained target perturbation model for state perturbation prediction, where the target perturbation model is composed of a target world model and a set of random residual models with the same model dimension, and the set of random residual models includes at least one neural network model. According to the output of the target world model, predict the target task scenario state of the currently executing task at the next moment. By performing parameter perturbation in the parameter space of the world model, the coverage of the simulated data of the world model is improved, so that the intelligent agent can see a large amount of data during the reinforcement learning training in the world model, improve its effect on out-of-distribution data, and ultimately ensure the effect in the real physical environment.
[0058] Embodiment 2
[0059] Figure 2 is a flowchart of a world model perturbation method provided by the second embodiment of the present invention. On the basis of the above embodiments, the training process of the target world model is specified. As Figure 2 shown, the method includes:
[0060] S201. Obtain sample offline data and the sample task scenario state corresponding to the sample offline data.
[0061] Among them, the sample offline data includes at least historical task scenario states and historical task execution actions.
[0062] Specifically, a large amount of state-action-new state-reward quadruple data needs to be collected. This can be achieved by having the agent interact with the environment randomly or according to a certain strategy. The dataset should be large enough and diverse to cover various possible states of the environment. Clean and preprocess the collected data, including but not limited to removing outliers and normalizing numerical features, so as to obtain sample offline data and sample task scenario states.
[0063] S202. Input the sample offline data into a preset training model for task scenario state prediction, and based on the output of the preset training model, obtain the output task scenario state.
[0064] Among them, the preset training model includes at least one of a regression model, a deep neural model, a generative adversarial network, or a variational autoencoder.
[0065] Specifically, use the sample offline data to train the selected preset training model. The training objective is to enable the preset training model to accurately predict the state at the next moment and obtain the output task scenario state.
[0066] Exemplarily, for a regression model, the mean squared error can be used as the loss function. For a probability model, the cross-entropy loss or other appropriate probability metrics can be used as the loss function.
[0067] S203. Determine the training error based on the output task scenario state and the sample task scenario state, and backpropagate the training error into the preset training model to adjust the network parameters in the preset training model.
[0068] Specifically, compare the output task scenario state with the sample task scenario state to determine the training error, and backpropagate the training error into the preset network model to adjust the network parameters in the preset network model until the preset convergence condition is met.
[0069] S204. When the preset convergence condition is met, determine that the training of the preset training model is completed, and obtain the target world model.
[0070] Specifically, the preset convergence condition can include when the number of iterations reaches a preset number or the training error converges.
[0071] When the preset convergence condition is met, determine that the training of the preset network model is completed. At this time, the preset network model that has completed training can be used as the target world model.
[0072] S205. Combine the target world model and the set of random residual models to construct the target perturbation model.
[0073] Figure 3This is the schematic diagram of the target perturbation model provided by the present invention. As Figure 3 shown, the target world model and the set of random residual models are combined in parallel. The task scenario state and the task execution action are respectively input into the target world model and the set of random residual models. The target world model outputs the predicted task scenario state at the next moment, while the set of random residual models randomly selects a neural network model and outputs an incremental task scenario state with the same dimension as the predicted task scenario state through the neural network model. Finally, the outputs of the two models are added together to obtain the target task scenario state with perturbation at the next moment.
[0074] S206. Obtain the task scenario state and the task execution action of the currently executing task.
[0075] S207. Input the task scenario state and the task execution action into the pre-trained target perturbation model for state perturbation prediction.
[0076] S208. Predict the target task scenario state of the currently executing task at the next moment according to the output of the target perturbation model.
[0077] The technical solution of the present invention trains the target world model through sample offline data and sample task scenario states to avoid high-cost trial and error of the reinforcement learning agent directly in the real physical world, and can ensure the effect of the agent in the real physical environment.
[0078] Embodiment III
[0079] Figure 4 This is the structural schematic diagram of a world model perturbation device provided by Embodiment III of the present invention. As Figure 4 shown, the device includes:
[0080] A task information acquisition module 301, configured to acquire the task scenario state and the task execution action of the currently executing task;
[0081] A task state perturbation module 302, configured to input the task scenario state and the task execution action into the pre-trained target perturbation model for state perturbation prediction, where the target perturbation model is composed of a target world model and a set of random residual models with the same model dimension, and the set of random residual models includes at least one neural network model;
[0082] A scenario state prediction module 303, configured to predict the target task scenario state of the currently executing task at the next moment according to the output of the target perturbation model.
[0083] The technical solution of the embodiment of the present invention obtains the task scenario state and task execution actions of the currently executing task. The task scenario state and the task execution actions are input into a pre-trained target perturbation model for state perturbation prediction, where the target perturbation model is composed of a target world model and a set of random residual models with the same model dimension, and the set of random residual models includes at least one neural network model. According to the output of the target world model, the target task scenario state of the currently executing task at the next moment is predicted. By performing parameter perturbation in the parameter space of the world model, the coverage of the simulation data of the world model is improved, so that the agent can see a large amount of data during the reinforcement learning training in the world model, improving its effect on out-of-distribution data, and ultimately ensuring the effect in the real physical environment.
[0084] Optionally, the task state perturbation module 302 includes:
[0085] The first scenario state prediction unit is configured to obtain a predicted task scenario state based on the task scenario state, the task execution action, and the target world model;
[0086] The second scenario state prediction unit is configured to determine an incremental task scenario state based on the task scenario state, the task execution action, and the set of random residual models;
[0087] The task state perturbation unit is configured to combine the predicted task scenario state and the incremental task scenario state for processing, so as to enable the target perturbation model to perform state perturbation prediction.
[0088] Optionally, the second scenario state prediction unit is specifically configured to:
[0089] Based on a preset screening method, a target neural network model is screened from the set of random residual models;
[0090] The task scenario state and the task execution action are input into the target neural network model for incremental calculation to obtain an incremental task scenario state.
[0091] Optionally, the task state perturbation unit is specifically configured to:
[0092] Add the predicted task scenario state and the incremental task scenario state, and determine the added result as the target task scenario state.
[0093] Optionally, the scenario state prediction module 303 is specifically configured to:
[0094] s′ = f(s, a; θ) + clip(g(s, a; φ i ), -c, c);
[0095] i ∼ uniform({1, 2,..., N});
[0096] Among them, f(s, a; θ) refers to the target world model, g(s, a; φ i ) refers to the set of stochastic residual models, N refers to the number of neural network models in the set of stochastic residual models, i refers to the i-th neural network model in the set of stochastic residual models, s refers to the task scenario state, a refers to the task execution action, θ refers to the model parameters of the target world model, φ i refers to the model parameters of the set of stochastic residual models, c refers to the range of randomized perturbations, and the clip() function is used to limit the output range of the neural network model.
[0097] Optionally, the device further includes:
[0098] A world model training module, configured to obtain sample offline data and the sample task scenario state corresponding to the sample offline data, where the sample offline data at least includes historical task scenario states and historical task execution actions; input the sample offline data into a preset training model for task scenario state prediction, and obtain an output task scenario state based on the output of the preset training model; determine a training error based on the output task scenario state and the sample task scenario state, and backpropagate the training error to the preset training model to adjust the network parameters in the preset training model; when a preset convergence condition is met, determine that the training of the preset training model ends, and obtain a target world model.
[0099] Optionally, the preset training model at least includes one of a regression model, a deep neural model, a generative adversarial network, or a variational autoencoder.
[0100] The world model perturbation device provided by the embodiments of the present invention can execute the world model perturbation method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0101] Embodiment 4
[0102] Figure 5FIG. 0 shows a schematic structural diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0103] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0104] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0105] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method of world model perturbation.
[0106] In some embodiments, the method of world model perturbation may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method of world model perturbation described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the method of world model perturbation by any other suitable means (e.g., by means of firmware).
[0107] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0108] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0109] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0110] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0111] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0112] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0113] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0114] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A perturbation method for a world model, applied to the fields of artificial intelligence and machine learning, characterized in that Including: Obtain the task scenario state and task execution action of the currently executing task; wherein, when the currently executing task is a maze navigation task, the task scenario state is the position coordinates of the agent in the maze, and the task execution action is the moving direction of the agent; Input the task scenario state and the task execution action into a pre-trained target perturbation model for state perturbation prediction, wherein the target perturbation model consists of a target world model and a set of random residual models with the same model dimension, and the set of random residual models includes at least one neural network model; Predict the target task scenario state of the currently executing task at the next moment according to the output of the target perturbation model; Among them, the inputting the task scenario state and the task execution action into a pre-trained target perturbation model for state perturbation prediction includes: Based on the task scenario state, the task execution action, and the target world model, obtain a predicted task scenario state; Based on a preset screening method, screen out a target neural network model from the set of random residual models; Input the task scenario state and the task execution action into the target neural network model for incremental calculation to obtain an incremental task scenario state; Combine the predicted task scenario state and the incremental task scenario state to enable the target perturbation model to perform state perturbation prediction.
2. The method according to claim 1, wherein The predicting the target task scenario state of the currently executing task at the next moment according to the output of the target world model includes: Add the predicted task scenario state and the incremental task scenario state, and determine the addition result as the target task scenario state.
3. The method according to claim 1, characterized in that The predicting the target task scenario state of the currently executing task at the next moment according to the output of the target world model includes: s′ = f(s, a; θ) + clip(g(s, a; φ i ), -c, c); i ∼ uniform({1, 2,..., N}); Among them, f(s,a;θ) refers to the target world model, g(s,a;φ i ) refers to the set of stochastic residual models, N refers to the number of neural network models in the set of stochastic residual models, i refers to the i-th neural network model in the set of stochastic residual models, uniform({1,2,...,N}) represents uniform random sampling from 1 to N, s refers to the task scenario state, a refers to the task execution action, θ refers to the model parameters of the target world model, φ i refers to the model parameters of the set of stochastic residual models, c refers to the range of randomized perturbations, and the clip() function is used to limit the output range of the neural network model.
4. The method according to claim 1, wherein The training process of the target world model includes: Obtain sample offline data and the corresponding sample task scenario state of the sample offline data, wherein the sample offline data at least includes historical task scenario states and historical task execution actions; Input the sample offline data into a preset training model for task scenario state prediction, and based on the output of the preset training model, obtain an output task scenario state; Determine a training error based on the output task scenario state and the sample task scenario state, and backpropagate the training error into the preset training model to adjust the network parameters in the preset training model; When a preset convergence condition is satisfied, determine that the training of the preset training model is completed, and obtain the target world model.
5. The method according to claim 4, characterized in that, The preset training model at least includes one of a regression model, a deep neural model, a generative adversarial network, or a variational autoencoder.
6. A world model perturbation device, applied to the fields of artificial intelligence and machine learning, characterized in that, Including: A task information acquisition module, configured to acquire the task scenario state and the task execution action of the currently executing task; wherein, when the currently executing task is a maze navigation task, the task scenario state is the position coordinates of the agent in the maze, and the task execution action is the moving direction of the agent; A task status perturbation module, configured to input the task scenario status and the task execution action into a target perturbation model obtained by pre-training for status perturbation prediction, wherein the target perturbation model is composed of a target world model and a random residual model set with the same model dimension, and the random residual model set includes at least one neural network model; A scenario status prediction module, configured to predict the target task scenario status of the currently executing task at the next moment according to the output of the target perturbation model; Wherein, the task status perturbation module includes: A first scenario status prediction unit, configured to obtain a predicted task scenario status based on the task scenario status, the task execution action, and the target world model; A second scenario status prediction unit, configured to screen and obtain a target neural network model from the random residual model set based on a preset screening method; input the task scenario status and the task execution action into the target neural network model for incremental calculation to obtain an incremental task scenario status; A task status perturbation unit, configured to combine the predicted task scenario status and the incremental task scenario status to enable the target perturbation model to perform status perturbation prediction.
7. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the world model perturbation method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the world model perturbation method according to any one of claims 1-5 when executed.
Citation Information
Patent Citations
Trend prediction method based on attention mechanism and reinforcement learning
CN114049222A
Model training method and device based on mixed disturbance, storage medium and server
CN115345301A