Machine action generation method and device based on environment image, equipment and medium
By acquiring machine trajectory data and optimizing the action generation model using reward signals from online reinforcement learning, the problem of insufficient accuracy of existing models under sudden environmental changes and unexpected situations is solved, achieving more efficient action generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-06
- Publication Date
- 2026-04-14
AI Technical Summary
Existing action generation models suffer from weak generalization ability, insufficient utilization of expert data, and unstable training due to sparse reward signals when faced with sudden changes in the market environment or unexpected situations during surgery, resulting in insufficient accuracy of machine action generation.
By acquiring machine trajectory data, generating behavioral actions using a pre-set action generation model, calculating the behavior cloning loss and optimizing the model, and combining reward signals and loss functions in online reinforcement learning, the action generation model is optimized to establish a reliable foundation for action generation, enhance autonomous exploration capabilities, and provide intensive guidance through multi-dimensional reward signals.
It improves the accuracy of machine motion generation, enabling the model to adapt to environmental changes more efficiently and optimize decisions, thereby enhancing the accuracy of motion generation in real-world scenarios.
Smart Images

Figure CN121859979A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent decision-making technology, and in particular to a method, apparatus, device, and medium for generating machine actions based on environmental images. Background Technology
[0002] Action generation refers to the process by which a visual-language-action model, after receiving input information (such as visual images, language instructions, environmental state data, etc.), outputs specific actions that conform to the task objectives based on its learned strategies.
[0003] In the fintech field, examples of action generation applications include intelligent trading systems generating buy, sell, or position adjustment instructions based on real-time market data (such as stock price fluctuations and changes in trading volume) and preset trading strategies.
[0004] In the healthcare field, motion generation can manifest as surgical robots generating precise instrument operation actions (such as the force and path of cutting and suturing) based on medical images (such as CT and MRI images) and surgical planning instructions, as well as intelligent monitoring devices generating response actions such as alarm prompts and medication dosage suggestions based on patient vital sign data (such as changes in heart rate and blood oxygen saturation).
[0005] However, existing action generation methods have significant shortcomings. In the fintech field, models can be used to generate automated trading decisions, but when faced with sudden changes in the market environment, the strategies become rigid because the models only fit historical expert data and cannot adapt to new market paradigms. In the healthcare field, action generation can be applied to surgical robots to perform precise operations based on medical images and preoperative planning, but it is also difficult to handle sudden intraoperative situations due to reliance on limited expert demonstrations. Furthermore, pure reinforcement learning methods suffer from unstable training and difficulty in convergence due to sparse reward signals.
[0006] Therefore, existing technologies for action generation suffer from problems such as weak generalization ability, insufficient utilization of expert data, sparse rewards, and unstable training, which in turn lead to insufficient accuracy of the model when generating machine actions based on environmental images. Summary of the Invention
[0007] This invention provides a method, apparatus, device, and medium for generating machine actions based on environmental images, in order to solve the technical problem of insufficient accuracy when generating machine actions using action generation models.
[0008] Firstly, a method for generating machine actions based on environmental images is provided, including: Acquire the machine trajectory data of the target machine, and use a preset action generation model to generate the behavioral actions corresponding to the environmental observation images in the machine trajectory data; Calculate the behavior cloning loss between the behavior action and the standard behavior action in the machine trajectory data, and optimize the action generation model based on the behavior cloning loss to obtain a preliminary optimized action generation model; Acquire an initial environment image in the target scene, and use the preliminary optimized action generation model to generate machine execution actions corresponding to the initial environment image; Real-time acquisition of updated environmental images corresponding to the actions performed by the machine; calculation of reward signals based on the updated environmental images and the real environmental images in the machine trajectory data. The interaction trajectory data of the target machine is determined based on the machine's actions and the reward signal, and the strategy loss, value loss, and behavior loss are calculated based on the interaction trajectory data and the machine trajectory data. The preliminary optimized action generation model is optimized based on the policy loss, the value loss, and the behavior loss to obtain the optimized action generation model. Acquire an actual environment image in the target scene, and use the optimized action generation model to generate machine actions corresponding to the actual environment image.
[0009] Secondly, a machine motion generation device based on environmental images is provided, comprising: The behavior and action generation module is used to acquire the machine trajectory data of the target machine and generate the behavior and action corresponding to the environmental observation image in the machine trajectory data using a preset action generation model. The preliminary optimization module for the action generation model is used to calculate the behavior cloning loss between the action and the standard action in the machine trajectory data, and to optimize the action generation model based on the behavior cloning loss to obtain the preliminary optimized action generation model. The machine action determination module is used to acquire an initial environment image in the target scene and generate machine actions corresponding to the initial environment image using the preliminary optimized action generation model. The reward signal calculation module is used to acquire the updated environmental image corresponding to the machine's action in real time, and calculate the reward signal based on the updated environmental image and the real environmental image in the machine trajectory data. The interaction trajectory data determination module is used to determine the interaction trajectory data of the target machine based on the machine's executed actions and the reward signal, and to calculate the strategy loss, value loss, and behavior loss based on the interaction trajectory data and the machine trajectory data. An optimized action generation model optimization module is used to optimize the preliminary optimized action generation model based on the policy loss, the value loss, and the behavior loss to obtain an optimized action generation model. The machine motion generation module is used to acquire actual environment images in the target scene and generate machine motions corresponding to the actual environment images using the optimized motion generation model.
[0010] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method for generating machine actions based on environmental images.
[0011] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described method for generating machine actions based on environmental images.
[0012] In the aforementioned scheme implemented by the machine action generation method, apparatus, device, and medium based on environmental images, machine trajectory data of the target machine can be obtained through a client. A preset action generation model is used to generate behavioral actions corresponding to the environmental observation images in the machine trajectory data. The behavioral cloning loss between the behavioral action and the standard behavioral action in the machine trajectory data is calculated, and the action generation model is optimized based on the behavioral cloning loss to obtain a preliminary optimized action generation model. An initial environmental image in the target scene is obtained, and the machine execution action corresponding to the initial environmental image is generated using the preliminary optimized action generation model. An updated environmental image corresponding to the machine execution action is obtained in real time, and a reward signal is calculated based on the updated environmental image and the real environmental image in the machine trajectory data. The interaction trajectory data of the target machine is determined based on the machine execution action and the reward signal, and a strategy loss, value loss, and behavioral loss are calculated based on the interaction trajectory data and the machine trajectory data. The preliminary optimized action generation model is optimized based on the strategy loss, the value loss, and the behavioral loss to obtain an optimized action generation model. Finally, an actual environmental image in the target scene is obtained, and the machine action corresponding to the actual environmental image is generated using the optimized action generation model. By feeding machine actions back to the client, this invention utilizes expert trajectory data to quickly establish a reliable foundation for action generation through behavior cloning, reducing the blindness of initial exploration. Subsequently, in the online reinforcement learning phase, dynamically decaying behavior cloning constraints ensure that the policy gradually enhances its autonomous exploration capabilities while learning stably. Meanwhile, the adaptive hybrid reward model provides dense guidance through multi-dimensional reward signals (perceptual alignment, grasping continuity, task completion), effectively solving the sparse reward problem. This enables the model to adapt to environmental changes more efficiently and optimize decisions, ultimately achieving more accurate machine actions in real-world scenarios and improving the accuracy of the action generation model when generating machine actions based on environmental images. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a schematic diagram of an application environment for a machine motion generation method based on environmental images according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a machine motion generation method based on environmental images in one embodiment of the present invention; Figure 3 yes Figure 2 A flowchart illustrating a specific implementation method of step S1; Figure 4 yes Figure 2 A flowchart illustrating a specific implementation of step S5; Figure 5 This is a schematic diagram of a machine motion generation device based on environmental images in one embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 7 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] The machine motion generation method based on environmental images provided in this invention can be applied to, for example... Figure 1In this application environment, the client communicates with the server via a network. The server can obtain machine trajectory data of the target machine through the client, generate behavioral actions corresponding to the environmental observation images in the machine trajectory data using a preset action generation model; calculate the behavioral cloning loss between the behavioral actions and the standard behavioral actions in the machine trajectory data, and optimize the action generation model based on the behavioral cloning loss to obtain a preliminary optimized action generation model; acquire an initial environmental image in the target scene, and generate machine execution actions corresponding to the initial environmental image using the preliminary optimized action generation model; acquire updated environmental images corresponding to the machine execution actions in real time, and calculate a reward signal based on the updated environmental image and the real environmental image in the machine trajectory data; determine the interaction trajectory data of the target machine based on the machine execution actions and the reward signal, and calculate policy loss, value loss, and behavioral loss based on the interaction trajectory data and the machine trajectory data; optimize the preliminary optimized action generation model based on the policy loss, the value loss, and the behavioral loss to obtain an optimized action generation model; acquire actual environmental images in the target scene, and generate machine actions corresponding to the actual environmental images using the optimized action generation model. This invention feeds machine actions back to the client. It utilizes expert trajectory data to quickly establish a reliable foundation for action generation through behavior cloning, reducing the blindness of initial exploration. Subsequently, in the online reinforcement learning phase, dynamically decaying behavior cloning constraints ensure that the policy gradually enhances its autonomous exploration capabilities while learning stably. An adaptive hybrid reward model provides dense guidance through multi-dimensional reward signals (perceptual alignment, grasping continuity, task completion), effectively solving the sparse reward problem. This allows the model to adapt to environmental changes more efficiently and optimize decisions, ultimately generating more accurate machine actions in real-world scenarios and improving the accuracy of the action generation model when generating machine actions based on environmental images. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention is described in detail below through specific embodiments.
[0017] Please see Figure 2 As shown, Figure 2 A flowchart illustrating a machine motion generation method based on environmental images provided in an embodiment of the present invention includes the following steps: S1. Obtain the machine trajectory data of the target machine, and use a preset action generation model to generate the behavior action corresponding to the environmental observation image in the machine trajectory data.
[0018] In this embodiment of the invention, the machine trajectory data of the target machine is a kind of expert trajectory data, that is, a multimodal data set related to the task recorded by the target machine (such as a robotic arm, intelligent execution device, etc.) in the process of completing a specific task. This data includes environmental observation images (such as real-time images of the machine operation scene observed by a camera, clearly showing the operation object, target area and the state of the machine itself), corresponding task instructions (such as grabbing an object and placing it in a designated position), and standard behavioral actions (such as the motion parameters of each execution component of the machine), and the three types of data are strictly aligned in time sequence.
[0019] In detail, the preset action generation model is a reinforcement learning policy model based on a vision-language-action (VLA) architecture, consisting of a shared OpenVLA backbone network, a policy head for action decision-making, and a value head for state value evaluation. The actions are specific execution instructions or motion parameters generated for the target machine to complete a specific task, based on the input environmental observation image and task instructions.
[0020] In the embodiments of the present invention, see Figure 3 As shown, the step of generating behavioral actions corresponding to environmental observation images in the machine trajectory data using a preset action generation model includes: S31. Use the convolutional layer in the preset action generation model to extract the visual feature vector of the environmental observation image in the machine trajectory data; S32. Determine the spatial association feature vector of the target machine based on the visual feature vector; S33. Concatenate the language feature vector corresponding to the task instruction in the machine trajectory data with the spatial association feature vector to obtain a joint feature vector; S34. Map the joint feature vector to the behavior of the target machine.
[0021] In detail, the preprocessed environmental observation image is input into the convolutional layer. The convolutional layer extracts local features of the image through a sliding window. After each convolution operation, nonlinearity is introduced through the ReLU activation function. At the same time, the feature map is downsampled through the max pooling layer to reduce the number of parameters and retain key features. After multiple rounds of convolution and pooling operations, the final two-dimensional feature map is flattened into a one-dimensional vector, which is the visual feature vector of the environmental observation image.
[0022] Specifically, the position information of the target machine in the environmental observation image (such as the coordinate range of the robotic arm in the image) is extracted and converted into spatial coordinates in the machine coordinate system. Then, the visual feature vector and the machine spatial coordinate information are input into a preset spatial association layer (composed of a fully connected layer and an attention mechanism). The attention mechanism highlights the visual features related to the machine's actions by calculating the association weights between each element in the visual features and the machine's spatial coordinates. Then, the fully connected layer converts the associated features into a fixed-dimensional vector, and finally obtains the spatial association feature vector between the visual features and the target machine.
[0023] Furthermore, the text encoding module in the model is invoked to convert the task instructions (such as "grab the red square") in the machine trajectory data into fixed-dimensional language feature vectors. Then, the language feature vectors and spatial correlation feature vectors are concatenated along the feature dimension to obtain a joint feature vector.
[0024] Next, the joint feature vector is input into the policy head of the action generation model. The joint features are then subjected to dimensionality transformation and nonlinear transformation through a fully connected layer. The output vector of the final fully connected layer is then input into the output layer. If the action is a continuous value (such as a joint angle), the output value is mapped to the physical range of the machine action through the Sigmoid or Tanh function. If it is a discrete action, the action probability distribution is calculated through the Softmax function, and the action with the highest probability is selected as the output, thus obtaining the target machine's action.
[0025] In this embodiment of the invention, relying on expert-level machine trajectory data, the pre-trained model can quickly establish a mapping relationship between environment, instructions, and actions, avoiding the low training efficiency caused by the model learning from scratch.
[0026] In the fintech field, acquiring historical trading trajectory data from intelligent trading systems and using preset action generation models to generate trading actions corresponding to market data images allows the models to quickly learn the operational logic of experienced traders and have the ability to generate reasonable trading actions based on market conditions.
[0027] In the field of healthcare, acquiring surgical trajectory data of surgical robots (including medical images of the surgical site, surgical instructions such as "remove a lesion with a diameter of 2cm", standard operating actions such as the movement path of the robotic arm and cutting force parameters), and using preset models to generate surgical actions corresponding to the medical images, allows the model to initially grasp the action logic of basic surgical operations.
[0028] S2. Calculate the behavior cloning loss between the behavior action and the standard behavior action in the machine trajectory data, and optimize the action generation model based on the behavior cloning loss to obtain a preliminary optimized action generation model.
[0029] In this embodiment of the invention, behavior cloning loss is an indicator used to quantify the difference between the behavior generated by the preset action generation model and the standard behavior in the machine trajectory data. Its core function is to measure the model's accuracy in imitating expert actions, and it is usually calculated using a loss function suitable for continuous or discrete actions.
[0030] In this embodiment of the invention, optimizing the action generation model based on the behavior cloning loss to obtain a preliminary optimized action generation model includes: If the behavior cloning loss is greater than a preset loss threshold, then the gradient of the weight parameters in the action generation model is calculated based on the behavior cloning loss. The weight parameters in the action generation model are updated according to the gradient, and an updated action generation model is generated according to the updated weight parameters. The update action generation model is used to generate update actions corresponding to the environmental observation image; Calculate the update behavior clone loss between the update behavior action and the standard behavior action until the update behavior clone loss is less than or equal to a preset loss threshold, and use the update action generation model corresponding to the update behavior clone loss as the initial optimized action generation model. If the behavior cloning loss is less than or equal to a preset loss threshold, then the action generation model corresponding to the behavior cloning loss is determined as the preliminary optimized action generation model.
[0031] In detail, a preset loss threshold is obtained (e.g., the preset loss threshold is 0.05 in the robotic arm action generation scenario). The calculated behavior clone loss is compared with this threshold. If the loss is greater than the threshold, gradient calculation is performed. Backpropagation is performed on the weight parameters in the preset action generation model (including the convolutional layer weights of the OpenVLA backbone network and the fully connected layer weights of the policy head) to calculate the partial derivative (i.e., gradient) of each weight parameter with respect to the loss value.
[0032] Specifically, obtain the learning rate (e.g., 0.001) for the pre-set parameter update, and then update each weight parameter according to the formula "new parameter value = old parameter value - learning rate × gradient of the parameter". After all weight parameters are updated, reload the updated parameters into the corresponding module of the original action generation model to obtain the updated action generation model.
[0033] Next, the environmental observation images and task instructions are re-input into the update action generation model. The model performs forward propagation according to the same process as the initial generated action, and finally generates an updated action that differs from the original action.
[0034] Furthermore, the update behavior cloning loss between the update action and the standard action in the machine trajectory data is calculated. Then, this loss is compared with a preset loss threshold. If it is still greater than the threshold, the gradient is calculated, the parameters are updated, the update action is generated, and the update loss is calculated repeatedly. The change trend of the loss value is recorded after each round of the loop until the update behavior cloning loss calculated in a certain round is less than or equal to the preset loss threshold. At this time, the iteration stops, and the current update action generation model is determined as the preliminary optimized action generation model.
[0035] Next, if the loss is directly less than or equal to the preset loss threshold, the action generation model has a good ability to imitate expert actions, and a preliminary optimized action generation model is obtained.
[0036] In this embodiment of the invention, by quantifying the differences in actions and adjusting the model parameters accordingly, the model can accurately learn expert-level action logic, significantly reducing the error in action generation. Through multiple rounds of parameter adjustment, it can stably generate behavioral actions that are less different from expert standard actions, thus meeting the optimization goals of the offline behavior cloning stage.
[0037] In the fintech field, the mean squared error loss of a trading action (such as selling 500 shares) generated by an intelligent trading model is compared with that of a standard trading action (such as selling 1,000 shares) in the historical trajectory. If the loss exceeds the standard, the parameters of the market feature extraction layer and the trading decision layer in the model are updated through gradient descent until the loss reaches the standard, thus obtaining a preliminary optimized trading action generation model.
[0038] In the medical and health field, the cross-entropy loss between the instrument movement angle generated by the surgical robot model and the standard surgical action angle is calculated. The model's ability to extract medical image features is optimized by updating parameters until the loss reaches the target. This yields a preliminary optimized surgical action model, which allows the surgical actions generated by the model to better match the operation of expert doctors and reduce the risk of surgical operation deviation.
[0039] S3. Obtain the initial environment image in the target scene, and use the preliminary optimized action generation model to generate the machine execution action corresponding to the initial environment image.
[0040] In this embodiment of the invention, the initial environment image refers to the image data of the target scene (such as the physical scene of robotic arm operation, the real-time market scene of financial transactions, or the surgical field scene of medical surgery) in its initial state during the online reinforcement learning phase.
[0041] In detail, machine action execution refers to the initial optimization of the action generation model in the online reinforcement learning phase, which generates actions to drive the target machine to interact with the environment based on the initial environmental image of the target scene and real-time task instructions.
[0042] Specifically, the initial environmental image of the target scene is acquired in real time by the visual sensor of the target machine (such as the camera at the end of the robotic arm), and the real-time task instructions in the online stage (such as "grab the part at the initial position to the right tray") are obtained. The text encoding module of the model is converted into language feature vectors. The preprocessed initial environmental image and language feature vectors are input into the preliminary optimization action generation model. The model generates action parameters that match the initial environment. The generated action parameters are format-converted and converted into control signals that can be recognized by the actuator of the target machine to obtain the machine's execution action.
[0043] In this embodiment of the invention, based on the model optimized in the offline stage, interactive actions adapted to the initial state of the online scene are quickly generated, avoiding invalid or dangerous actions (such as the robotic arm colliding with obstacles) generated in the online stage due to the model's unfamiliarity with the scene.
[0044] In the fintech field, the initial market data (including opening price and opening volume) at the start of stock trading can be obtained. The first trading action after the market opens can be generated using a preliminarily optimized trading action model. This action can avoid losses caused by blind trading due to the model's unfamiliarity with real-time market data.
[0045] In the medical and health field, the initial surgical field image (including the surgical site and the initial position of instruments) is obtained at the start of the operation. The first surgical action (such as "moving the scalpel to 1cm from the edge of the lesion") is generated using a preliminarily optimized surgical action model. This action can ensure that the initial operation of the operation conforms to the expert standard and reduce the surgical risk.
[0046] S4. Real-time acquisition of the updated environmental image corresponding to the machine's action, and calculation of the reward signal based on the updated environmental image and the real environmental image in the machine trajectory data.
[0047] In this embodiment of the invention, the reward signal is a quantitative indicator used in the online reinforcement learning stage to evaluate the contribution of the machine's actions to the completion of the task. It is calculated based on the updated environment image after the machine performs the action and the real environment image in the machine trajectory data, and includes accuracy reward, incentive reward and task completion reward.
[0048] In this embodiment of the invention, calculating the reward signal based on the real environment image in the updated environment image and machine trajectory data includes: Calculate the visual similarity between the updated environment image and the real environment image, and determine the accuracy reward based on the visual similarity; The continuity of the machine's actions is determined based on the updated environment image and the real environment image, and an incentive reward is determined based on the continuity. The degree of completion of the machine's action task is determined based on the updated environment image and the real environment image, and the task completion reward is determined based on the degree of completion of the action task. The accuracy reward, the incentive reward, and the task completion reward are weighted and merged into a reward signal.
[0049] In detail, the real environment image corresponding to the updated environment image is extracted from the machine trajectory data (e.g., after the machine performs the "grabbing action", the corresponding environment image in the expert data is "grabbing action completed"). The two images are grayscaled, and the visual similarity between them is calculated using the structural similarity index algorithm. The accuracy reward is obtained according to the calculation rules of the set accuracy reward (e.g., accuracy reward = visual similarity × preset reward coefficient).
[0050] Specifically, the interaction state between the target machine and the operated object is identified from the updated environmental image. The positions of the key components of the machine and the operated object are located by the target detection algorithm, and the distance between them is calculated. If the distance is less than a preset threshold (e.g., 5mm) and the state is maintained for N consecutive frames (e.g., 3 frames) (indicating stable interaction), the continuity of the machine's actions is determined to be high. If the distance fluctuates greatly or only a single frame meets the threshold requirement, the continuity is low. According to the set incentive reward rules: high continuity corresponds to a reward value (e.g., 3), medium continuity corresponds to a medium reward value (e.g., 1), and low continuity corresponds to 0 reward. The incentive reward is the accumulation of rewards for multiple consecutive actions.
[0051] Furthermore, the criteria for judging task completion are obtained (e.g., the distance between the center of the part and the center of the target area in a robotic arm task is less than 2mm). Then, task completion-related features are extracted from the updated environment image, the task completion rate is calculated, and according to the set task completion reward rules, a high reward (e.g., 20) is given if the completion rate is 100%; a medium reward (e.g., 5-15, increasing linearly with the completion rate) is given if the completion rate is 50%-99%; and a reward of 0 is given if the completion rate is less than 50%.
[0052] Furthermore, the total reward signal is calculated according to the formula: Reward Signal = Accuracy Reward × Reward Weight 1 + Incentive Reward × Reward Weight 2 + Task Completion Reward × Reward Weight 3. For example, if the accuracy reward is 4 (weight 0.4), the incentive reward is 3 (weight 0.4), and the task completion reward is 0 (weight 0.2), then the reward signal = 4 × 0.4 + 3 × 0.4 + 0 × 0.2 = 2.8. After the calculation is completed, the reward signal is associated with and stored in relation to the corresponding machine action and the updated environment image.
[0053] In this embodiment of the invention, by dynamically weighting and fusing three types of complementary rewards, dense and accurate feedback is provided for the machine to perform actions. This solves the problem of sparse rewards in traditional reinforcement learning, which relies solely on final rewards and results in unguided intermediate steps. At the same time, rewards are calculated based on real-world images to ensure that the reward signal is consistent with the logic of the expert task, guiding the model to optimize towards expert-level action strategies and improving the stability and efficiency of online training.
[0054] In the fintech field, after an intelligent trading device executes the action of "buying 1,000 shares", it acquires the updated market data image and calculates the visual similarity with the "expected market data image after purchase" in the expert trajectory to obtain an accuracy reward; it calculates the continuity based on "whether the stock is held continuously and the stock price has not fallen significantly" to obtain an incentive reward; and it calculates the task completion reward based on "whether the preset profit target has been achieved". After weighted fusion, a reward signal is obtained to guide the model to optimize trading timing and position control.
[0055] In the field of healthcare, after a surgical robot performs a "lesion cutting" action, it acquires an updated surgical field image (including the edge of the lesion after cutting and the extent of tissue bleeding). The similarity between this image and the standard surgical field image after cutting by an expert surgeon is calculated to obtain an accuracy reward. The continuity of the cutting action is calculated based on whether it is continuous and without interruption, and an incentive reward is obtained based on whether the lesion is completely removed. After weighted fusion, the model is guided to optimize the cutting path and force to improve the surgical outcome.
[0056] S5. Determine the interaction trajectory data of the target machine based on the machine's executed actions and the reward signal, and calculate the strategy loss, value loss, and behavior loss based on the interaction trajectory data and the machine trajectory data.
[0057] In this embodiment of the invention, the interaction trajectory data refers to the time-series data set of the online reinforcement learning stage, including the environmental state of each interaction step (initial environmental image, updated environmental image, etc.), machine execution actions (actions corresponding to each environmental state step), reward signals (feedback rewards after each action step is executed), and the next environmental state (updated environmental image generated after the action is executed), and the data is strictly arranged in the order of interaction time.
[0058] In detail, based on the current environmental state (image input and task instructions), a machine action is predicted and output. After executing the action, the state is updated and a new image observation is returned. The calculated reward signal is obtained, and the state, action, and reward value at each time step are continuously recorded to form a complete interaction trajectory sequence τ=(l, o1, a1, r1, o2, a2, r2, ..., o t ,a t ,r t), where l represents the task instruction, o represents the environmental observation image, a is the policy output, and r is the mixed reward signal.
[0059] In the embodiments of the present invention, see Figure 4 As shown, the calculation of policy loss, value loss, and behavior loss based on the interaction trajectory data and the machine trajectory data includes: S41. Calculate the policy loss of the preliminary optimized action generation model based on the environmental image, machine action, and reward signal in the interaction trajectory data; S42. Calculate the value loss of the preliminary optimized action generation model based on the task instructions, environmental images and reward signals in the interaction trajectory data; S43. Calculate the behavioral loss of the preliminary optimized action generation model based on the interaction trajectory data and the machine trajectory data.
[0060] In detail, the environmental state, machine action, and corresponding reward signal for each step are extracted from the interaction trajectory data. Based on the policy network parameters (i.e., policy head parameters) of the current action generation model and the old policy parameters optimized in the previous round, the probability of generating the machine action (denoted as P_new) is calculated through the current policy network, and the probability of generating the same action (denoted as P_old) is calculated through the old policy parameters. The probability ratio r = P_new / P_old is calculated. Then, according to the pruning rule of the PPO algorithm, r is pruned to the range of [1-ε, 1+ε] (ε is usually set to 0.2) to obtain the pruned ratio r_clip. At the same time, the advantage function value A is calculated (through the temporal difference method). The policy loss is obtained by calculating the average of min(r×A, r_clip×A) and taking the negative (minimizing the loss corresponds to the optimal policy).
[0061] Specifically, the environmental state and corresponding reward signal for each step are extracted from the interaction trajectory data. Based on the model value network (including value head parameters), the current environmental state is forward propagated through the value network to output the value prediction value of the state. Then, the target value is calculated (using the temporal difference algorithm), and the value loss is calculated using the mean squared error (MSE) loss function.
[0062] Furthermore, the machine-executed actions for each step are extracted from the interaction trajectory data, and the standard behavioral actions corresponding to the interaction steps are extracted from the machine trajectory data to ensure that the two types of actions are aligned in terms of steps and dimensions. Then, the same loss function as in the offline stage is used (e.g., MSE for continuous actions and cross-entropy for discrete actions) to calculate the difference between the machine-executed actions and the standard behavioral actions. At the same time, the weight of the behavioral loss is adjusted according to the number of online training steps using a preset scheduling function (e.g., linear decay function) (e.g., weight 0.5 in the early stage of training and weight 0.1 in the later stage) to obtain the weighted behavioral loss in the online stage.
[0063] In this embodiment of the invention, the interaction trajectory data fully records the interaction process between the model and the environment, providing comprehensive data support for loss calculation; the policy loss can constrain the stability of policy updates, the value loss can improve the accuracy of value prediction, and the online behavior loss can combine expert data to constrain the policy, balancing exploration and stability. The three work together to ensure the efficiency and safety of model optimization in the online stage.
[0064] In the fintech field, based on the interaction trajectory data of intelligent trading, strategy loss is calculated to optimize trading strategies, value loss is calculated to improve market value prediction, and behavioral loss is calculated to constrain trading actions to avoid deviating from expert trading logic. These three types of losses work together to optimize the trading model.
[0065] In the field of healthcare, based on the interaction trajectory data of surgical robots, the strategy loss is calculated to optimize the cutting strategy, the value loss is calculated to improve the prediction of the surgical field state value, and the behavior loss is calculated to constrain actions to not deviate from expert surgical standards, thereby improving the accuracy of surgical models.
[0066] S6. Optimize the preliminary optimized action generation model based on the strategy loss, the value loss, and the behavior loss to obtain the optimized action generation model.
[0067] In this embodiment of the invention, the optimized action generation model refers to the VLA model obtained by controlling the policy loss, value loss, and behavior loss within a preset range through multiple rounds of optimization and iteration during the online reinforcement learning stage, and finally obtaining a model with high generalization ability and environmental adaptability.
[0068] In this embodiment of the invention, optimizing the preliminary optimized action generation model based on the policy loss, the value loss, and the behavioral loss to obtain the optimized action generation model includes: Update the policy parameters of the policy head in the preliminary optimized action generation model based on the behavioral loss and the policy loss; Update the core parameters in the backbone network of the preliminary optimized action generation model based on the policy loss and the value loss; Update the value parameters of the value heads in the preliminary optimized network based on the value loss; The preliminary optimized action generation model is updated based on the strategy parameters, the core parameters, and the value parameters to obtain the updated action generation model. The updated action generation model is used as a preliminary optimized action generation model, and the process returns to the step of generating machine-executed actions corresponding to the initial environment image using the preliminary optimized action generation model; The number of iterations of the initial optimized action generation model is accumulated until the number of iterations reaches a preset iteration threshold. The updated action generation model at this point is then determined as the optimized action generation model.
[0069] In detail, the behavior loss and policy loss in the online phase are weighted and summed according to preset weights (e.g., behavior loss weight 0.3, policy loss weight 0.7) to obtain the total optimization loss of the policy head. The gradient of the total loss with respect to the policy head parameters (e.g., fully connected layer weights, biases) is calculated, and the gradient descent algorithm is used to update each policy head parameter.
[0070] Next, the policy loss and value loss are weighted and summed according to preset weights (e.g., policy loss 0.4, value loss 0.6) to obtain the total optimization loss of the backbone network. The gradient of the total loss with respect to these core parameters is calculated, and the parameters are updated using a small learning rate. The core parameters refer to the weight parameters and bias parameters of each layer in the backbone network (e.g., the weights and biases of each convolutional kernel).
[0071] Specifically, with value loss as the sole optimization objective, the gradient of value loss with respect to value head parameters (such as fully connected layer weights and activation function parameters) is calculated, a learning rate (such as 0.001) is set, and the parameters are updated using the gradient descent algorithm.
[0072] Next, the updated policy head parameters, backbone network core parameters, and value head parameters are reintegrated into the corresponding modules of the preliminary optimized action generation model, replacing the old parameters of the original modules: the updated convolutional layer and text encoding module parameters are loaded into the OpenVLA backbone network, the updated policy parameters are loaded into the policy head, and the updated value parameters are loaded into the value head, resulting in the updated action generation model.
[0073] Furthermore, the updated action generation model is relabeled as a new preliminary optimized action generation model, overwriting the parameter storage path of the original preliminary optimized model. Then, the target machine's vision sensor is invoked to re-acquire the environmental image of the current target scene. The new environmental image and real-time task instructions are input into the new preliminary optimized action generation model, and the machine's actions are repeatedly generated, updated environmental images are acquired, reward signals are calculated, interaction trajectory data is determined, and the online iterative process of calculating three types of losses and parameter updates is performed. After each iteration, the model parameters are continuously optimized, and the interaction trajectory data is continuously accumulated until the preset iteration stopping condition is met.
[0074] Next, after each round of model optimization is completed, the current iteration number is recorded by a counter and compared with a preset iteration number threshold (e.g., 1000 times, set according to task complexity). If the current iteration number is less than the threshold, the above iteration process continues. If the iteration number reaches the threshold, the iteration stops. At this time, the action generation model updated in the last round is read from the parameter storage path and the model is determined as the final optimized action generation model.
[0075] In this embodiment of the invention, by updating parameters in a modular manner (strategy head, backbone network, value head), the performance of each module of the model can be accurately improved, avoiding the "one-size-fits-all" problem of parameter updates; multi-round iterative optimization combined with iteration number control can ensure that the model fully learns the environmental adaptability in interactive data, while preventing overfitting caused by overtraining.
[0076] In the fintech field, based on trading strategy loss, value loss, and behavioral loss, the strategy head, backbone network, and value head of the intelligent trading model are updated. After 1,000 iterations, the model can generate stable and profitable trading actions in different market conditions.
[0077] In the field of healthcare, based on surgical strategy loss, value loss, and behavioral loss, the strategy head, backbone network, and value head of the surgical robot model are updated. After 5,000 iterations, the model can perform precise surgery in scenarios with individual anatomical differences among different patients.
[0078] S7. Obtain the actual environment image in the target scene, and use the optimized action generation model to generate the machine action corresponding to the actual environment image.
[0079] In this embodiment of the invention, the actual environment image refers to the image data that the target machine collects in real time through a visual sensor in a real application scenario (non-training / testing scenario, such as an industrial production workshop, a financial trading site, or a hospital operating room), reflecting the real state of the scenario.
[0080] In detail, machine motion refers to the final motion generated by an optimized motion generation model based on actual environmental images and actual task instructions, which can be directly used for actual task execution by the target machine.
[0081] In this embodiment of the invention, generating machine actions corresponding to the actual environment image using the optimized action generation model includes: Obtain the actual task instructions in the target scene and convert the actual task instructions into a task feature vector; The visual information vector in the actual environment image is extracted using the convolutional layer in the optimized action generation model. The spatial structure vector of the target machine is determined based on the visual information vector. The task feature vector and the spatial structure vector are concatenated to form the fusion information vector of the target machine; The fused information vector is nonlinearly transformed into the machine actions of the target machine.
[0082] In detail, the actual task instructions input by the user are received through a human-computer interaction interface (such as an industrial control screen or a medical control console). The instructions are preprocessed into text. The preprocessed actual task instructions are input into the text encoding module and converted into a fixed-dimensional task feature vector through the calculation of the embedding layer and the self-attention layer.
[0083] Next, real-time images of the actual environment are acquired, and the preprocessed images are input into the convolutional layer of the optimized action generation model. The convolutional layer performs feature extraction according to the pre-trained parameter configuration. After each convolution, the nonlinearity is enhanced by the ReLU activation function, and downsampling is performed through the max pooling layer. After multiple rounds of convolution-pooling operations, a visual information vector is obtained.
[0084] Specifically, the key components of the target machine are identified from the actual environment image, the pixel coordinates of the components in the image are obtained, the pixel coordinates are converted into spatial coordinates in the machine coordinate system, features related to the operation object are extracted from the visual information vector, the spatial coordinates of the operation object in the machine coordinate system are calculated, and then the visual information vector, the spatial coordinates of the key machine components, and the spatial coordinates of the operation object are input into the spatial association layer of the model. The spatial coordinates are converted into feature vectors through the fully connected layer, and then concatenated with the visual information vector to obtain the spatial structure vector.
[0085] Furthermore, the task feature vector and spatial structure vector are concatenated into a fused information vector. This fused information vector is input into the policy head of the optimized action generation model. First, it undergoes dimensionality transformation and nonlinear transformation through three fully connected layers. Then, the output is processed according to the type of machine action: for continuous actions (such as joint angles), the Tanh function maps the output value to the physical range of the machine action; for discrete actions, the Softmax function calculates the action probability distribution, selects the action type with the highest probability, and combines it with the task instructions to determine the action parameters. Finally, the processed action parameters are converted into control signals recognizable by the target machine's actuator to obtain the final machine action.
[0086] In this embodiment of the invention, the trained and optimized model is directly applied to the real-time processing of real-world images and the accurate generation of machine actions, which can ensure that the machine can stably complete tasks in real and complex environments.
[0087] In the fintech field, smart trading terminals use cameras to capture images of actual stock market data, input them into optimized trading models, and the models generate machine actions such as "selling 1,000 shares when the stock price reaches 45 yuan".
[0088] In the field of healthcare, surgical robots use endoscopes to collect actual surgical field images of the patient's surgical site (including variations in blood vessel distribution and lesion location), input them into an optimized surgical model, and the model generates a machine action of "cutting along the edge of the lesion at 0.5cm with a cutting force of 20N", which is converted into instrument-driven signals and executed to precisely remove the lesion.
[0089] As can be seen, in the above scheme, the model can quickly learn expert logic by optimizing the loss based on machine trajectory data and behavior cloning, thus solving the problems of weak generalization and low utilization of expert data. In the online stage, the reward is calculated based on real environment images to alleviate the problems of sparse rewards and unstable training in reinforcement learning. The model is optimized in modules based on the calculation of three types of losses based on interaction trajectories to balance stability and exploration. Finally, the optimized model is used to generate actual actions, reducing the dependence on a large number of interaction samples, saving resources, and being universally adaptable to multiple scenarios, thereby improving the accuracy of the action generation model when generating machine actions based on environment images.
[0090] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0091] In one embodiment, a machine motion generation apparatus based on environmental images is provided, which corresponds one-to-one with the machine motion generation method based on environmental images in the above embodiments. For example... Figure 5 As shown, the machine motion generation device 100 based on environmental images includes a behavior motion generation module 101, a preliminary optimized motion generation model optimization module 102, a machine execution motion determination module 103, a reward signal calculation module 104, an interaction trajectory data determination module 105, an optimized motion generation model optimization module 106, and a machine motion generation module 107. Detailed descriptions of each functional module are as follows: The behavior and action generation module 101 is used to acquire the machine trajectory data of the target machine and generate the behavior and action corresponding to the environmental observation image in the machine trajectory data using a preset action generation model. The preliminary optimization module 102 for the action generation model is used to calculate the behavior cloning loss between the action and the standard action in the machine trajectory data, and optimize the action generation model based on the behavior cloning loss to obtain the preliminary optimized action generation model. The machine action determination module 103 is used to acquire an initial environment image in the target scene and generate machine actions corresponding to the initial environment image using the preliminary optimized action generation model. The reward signal calculation module 104 is used to acquire the updated environmental image corresponding to the machine's action in real time, and calculate the reward signal based on the updated environmental image and the real environmental image in the machine trajectory data. The interaction trajectory data determination module 105 is used to determine the interaction trajectory data of the target machine based on the machine's executed actions and the reward signal, and to calculate the strategy loss, value loss and behavior loss based on the interaction trajectory data and the machine trajectory data. The action generation model optimization module 106 is used to optimize the preliminary optimized action generation model based on the policy loss, the value loss and the behavior loss to obtain the optimized action generation model. The machine motion generation module 107 is used to acquire an actual environment image in the target scene and generate machine motions corresponding to the actual environment image using the optimized motion generation model.
[0092] In one embodiment, the behavior / action generation module 101, when generating behavior / actions corresponding to environmental observation images in the machine trajectory data using a preset action generation model, is configured to: Visual feature vectors of environmental observation images in the machine trajectory data are extracted using convolutional layers in a preset action generation model. Determine the spatial association feature vector of the target machine based on the visual feature vector; The language feature vector corresponding to the task instruction in the machine trajectory data is concatenated with the spatial association feature vector to obtain a joint feature vector; The joint feature vector is mapped to the behavior of the target machine.
[0093] In one embodiment, the preliminary optimization module 102 for the action generation model, when performing optimization of the action generation model based on the behavior cloning loss to obtain a preliminary optimized action generation model, is used to: If the behavior cloning loss is greater than a preset loss threshold, then the gradient of the weight parameters in the action generation model is calculated based on the behavior cloning loss. The weight parameters in the action generation model are updated according to the gradient, and an updated action generation model is generated according to the updated weight parameters. The update action generation model is used to generate update actions corresponding to the environmental observation image; Calculate the update behavior clone loss between the update behavior action and the standard behavior action until the update behavior clone loss is less than or equal to a preset loss threshold, and use the update action generation model corresponding to the update behavior clone loss as the initial optimized action generation model. If the behavior cloning loss is less than or equal to a preset loss threshold, then the action generation model corresponding to the behavior cloning loss is determined as the preliminary optimized action generation model.
[0094] In one embodiment, the reward signal calculation module 104, when performing the calculation of the reward signal based on the real environment image in the updated environment image and machine trajectory data, is used to: Calculate the visual similarity between the updated environment image and the real environment image, and determine the accuracy reward based on the visual similarity; The continuity of the machine's actions is determined based on the updated environment image and the real environment image, and an incentive reward is determined based on the continuity. The degree of completion of the machine's action task is determined based on the updated environment image and the real environment image, and the task completion reward is determined based on the degree of completion of the action task. The accuracy reward, the incentive reward, and the task completion reward are weighted and merged into a reward signal.
[0095] In one embodiment, the interaction trajectory data determination module 105, when performing the calculation of strategy loss, value loss, and behavior loss based on the interaction trajectory data and the machine trajectory data, is configured to: The policy loss of the preliminary optimized action generation model is calculated based on the environmental image, machine-executed actions, and reward signals in the interaction trajectory data. The value loss of the preliminary optimized action generation model is calculated based on the task instructions, environmental images, and reward signals in the interaction trajectory data. The behavioral loss of the preliminary optimized action generation model is calculated based on the interaction trajectory data and the machine trajectory data.
[0096] In one embodiment, the action generation model optimization module 106, when performing the optimization of the preliminary action generation model based on the policy loss, the value loss, and the behavior loss to obtain an optimized action generation model, is used to: Update the policy parameters of the policy head in the preliminary optimized action generation model based on the behavioral loss and the policy loss; Update the core parameters in the backbone network of the preliminary optimized action generation model based on the policy loss and the value loss; Update the value parameters of the value heads in the preliminary optimized network based on the value loss; The preliminary optimized action generation model is updated based on the strategy parameters, the core parameters, and the value parameters to obtain the updated action generation model. The updated action generation model is used as a preliminary optimized action generation model, and the process returns to the step of generating machine-executed actions corresponding to the initial environment image using the preliminary optimized action generation model; The number of iterations of the initial optimized action generation model is accumulated until the number of iterations reaches a preset iteration threshold. The updated action generation model at this point is then determined as the optimized action generation model.
[0097] In one embodiment, the machine motion generation module 107, when generating machine motions corresponding to the actual environment image using the optimized motion generation model, is configured to: Obtain the actual task instructions in the target scene and convert the actual task instructions into a task feature vector; The visual information vector in the actual environment image is extracted using the convolutional layer in the optimized action generation model. The spatial structure vector of the target machine is determined based on the visual information vector. The task feature vector and the spatial structure vector are concatenated to form the fusion information vector of the target machine; The fused information vector is nonlinearly transformed into the machine actions of the target machine.
[0098] This invention provides a machine action generation device based on environmental images. It optimizes the model by combining machine trajectory data with behavior cloning loss, enabling the model to quickly learn expert logic and addressing issues of weak generalization and low utilization of expert data. In the online phase, it calculates rewards based on real-world environmental images, alleviating the problems of sparse rewards and unstable training in reinforcement learning. It calculates three types of losses based on interaction trajectories and optimizes the model in modules, balancing stability and exploratory capabilities. Finally, it uses the optimized model to generate actual actions, reducing reliance on a large number of interaction samples, saving resources, and providing universal applicability to multiple scenarios, thus improving the accuracy of the action generation model in generating machine actions based on environmental images.
[0099] Specific limitations regarding the machine motion generation device based on environmental images can be found in the limitations of the machine motion generation method based on environmental images mentioned above, and will not be repeated here. Each module in the aforementioned machine motion generation device based on environmental images can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0100] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a machine motion generation method based on environmental images on the server side.
[0101] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a machine motion generation method based on environmental images.
[0102] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire the machine trajectory data of the target machine, and use a preset action generation model to generate the behavioral actions corresponding to the environmental observation images in the machine trajectory data; Calculate the behavior cloning loss between the behavior action and the standard behavior action in the machine trajectory data, and optimize the action generation model based on the behavior cloning loss to obtain a preliminary optimized action generation model; Acquire an initial environment image in the target scene, and use the preliminary optimized action generation model to generate machine execution actions corresponding to the initial environment image; Real-time acquisition of updated environmental images corresponding to the actions performed by the machine; calculation of reward signals based on the updated environmental images and the real environmental images in the machine trajectory data. The interaction trajectory data of the target machine is determined based on the machine's actions and the reward signal, and the strategy loss, value loss, and behavior loss are calculated based on the interaction trajectory data and the machine trajectory data. The preliminary optimized action generation model is optimized based on the policy loss, the value loss, and the behavior loss to obtain the optimized action generation model. Acquire an actual environment image in the target scene, and use the optimized action generation model to generate machine actions corresponding to the actual environment image.
[0103] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Acquire the machine trajectory data of the target machine, and use a preset action generation model to generate the behavioral actions corresponding to the environmental observation images in the machine trajectory data; Calculate the behavior cloning loss between the behavior action and the standard behavior action in the machine trajectory data, and optimize the action generation model based on the behavior cloning loss to obtain a preliminary optimized action generation model; Acquire an initial environment image in the target scene, and use the preliminary optimized action generation model to generate machine execution actions corresponding to the initial environment image; Real-time acquisition of updated environmental images corresponding to the actions performed by the machine; calculation of reward signals based on the updated environmental images and the real environmental images in the machine trajectory data. The interaction trajectory data of the target machine is determined based on the machine's actions and the reward signal, and the strategy loss, value loss, and behavior loss are calculated based on the interaction trajectory data and the machine trajectory data. The preliminary optimized action generation model is optimized based on the policy loss, the value loss, and the behavior loss to obtain the optimized action generation model. The actual environment image in the target scene is acquired, and the machine action corresponding to the actual environment image is generated using the optimized action generation model. It should be noted that the functions or steps described above regarding the computer-readable storage medium or computer device can be found in the relevant descriptions on the server side and client side of the foregoing method embodiments; to avoid repetition, they will not be described in detail here.
[0104] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0105] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0106] It should be noted that if any software tools or components not belonging to our company appear in the embodiments of this application, they are merely for illustrative purposes and do not represent actual use.
[0107] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for generating machine actions based on environmental images, characterized in that, include: Acquire the machine trajectory data of the target machine, and use a preset action generation model to generate the behavioral actions corresponding to the environmental observation images in the machine trajectory data; Calculate the behavior cloning loss between the behavior action and the standard behavior action in the machine trajectory data, and optimize the action generation model based on the behavior cloning loss to obtain a preliminary optimized action generation model; Acquire an initial environment image in the target scene, and use the preliminary optimized action generation model to generate machine execution actions corresponding to the initial environment image; Real-time acquisition of updated environmental images corresponding to the actions performed by the machine; calculation of reward signals based on the updated environmental images and the real environmental images in the machine trajectory data. The interaction trajectory data of the target machine is determined based on the machine's actions and the reward signal, and the strategy loss, value loss, and behavior loss are calculated based on the interaction trajectory data and the machine trajectory data. The preliminary optimized action generation model is optimized based on the policy loss, the value loss, and the behavior loss to obtain the optimized action generation model. Acquire an actual environment image in the target scene, and use the optimized action generation model to generate machine actions corresponding to the actual environment image.
2. The machine motion generation method based on environmental images as described in claim 1, characterized in that, The step of generating behavioral actions corresponding to environmental observation images in the machine trajectory data using a preset action generation model includes: Visual feature vectors of environmental observation images in the machine trajectory data are extracted using convolutional layers in a preset action generation model. Determine the spatial association feature vector of the target machine based on the visual feature vector; The language feature vector corresponding to the task instruction in the machine trajectory data is concatenated with the spatial association feature vector to obtain a joint feature vector; The joint feature vector is mapped to the behavior of the target machine.
3. The machine motion generation method based on environmental images as described in claim 1, characterized in that, The step of optimizing the action generation model based on the behavior cloning loss to obtain a preliminary optimized action generation model includes: If the behavior cloning loss is greater than a preset loss threshold, then the gradient of the weight parameters in the action generation model is calculated based on the behavior cloning loss. The weight parameters in the action generation model are updated according to the gradient, and an updated action generation model is generated according to the updated weight parameters. The update action generation model is used to generate update actions corresponding to the environmental observation image; Calculate the update behavior clone loss between the update behavior action and the standard behavior action until the update behavior clone loss is less than or equal to a preset loss threshold, and use the update action generation model corresponding to the update behavior clone loss as the initial optimized action generation model. If the behavior cloning loss is less than or equal to a preset loss threshold, then the action generation model corresponding to the behavior cloning loss is determined as the preliminary optimized action generation model.
4. The machine motion generation method based on environmental images as described in claim 1, characterized in that, The step of calculating the reward signal based on the updated environmental image and the real environmental image in the machine trajectory data includes: Calculate the visual similarity between the updated environment image and the real environment image, and determine the accuracy reward based on the visual similarity; The continuity of the machine's actions is determined based on the updated environment image and the real environment image, and an incentive reward is determined based on the continuity. The degree of completion of the machine's action task is determined based on the updated environment image and the real environment image, and the task completion reward is determined based on the degree of completion of the action task. The accuracy reward, the incentive reward, and the task completion reward are weighted and merged into a reward signal.
5. The machine motion generation method based on environmental images as described in claim 1, characterized in that, The calculation of strategy loss, value loss, and behavior loss based on the interaction trajectory data and the machine trajectory data includes: The policy loss of the preliminary optimized action generation model is calculated based on the environmental image, machine-executed actions, and reward signals in the interaction trajectory data. The value loss of the preliminary optimized action generation model is calculated based on the task instructions, environmental images, and reward signals in the interaction trajectory data. The behavioral loss of the preliminary optimized action generation model is calculated based on the interaction trajectory data and the machine trajectory data.
6. The machine motion generation method based on environmental images as described in claim 1, characterized in that, The step of optimizing the preliminary optimized action generation model based on the policy loss, the value loss, and the behavioral loss to obtain the optimized action generation model includes: Update the policy parameters of the policy head in the preliminary optimized action generation model based on the behavioral loss and the policy loss; Update the core parameters in the backbone network of the preliminary optimized action generation model based on the policy loss and the value loss; Update the value parameters of the value heads in the preliminary optimized network based on the value loss; The preliminary optimized action generation model is updated based on the strategy parameters, the core parameters, and the value parameters to obtain the updated action generation model. The updated action generation model is used as a preliminary optimized action generation model, and the process returns to the step of generating machine-executed actions corresponding to the initial environment image using the preliminary optimized action generation model; The number of iterations of the initial optimized action generation model is accumulated until the number of iterations reaches a preset iteration threshold. The updated action generation model at this point is then determined as the optimized action generation model.
7. The machine motion generation method based on environmental images as described in claim 1, characterized in that, The step of generating machine actions corresponding to the actual environment image using the optimized action generation model includes: Obtain the actual task instructions in the target scene and convert the actual task instructions into a task feature vector; The visual information vector in the actual environment image is extracted using the convolutional layer in the optimized action generation model. The spatial structure vector of the target machine is determined based on the visual information vector. The task feature vector and the spatial structure vector are concatenated to form the fusion information vector of the target machine; The fused information vector is nonlinearly transformed into the machine actions of the target machine.
8. A machine motion generation device based on environmental images, characterized in that, include: The behavior and action generation module is used to acquire the machine trajectory data of the target machine and generate the behavior and action corresponding to the environmental observation image in the machine trajectory data using a preset action generation model. The preliminary optimization module for the action generation model is used to calculate the behavior cloning loss between the action and the standard action in the machine trajectory data, and to optimize the action generation model based on the behavior cloning loss to obtain the preliminary optimized action generation model. The machine action determination module is used to acquire an initial environment image in the target scene and generate machine actions corresponding to the initial environment image using the preliminary optimized action generation model. The reward signal calculation module is used to acquire the updated environmental image corresponding to the machine's action in real time, and calculate the reward signal based on the updated environmental image and the real environmental image in the machine trajectory data. The interaction trajectory data determination module is used to determine the interaction trajectory data of the target machine based on the machine's executed actions and the reward signal, and to calculate the strategy loss, value loss, and behavior loss based on the interaction trajectory data and the machine trajectory data. An optimized action generation model optimization module is used to optimize the preliminary optimized action generation model based on the policy loss, the value loss, and the behavior loss to obtain an optimized action generation model. The machine motion generation module is used to acquire actual environment images in the target scene and generate machine motions corresponding to the actual environment images using the optimized motion generation model.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the machine motion generation method based on environmental images as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the machine motion generation method based on environmental images as described in any one of claims 1 to 7.