Real-time task decision-making method, device and equipment and storage medium thereof

By using a pre-trained VLA model and a progressive learning strategy, structured action-thought chain execution data is generated, which solves the problems of response latency and weak dynamic environment understanding of vision-language-action models in real-time task decision-making, and achieves efficient intelligent decision-making, which can be applied to autonomous driving, robot control and role-playing games.

CN120985639APending Publication Date: 2025-11-21PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511065540.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing vision-language-action models suffer from high response latency, weak understanding of dynamic environments, and scarcity of training data in real-time task decision-making scenarios, making it impossible to make efficient and intelligent decisions.

Method used

A pre-trained VLA model is used to generate structured action thought chain execution data by collecting target operation instructions in real time, and to generate action screen changes on the target display interface. The model is trained by combining progressive learning strategies and expert experience to improve the model's generalization ability and decision-making efficiency.

Benefits of technology

It enables efficient and intelligent decision-making in autonomous driving, robot control, and role-playing games, especially in the medical field to assist in disease diagnosis and surgical procedures, as well as in the field of financial technology, improving the accuracy and efficiency of decision-making in real-time tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120985639A_ABST
    Figure CN120985639A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to a real-time task decision method, device and equipment and a storage medium thereof. Inputting the target operation instruction into a pre-trained VLA model, and generating structured action thinking chain execution data; and performing action picture change generation in a target display interface according to the structured action thinking chain execution data. The real-time task decision-making method is applied to automatic driving, robot control and role playing game control scenes, for example, in the medical application field, a robot is controlled to assist or replace a doctor to carry out etiological examination, surgical operation and the like on a target individual, and in the financial science and technology application field; the corresponding task execution robot is trained to carry out operations such as automatic position holding, flattening and investment operation, and efficient and intelligent decision making of real-time tasks is realized by fully understanding interaction and sequential processing relations among vision, languages and actions.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and is applied to the scenes of automatic driving, robot control and role-playing game control, wherein the robot control, for example, in the medical application field, controls the robot to assist or replace the doctor to perform etiological examination, surgical operation and the like on a target individual, and relates to a real-time task decision method, device, equipment and storage medium thereof. BACKGROUND

[0002] With the rapid development of robot technology, robots are increasingly widely used in industries, medical treatment, services and the like. However, the existing technology still has many challenges in dealing with complex life scenes. In modern life scenes, the robot control system faces the challenge of processing a sample quantity far greater than the training sample. The existing multi-modal large model (MLMS) has poor generalization ability when dealing with various complex and variable life scenes.

[0003] The visual-language-action (VLA) model marks a revolutionary progress in artificial intelligence, aiming to unify perception, natural language understanding and specific action in one computing framework. Currently, the visual-language-action model still has three key bottlenecks in the real-time task decision scene, namely high response delay, weak dynamic environment understanding and training data scarcity. These key bottlenecks result in the current inability to make efficient intelligent decisions for real-time tasks. Therefore, how to improve the efficient intelligent decision-making of real-time tasks has become a technical problem to be solved. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide a real-time task decision method, device, equipment and storage medium to improve the efficient intelligent decision-making of real-time tasks.

[0005] In a first aspect, the embodiments of the present application provide a real-time task decision method, which adopts the technical solution as follows:

[0006] A real-time task decision method includes the following steps:

[0007] Real-time acquisition of a target operation instruction;

[0008] Inputting the target operation instruction into a pre-trained VLA model to generate structured action thought chain execution data, wherein the structured action thought chain execution data is used to control the target video frame picture to change the picture action according to the target operation instruction;

[0009] Generating action picture changes in a target display interface according to the structured action thought chain execution data.

[0010] In a second aspect, the embodiments of the present application further provide a real-time task decision device, which adopts the technical scheme as follows:

[0011] A real-time task decision device comprises:

[0012] An operation instruction real-time collection module is configured to collect a target operation instruction in real time.

[0013] An action execution data generation module is configured to input the target operation instruction into a pre-trained VLA model to generate structured action thought chain execution data, wherein the structured action thought chain execution data is used to control a target video frame picture to change in picture action along with the target operation instruction.

[0014] An action picture change generation module is configured to generate action picture change in a target display interface according to the structured action thought chain execution data.

[0015] In a third aspect, the embodiments of the present application further provide a computer device, which adopts the technical scheme as follows:

[0016] A computer device comprises a memory and a processor, the memory stores computer readable instructions, and the processor implements the steps of the real-time task decision method as described above when executing the computer readable instructions.

[0017] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, which adopts the technical scheme as follows:

[0018] A computer readable storage medium stores computer readable instructions, and the computer readable instructions are executed by a processor to implement the steps of the real-time task decision method as described above.

[0019] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0020] The real-time task decision method described in the application generates structured action thought chain execution data by inputting the target operation instruction into the pre-trained VLA model, and generates action picture changes in the target display interface according to the structured action thought chain execution data. The real-time task decision method is applied to automatic driving, robot control, and role-playing game control scenarios, such as medical application fields, where a robot is controlled to assist or replace a doctor in performing cause examination and surgical operation on a target individual, and financial technology application fields, where a corresponding task execution robot is trained to perform automatic holding, flat, and investment operation, etc. By fully understanding the interaction and time sequence processing relationship between vision, language, and action, efficient intelligent decision of real-time tasks is realized. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the schemes in the application, the following will briefly introduce the drawings needed in the description of the embodiments of the application. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0022] Figure 1 is an exemplary system architecture diagram to which the application can be applied;

[0023] Figure 2 is a flowchart of an embodiment of a real-time task decision method according to the application;

[0024] Figure 3 is a flowchart of a specific embodiment of pre-training of a VLA model in the processing resource adjustment method described in the application;

[0025] Figure 4 is Figure 3 is a flowchart of a specific embodiment of step 301 shown in the figure;

[0026] Figure 5 is Figure 3 is a flowchart of a specific embodiment of step 302 shown in the figure;

[0027] Figure 6 is Figure 3 is a flowchart of a specific embodiment of step 303 shown in the figure;

[0028] Figure 7 is a flowchart of a specific embodiment of updating and training of a VLA model in the processing resource adjustment method described in the application;

[0029] Figure 8is a structural schematic diagram of an embodiment of a real-time task decision device according to the present application;

[0030] Figure 9 is a structural schematic diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application; the use herein of terms such as "comprise" and "have" and any variations thereof are intended to cover a non-exclusive inclusion; the use herein of terms such as "first", "second" and the like are intended to distinguish between similar objects unless the context indicates otherwise.

[0032] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase that the phrase in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. It is explicitly contemplated that embodiments described herein can be combined with each other.

[0033] In order to make the technical personnel in the art better understand the scheme of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings.

[0034] As shown in Figure 1 The system architecture 100 can include a terminal device 101, a network 102 and a server 103, and the terminal device 101 can be a notebook computer 1011, a tablet computer 1012 or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0035] A user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0036] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing, in addition to the notebook computer 1011, the tablet computer 1012 or the mobile phone 1013, the terminal device 101 can also be an electronic book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer and a desktop computer, etc.

[0037] The server 103 can be a server providing various services, for example, a background server supporting a page displayed on the terminal device 101.

[0038] It should be noted that the real-time task decision method provided by the embodiment of the present application is generally executed by a server, and accordingly, the real-time task decision device is generally arranged in the server.

[0039] It should be understood that Figure 1 The number of terminal devices, networks and servers in

[0040] With reference to Figure 2 , a flow chart of one embodiment of a real-time task decision method according to the present application is shown. The real-time task decision method includes the following steps:

[0041] Step 201, collecting a target operation instruction in real time.

[0042] In the embodiment, the target operation instruction includes an operation instruction issued by an instruction issuing object through a human-computer interaction mode, for example, an operation instruction issued by the instruction issuing object through an operation keyboard, for example, an operation instruction issued by the instruction issuing object through an operation button / node, for example, an operation instruction issued by the instruction issuing object through a gesture or the like.

[0043] Specifically, the real-time collection of the target operation instruction includes the real-time collection of the target operation instruction by using a preset instruction collection component, and the instruction collection component includes a collection interface, a scanning interface, etc.

[0044] Specifically, in the execution of the real-time collection of the target operation instruction, a display screen of a target display interface can also be collected synchronously, and the display screen is used as a screen to be changed in action, so as to realize the real-time change processing of the target display screen according to the target operation instruction.

[0045] Step 202, input the target operation instruction into the pre-trained VLA model to generate structured action thought chain execution data, wherein the structured action thought chain execution data is used to control the target video frame picture to change the picture action with the target operation instruction.

[0046] In this embodiment, the pre-trained VLA (Vision-Language-Action) model, wherein the VLA model is a multi-modal artificial intelligence framework that integrates visual perception, language understanding and action control, mainly applied to automatic driving and robot control fields, realizing end-to-end intelligent decision and execution, such as pre-trained OpenVLA model, CEED-VLA model, etc.

[0047] Specifically, in this application, the pre-trained VLA model can control the corresponding target display interface to change the action picture according to the target operation instruction.

[0048] In this embodiment, the target operation instruction includes a series of operation instructions, and the pre-trained VLA (Vision-Language-Action) model can generate structured action thought chain execution data according to the series of operation instructions in time sequence. If the structured action thought chain execution data is generated according to one operation instruction, it can be understood as a picture change generation operation. If the structured action thought chain execution data is generated according to a series of operation instructions, the structured action thought chain execution data can control a series of continuous action picture changes to be generated.

[0049] Step 203, generating action picture changes in the target display interface according to the structured action thought chain execution data.

[0050] In this embodiment, the target display interface refers to the action picture display interface.

[0051] In this embodiment, the target operation instruction is collected in real time; the target operation instruction is input into the pre-trained VLA model to generate structured action thought chain execution data; and the action picture changes are generated in the target display interface according to the structured action thought chain execution data. The real-time task decision method is applied to automatic driving, robot control and role-playing game control scenes, such as medical application field, controlling robot to assist or replace doctor to examine the cause of the target individual, operation, etc., financial technology application field, training corresponding task execution robot to automatically hold, flat, investment operation, etc. Through fully understanding the interaction and time sequence processing relationship among vision, language and action, efficient intelligent decision for real-time task is realized.

[0052] With reference to the foregoing Figure 3 In some optional implementations, before step 202 is performed, the method further includes a step of pre-training a VLA model, Figure 3 is a flowchart of one specific embodiment of the pre-training of the VLA model in the processing resource adjustment method described in the present application, including the following steps:

[0053] Step 301, real-time collection of batch sets of target action pictures and target operation instructions;

[0054] Specifically, by real-time collection of batch sets of target action pictures and target operation instructions, the graphic-text group data is generated according to the actual action picture and operation instruction change relationship, and then the VLA model learning training is performed according to the graphic-text group data, so as to ensure that the pre-trained VLA model fully combines the actual processing task.

[0055] Step 302, cleaning processing of the batch sets of target action pictures and target operation instructions with group-based differential identification to obtain a model training data set;

[0056] Specifically, by cleaning processing of the batch sets of target action pictures and target operation instructions with group-based differential identification, the high availability of the model training data set is ensured, the model training efficiency is improved, and the misleading training of dirty data is also avoided.

[0057] Step 303, inputting the model training data set into a VLA model to be trained, and adopting a progressive learning strategy to learn and train the VLA model to be trained to obtain a pre-trained VLA model.

[0058] In this embodiment, the progressive learning strategy is adopted, that is, the segmented learning, for example: when the VLA model is pre-trained, the action picture data, the operation instruction data and the text language data are involved, therefore, the progressive learning strategy here can be understood as learning and training the different modal data, that is, the action picture data, the operation instruction data and the text language data respectively, and then comprehensively training according to the correlation between them, adopting the progressive learning strategy, so that the learning and training of the visual encoder, the text encoder and the action decoder are not interfered with each other when learning and training the VLA model, and then the learning and training results of different encoders are integrated and trained to obtain the pre-trained VLA model.

[0059] In this embodiment, the target action picture includes a streaming change picture in a target display interface, and the target operation instruction includes operation instruction data of a target operation terminal.

[0060] With reference to the foregoingFigure 4 , Figure 4 is Figure 3 a flowchart of one specific embodiment of step 301, comprising:

[0061] Step 401, through the preset action tracking component, real-time tracking of the target display interface in the flow changes in the picture and the target operation terminal operation instruction data, and record real-time tracking time.

[0062] Among them, the preset action tracking component includes self-developed action tracking component, at least including two tracking processing threads, one of which is used to track the video frame picture of the target display interface in the flow changes in the picture, and the other is used to synchronize the operation instruction of the operation instruction data of the target operation terminal.

[0063] By limiting the action tracking component to include at least two tracking processing threads, one of which is used to track the video frame picture of the target display interface in the flow changes in the picture, and the other is used to synchronize the operation instruction of the operation instruction data of the target operation terminal, so that when training data collection is performed, a corresponding operation instruction can be obtained for each collected changed video frame picture, ensuring the sufficiency and reliability of the learning training data.

[0064] Step 402, the video frame picture and operation instruction data obtained by synchronous tracking are processed by group differentiation marking to obtain group data after group differentiation marking.

[0065] In this embodiment, the two tracking processing threads in the preset action tracking component adopt a millisecond-level synchronous tracking collection mode of multi-thread architecture, so as to collect the flow changes in the picture in the target display interface and the operation instruction data of the target operation terminal with millisecond-level collection efficiency according to the two tracking processing threads, thereby realizing real-time decision of action execution task.

[0066] In this embodiment, after the step of performing real-time collection of batch group target action picture and target operation instruction, i.e. step 301, and before the step of performing cleaning processing on the batch group target action picture and target operation instruction with group differentiation identification to obtain model training data set, i.e. step 302, the method further comprises: according to the real-time tracking time, all group data are time-sequenced and arranged to obtain time-sequenced group data; the time-sequenced group data is updated to the batch group target action picture and target operation instruction.

[0067] Specifically, between step 301 and step 302, all collected group data is also time-sequenced according to the real-time tracking time, time-sequenced group data is obtained, and the time-sequenced group data is updated as the batch group target action picture and target operation instruction, so that when subsequent learning and training of the VLA model is performed, the VLA model can fully learn the time sequence change relationship between actions, pictures, and actions and pictures, and the high availability and generalization processing capability of the VLA model are improved.

[0068] With reference to the foregoing Figure 5 , Figure 5 is Figure 3 a flowchart of one specific embodiment of step 302, including:

[0069] Step 501, according to the preset cleaning rule and time-sequencing relationship, deleting the group data obviously existing error execution or delay execution in the batch group target action picture and target operation instruction, obtaining the optimized time-sequenced group data;

[0070] Specifically, the preset cleaning rule is a rule set according to the relationship between actual operation instruction and picture, for example: when the current picture obviously changes compared with the previous picture, the operation instruction corresponding to the current picture is obviously changed compared with the operation instruction corresponding to the previous picture, obviously, at this time, there is an error execution; or, the current operation instruction has been issued, but the current corresponding picture has not produced any change, here, there is obviously a delay execution problem, that is, according to the error execution or delay execution, the group data obviously existing error execution or delay execution in the batch group target action picture and target operation instruction is deleted, so as to reduce the misleading of the model learning and training by the defective data when learning and training the VLA model.

[0071] Step 502, setting the optimized time-sequenced group data as the model training data set.

[0072] In this embodiment, after the step of performing cleaning processing on the batch group target action picture and target operation instruction to distinguish and identify by group to obtain the model training data set, the method further includes: adding text knowledge description to each group of training samples in the model training data set according to the preset expert experience, wherein each group of training samples contains a video frame picture and an operation instruction synchronously tracked and collected with the video frame picture.

[0073] Specifically, the textualized knowledge description includes only textualized description of the video frame pictures in the training samples, only textualized description of the operation instructions in the training samples, or simultaneous textualized description of the video frame pictures and the operation instructions in the training samples.

[0074] By adding the textualized knowledge description to each group of training samples in the model training data set according to the preset expert experience, more languageized feature texts for the operation instructions and the picture changes are learned during the learning training of the VLA model, so that the pre-trained VLA model can ensure sufficient semantic understanding and recognition in the visual-language-action.

[0075] With reference to Figure 6 , Figure 6 is Figure 3 a flowchart of one specific embodiment of the step 303, including:

[0076] In step 601, the visual encoder in the VLA model to be trained is used to perform visual feature coding processing on all video frame pictures in the model training data set, to obtain corresponding visual features of all video frame pictures.

[0077] In step 602, the time sequence changes of the visual features corresponding to all video frame pictures are learned according to the time sequence of the video frame pictures in the model training data set.

[0078] In essence, steps 601 and 602 are learning training of the visual encoder in the VLA model, so that the VLA model finally learns the time sequence changes of the visual features corresponding to all video frame pictures.

[0079] In step 603, the action decoder in the VLA model to be trained is used to perform decoding processing on all operation instructions in the model training data set, to obtain corresponding picture change actions of all operation instructions.

[0080] In step 604, the picture action time sequence changes corresponding to all operation instructions are learned according to the time sequence of the operation instructions in the model training data set.

[0081] In essence, steps 603 and 604 are learning training of the action decoder in the VLA model, so that the VLA model finally learns the picture action time sequence changes corresponding to all operation instructions.

[0082] Step 605, according to the time sequence changes of the picture motion corresponding to all operation instructions and the time sequence changes of the visual features corresponding to all video frame pictures and the mapping relationship between the video frame pictures and the operation instructions, learning the picture motion change relationship of all video frame pictures with the operation instructions;

[0083] Specifically, the mapping relationship between the video frame pictures and the operation instructions during real-time data acquisition is combined to learn the picture motion change relationship of all video frame pictures with the operation instructions.

[0084] Step 606, deploying the picture motion change relationship of all video frame pictures with the operation instructions as action thinking chain execution knowledge to the VLA model to obtain a pre-trained VLA model.

[0085] By deploying the learned picture motion change relationship of all video frame pictures with the operation instructions as action thinking chain execution knowledge to the VLA model, the picture motion change decision of the current video frame picture can be realized directly according to the collected operation instructions when the trained VLA model is actually used subsequently.

[0086] Continue to refer to Figure 7 In some optional implementations, after step 303 is executed, the method further includes a step of updating and training the VLA model, Figure 7 is a flowchart of one specific embodiment of the processing resource adjustment method described in the present application, including the following steps:

[0087] Step 701, using the language encoder in the pre-trained VLA model to perform text feature encoding processing on the text-based knowledge description corresponding to each group of training samples in the model training data set respectively, to obtain the text features corresponding to all training samples respectively;

[0088] Specifically, the language encoder in the VLA model is used to perform text feature encoding processing on the text-based knowledge description corresponding to each group of training samples in the model training data set respectively, to obtain the text features corresponding to all training samples respectively.

[0089] Step 702, learning the association mapping relationship between the visual features and the text features according to the visual features corresponding to all video frame pictures respectively, the text features corresponding to all training samples respectively, and the training samples to which all video frame pictures belong;

[0090] Specifically, the association mapping relationship between the visual features and the text features is learned according to the visual features corresponding to all video frame pictures respectively, the text features corresponding to all training samples respectively, and the training samples to which all video frame pictures belong.

[0091] At step 703, the association mapping relationship between the picture action time sequence change and the text feature is learned according to the picture action time sequence change corresponding to all operation instructions, the text feature corresponding to all training samples respectively, and the training sample to which all operation instructions belong.

[0092] Specifically, the association mapping relationship between the picture action time sequence change and the text feature is learned.

[0093] At step 704, the association mapping relationship between the picture action time sequence change and the text feature is deployed into the pre-trained VLA model as action thinking chain execution supplement knowledge, to obtain an updated VLA model.

[0094] By deploying the association mapping relationship between the picture action time sequence change and the text feature into the pre-trained VLA model as action thinking chain execution supplement knowledge, when generating the picture action change according to the operation instruction subsequently, the text semantic feature of the operation instruction is combined on the basis of the picture action change relationship of all previous video frames, so that the accurate execution of the final VLA model is ensured, and the efficient and accurate decision-making of the real-time decision task is improved.

[0095] In the embodiment, the target operation instruction is collected in real time, the target operation instruction is input into the pre-trained VLA model to generate structured action thinking chain execution data, and action picture change is generated in a target display interface according to the structured action thinking chain execution data. The real-time task decision method is applied to automatic driving, robot control, and role-playing game control scenes, for example, in the medical application field, a robot is controlled to assist or replace a doctor to perform cause examination, operation, and the like on a target individual, in the financial technology application field, a corresponding task execution robot is trained to perform automatic holding, flat, investment, and the like, and by fully understanding the interaction and time sequence processing relationship among vision, language, and action, efficient intelligent decision-making for real-time tasks is realized.

[0096] The embodiments of the application can acquire and process related data based on artificial intelligence technology. Artificial intelligence (AI) is a theory, method, technology, and application system for simulating, extending, and expanding human intelligence by using a digital computer or a machine controlled by a digital computer, perceiving an environment, acquiring knowledge, and using the knowledge to obtain optimal results.

[0097] The artificial intelligence basic technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. The artificial intelligence software technology mainly includes computer vision technology, robot technology, biometric identification technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.

[0098] In this embodiment, the target operation instruction is collected in real time; the target operation instruction is input into the pre-trained VLA model to generate structured action thought chain execution data; and action picture change generation is performed in a target display interface according to the structured action thought chain execution data. The real-time task decision method is applied to automatic driving, robot control, and role-playing game control scenarios, for example: in the medical application field, a robot is controlled to assist or replace a doctor to perform etiological examination, surgical operation, etc. on a target individual; in the financial technology application field, a corresponding task execution robot is trained to perform automatic holding, flat, investment, etc. operation. By fully understanding the interaction and timing processing relationship among vision, language and action, efficient intelligent decision of a real-time task is realized.

[0099] Further referring to Figure 8 , as an implementation of the method shown in Figure 2 , the present application provides an embodiment of a real-time task decision device, which corresponds to the method embodiment shown in Figure 2 . The device can be applied to various electronic devices.

[0100] As shown in Figure 8 , the real-time task decision device 800 described in this embodiment includes an operation instruction real-time collection module 801, an action execution data generation module 802, and an action picture change generation module 803. Among them:

[0101] The operation instruction real-time collection module 801 is configured to collect a target operation instruction in real time.

[0102] The action execution data generation module 802 is configured to input the target operation instruction into a pre-trained VLA model to generate structured action thought chain execution data, wherein the structured action thought chain execution data is used to control a target video frame picture to change in picture action along with the target operation instruction.

[0103] The action picture change generation module 803 is configured to generate action picture change in a target display interface according to the structured action thought chain execution data.

[0104] The application collects target operation instructions in real time; inputs the target operation instructions into a pre-trained VLA model to generate structured action thought chain execution data; and generates action picture changes in a target display interface according to the structured action thought chain execution data. The real-time task decision method is applied to automatic driving, robot control and role-playing game control scenarios, for example: in the medical application field, a robot is controlled to assist or replace a doctor to perform etiological examination and surgical operation on a target individual; in the financial technology application field, a corresponding task execution robot is trained to perform automatic holding, closing and investment operation; and through full understanding of the interaction and time sequence processing relationship among vision, language and action, efficient intelligent decision of a real-time task is realized.

[0105] In this embodiment, the real-time task decision device 800 further includes a real-time batch acquisition module, a model training data set obtaining module and a VLA model pre-training module. Among them:

[0106] The real-time batch acquisition module is used for real-time acquisition of batched target action pictures and target operation instructions;

[0107] The model training data set obtaining module is used for cleaning and processing the batched target action pictures and target operation instructions with group-based distinguishing marks to obtain a model training data set;

[0108] The VLA model pre-training module is used for inputting the model training data set into a VLA model to be trained, using a progressive learning strategy to learn and train the VLA model to be trained, and obtaining a pre-trained VLA model.

[0109] In this embodiment, the real-time batch acquisition module includes a real-time tracking acquisition unit and a group-based processing unit. Among them:

[0110] The real-time tracking acquisition unit is used for real-time tracking of streaming change pictures in a target display interface and operation instruction data of a target operation terminal through a preset action tracking component, and recording real-time tracking time, wherein the preset action tracking component includes a self-developed action tracking component and at least two tracking processing threads, one of which is used for tracking video frame pictures of streaming change pictures in the target display interface, and the other is used for synchronously tracking operation instructions of operation instruction data of the target operation terminal;

[0111] The group-based processing unit is used for group-based distinguishing mark processing of the video frame pictures and the operation instruction data obtained by synchronous tracking to obtain group-based distinguishing mark processed grouped data.

[0112] In this embodiment, the real-time task decision device 800 further comprises a time sequence arrangement module and a data updating module. Among them:

[0113] The time sequence arrangement module is configured to arrange all the grouped data in time sequence according to the real-time tracking time, to obtain time sequence grouped data.

[0114] The data updating module is configured to update the time sequence grouped data to the batch grouped target action screen and target operation instruction.

[0115] In this embodiment, the real-time task decision device 800 further comprises a data optimization processing module and a text knowledge description adding module. Among them:

[0116] The data optimization processing module is configured to delete the grouped data that obviously exists in the batch grouped target action screen and target operation instruction according to the preset cleaning rule and time sequence relationship, to obtain the optimized time sequence grouped data; and is further configured to set the optimized time sequence grouped data as the model training data set.

[0117] The text knowledge description adding module is configured to add text knowledge description to each training sample in the model training data set according to the preset expert experience, wherein each training sample comprises a video frame screen and an operation instruction collected synchronously with the video frame screen.

[0118] In this embodiment, the VLA model pre-training module comprises a visual feature coding unit, a first learning unit, an operation instruction decoding unit, a second learning unit, a first comprehensive learning unit, and an action execution knowledge deployment unit. Among them:

[0119] The visual feature coding unit is configured to use the visual encoder in the VLA model to be trained to perform visual feature coding processing on all video frame screens in the model training data set, to obtain the visual features corresponding to all video frame screens respectively.

[0120] The first learning unit is configured to learn the time sequence change of the visual features corresponding to all video frame screens according to the time sequence of all video frame screens in the model training data set.

[0121] The operation instruction decoding unit is configured to use the action decoder in the VLA model to be trained to decode all operation instructions in the model training data set, to obtain the screen change actions corresponding to all operation instructions respectively.

[0122] The second learning unit is configured to learn the time sequence change of the screen action corresponding to all operation instructions according to the time sequence of all operation instructions in the model training data set.

[0123] The first comprehensive learning unit is configured to learn the screen action change relationship of all video frame screens with operation instructions according to the screen action time sequence change corresponding to all operation instructions, the time sequence change of the visual features corresponding to all video frame screens, and the mapping relationship between the video frame screens and the operation instructions.

[0124] The action execution knowledge deployment unit is configured to deploy the screen action change relationship of all video frame screens with operation instructions as action thought chain execution knowledge to the VLA model to obtain a pre-trained VLA model.

[0125] In this embodiment, the real-time task decision device 800 further includes a text feature encoding module, a first supplementary learning module, a second supplementary learning module, and a VLA model updating module. Specifically:

[0126] The text feature encoding module is configured to perform text feature encoding processing on the text knowledge description corresponding to each training sample in the model training data set by using a language encoder in the pre-trained VLA model, to obtain the text features corresponding to all training samples.

[0127] The first supplementary learning module is configured to learn the association mapping relationship between the visual features and the text features according to the visual features corresponding to all video frame screens, the text features corresponding to all training samples, and the training samples to which all video frame screens belong.

[0128] The second supplementary learning module is configured to learn the association mapping relationship between the screen action time sequence change and the text features according to the screen action time sequence change corresponding to all operation instructions, the text features corresponding to all training samples, and the training samples to which all operation instructions belong.

[0129] The VLA model updating module is configured to deploy the association mapping relationship between the screen action time sequence change and the text features as action thought chain execution supplementary knowledge to the pre-trained VLA model to obtain an updated VLA model.

[0130] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by computer readable instructions instructing related hardware, and the computer readable instructions can be stored in a computer readable storage medium. When the program is executed, the processes of the above-mentioned embodiments can be included. The storage medium can be a non-volatile storage medium such as a magnetic disc, an optical disc, a read-only memory (ROM), or a random access memory (RAM).

[0131] It should be understood that although each step in the flowchart of the accompanying drawings is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless explicitly stated herein, the execution of these steps is not strictly limited in sequence, and they can be executed in other sequences. Moreover, at least part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or sub-steps or stages of other steps.

[0132] To solve the above technical problems, the embodiments of the present application further provide a computer device. For details, please refer to Figure 9 Figure 9 The basic structure block diagram of the computer device of the present embodiment is shown in FIG. 1.

[0133] The computer device 9 comprises a memory 9a, a processor 9b and a network interface 9c which are connected to each other through a system bus. It should be pointed out that Figure 9 The computer device 9 with the components of the memory 9a, the processor 9b and the network interface 9c is only shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, the computer device herein is a device capable of automatically performing numerical calculation and / or information processing according to pre-set or stored instructions, which hardware includes but is not limited to microprocessor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), digital signal processor (DSP), embedded device, etc.

[0134] The computer device can be a desktop computer, a notebook computer, a palm computer and a cloud server, etc. The computer device can interact with the user through a keyboard, a mouse, a remote controller, a touchpad or a voice control device, etc.

[0135] ​The memory 9a includes at least one type of readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 9a can be an internal storage unit of the computer device 9, such as a hard disk or a memory of the computer device 9. In other embodiments, the memory 9a can also be an external storage device of the computer device 9, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 9. Of course, the memory 9a can also include both an internal storage unit and an external storage device of the computer device 9. In this embodiment, the memory 9a is generally used to store an operating system and various application software installed on the computer device 9, such as computer readable instructions of a real-time task decision method, etc. In addition, the memory 9a can also be used to temporarily store various data that have been output or will be output.

[0136] The processor 9b can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip in some embodiments. The processor 9b is generally used to control the overall operation of the computer device 9. In this embodiment, the processor 9b is used to run computer readable instructions or process data stored in the memory 9a, such as computer readable instructions of the real-time task decision method.

[0137] The network interface 9c can include a wireless network interface or a wired network interface, and is generally used to establish a communication connection between the computer device 9 and other electronic devices.

[0138] The computer device provided in the embodiment belongs to the technical field of artificial intelligence, and is applied to automatic driving, robot control and role-playing game control scenes, wherein the robot control is, for example, in the medical application field, the robot is controlled to assist or replace doctors to perform cause examination, operation and the like on target individuals. The application generates structured action thought chain execution data by inputting the target operation instruction into the pre-trained VLA model, and generates action picture changes in the target display interface according to the structured action thought chain execution data. The real-time task decision method is applied to automatic driving, robot control and role-playing game control scenes, for example, in the medical application field, the robot is controlled to assist or replace doctors to perform cause examination, operation and the like on target individuals, in the financial technology application field, the corresponding task execution robot is trained to perform automatic holding, flat, investment and operation and the like, and through fully understanding the interaction and time sequence processing relationship among vision, language and action, efficient intelligent decision for real-time tasks is realized.

[0139] The application also provides another implementation, that is, a computer readable storage medium storing computer readable instructions, the computer readable instructions can be executed by a processor to make the processor execute the steps of a real-time task decision method as described above.

[0140] The computer readable storage medium provided in the embodiment belongs to the technical field of artificial intelligence, and is applied to automatic driving, robot control and role-playing game control scenes, wherein the robot control is, for example, in the medical application field, the robot is controlled to assist or replace doctors to perform cause examination, operation and the like on target individuals. The application generates structured action thought chain execution data by inputting the target operation instruction into the pre-trained VLA model, and generates action picture changes in the target display interface according to the structured action thought chain execution data. The real-time task decision method is applied to automatic driving, robot control and role-playing game control scenes, for example, in the medical application field, the robot is controlled to assist or replace doctors to perform cause examination, operation and the like on target individuals, in the financial technology application field, the corresponding task execution robot is trained to perform automatic holding, flat, investment and operation and the like, and through fully understanding the interaction and time sequence processing relationship among vision, language and action, efficient intelligent decision for real-time tasks is realized.

[0141] Those skilled in the art can clearly understand the above-mentioned method of the embodiment can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) execute the method described in each embodiment of the present application.

[0142] Obviously, the above-described embodiments are only some of the embodiments of the present application, not all the embodiments, and the drawings give the preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms, and conversely, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing specific embodiments, or make equivalent replacements to some technical features. Any equivalent structure made by using the content of the specification and drawings, directly or indirectly applied to other related technical fields, is also within the scope of the patent protection of the present application. The non-company software tools or components appearing in the embodiments of the present application are only illustrative, not representing the actual use.

Claims

1. A method for real-time task decision-making, characterized in that, The method comprises the following steps: Real-time acquisition of target operation instructions; Input the target operation instructions into the pre-trained VLA model to generate structured action thought chain execution data, wherein the structured action thought chain execution data is used to control the target video frame picture to change the picture action with the target operation instructions; According to the structured action thought chain execution data, the action picture change is generated in the target display interface.

2. The real-time task decision method according to claim 1, characterized in that, Before the step of inputting the target operation instructions into the pre-trained VLA model to generate structured action thought chain execution data, the method further comprises: Real-time acquisition of batch grouped target action pictures and target operation instructions; Cleaning processing is performed on the batch grouped target action pictures and target operation instructions with group-based differentiated identification to obtain a model training data set; The model training data set is input into the VLA model to be trained, and a progressive learning strategy is used to learn and train the VLA model to be trained to obtain a pre-trained VLA model.

3. The real-time task decision method of claim 2, wherein, The target action picture includes a streaming change picture in the target display interface, and the target operation instruction includes operation instruction data of a target operation terminal. The step of real-time acquisition of batch grouped target action pictures and target operation instructions specifically comprises: Through a preset action tracking component, real-time tracking of streaming change pictures in the target display interface and operation instruction data of the target operation terminal is performed, and real-time tracking time is recorded. The preset action tracking component includes a self-developed action tracking component and at least two tracking processing threads. One of the tracking processing threads is used to track the video frame picture of the streaming change picture in the target display interface, and the other tracking processing thread is used to synchronously track the operation instruction data of the operation instruction of the target operation terminal. Grouped data after group-based differentiated marking processing is obtained by performing group-based differentiated marking processing on the synchronously tracked video frame picture and operation instruction data.

4. The real-time task decision method of claim 3, wherein, The two tracking processing threads in the preset action tracking component adopt a millisecond-level synchronous tracking acquisition mode of a multi-thread architecture. After the step of real-time acquisition of batch grouped target action pictures and target operation instructions is performed, and before the step of cleaning processing of the batch grouped target action pictures and target operation instructions with group-based differentiated identification is performed to obtain a model training data set, the method further comprises: According to the real-time tracking time, all grouped data is time-sequenced to obtain time-sequenced grouped data; The time-sequenced grouped data is updated as the batch grouped target action pictures and target operation instructions.

5. The real-time task decision method according to any one of claims 2 or 4, characterized in that, The step of cleaning processing of the batch grouped target action pictures and target operation instructions with group-based differentiated identification to obtain a model training data set specifically comprises: According to a preset cleaning rule and a time-sequencing relationship, the grouped data with obvious error execution or delay execution in the batch grouped target action pictures and target operation instructions is deleted to obtain optimized time-sequenced grouped data; set the sequenced grouped data after the optimization as the model training dataset; After the step of performing the cleaning processing on the batched target action screen and target operation instruction to distinguish and identify by group, obtaining the model training dataset, the method further comprises: According to the preset expert experience, add a text knowledge description to each training sample in the model training dataset, wherein each training sample contains a video frame screen and an operation instruction collected synchronously with the video frame screen.

6. The real-time task decision method of claim 5, wherein, The step of inputting the model training dataset into the VLA model to be trained, adopting a progressive learning strategy to learn and train the VLA model to be trained, and obtaining a pre-trained VLA model, specifically comprises: Using the visual encoder in the VLA model to be trained, the visual features of all video frame screens in the model training dataset are coded to obtain the corresponding visual features of all video frame screens; According to the time sequence of all video frame screens in the model training dataset, the time sequence changes of the visual features corresponding to all video frame screens are learned; Using the action decoder in the VLA model to be trained, all operation instructions in the model training dataset are decoded to obtain the corresponding screen change actions of all operation instructions; According to the time sequence of all operation instructions in the model training dataset, the time sequence changes of the screen actions corresponding to all operation instructions are learned; According to the time sequence changes of the screen actions corresponding to all operation instructions, the time sequence changes of the visual features corresponding to all video frame screens, and the mapping relationship between the video frame screens and the operation instructions, the screen action change relationship of all video frame screens with the operation instructions is learned; The screen action change relationship of all video frame screens with the operation instructions is deployed as an action thinking chain execution knowledge into the VLA model to obtain a pre-trained VLA model.

7. The real-time task decision method of claim 6, wherein, After the step of inputting the model training dataset into the VLA model to be trained, adopting a progressive learning strategy to learn and train the VLA model to be trained, and obtaining a pre-trained VLA model, the method further comprises: Using the language encoder in the pre-trained VLA model, the text knowledge description corresponding to each training sample in the model training dataset is textually coded to obtain the corresponding text features of all training samples; According to the corresponding visual features of all video frame screens, the corresponding text features of all training samples, and the training samples to which all video frame screens belong, the association mapping relationship between the visual features and the text features is learned; According to the time sequence changes of the screen actions corresponding to all operation instructions, the corresponding text features of all training samples, and the training samples to which all operation instructions belong, the association mapping relationship between the time sequence changes of the screen actions and the text features is learned; The mapping relationship between the picture motion time sequence change and the text features is taken as action thinking chain execution supplement knowledge and deployed to the pre-trained VLA model to obtain an updated VLA model.

8. A real-time task decision device, characterized by comprising: The method comprises the following steps: An operation instruction real-time acquisition module is configured to acquire a target operation instruction in real time. An action execution data generation module is configured to input the target operation instruction into a pre-trained VLA model to generate structured action thinking chain execution data, wherein the structured action thinking chain execution data is used to control a target video frame picture to change in picture motion in accordance with the target operation instruction. An action picture change generation module is configured to generate an action picture change in a target display interface according to the structured action thinking chain execution data.

9. A computer device, comprising: The method comprises the following steps: a memory and a processor, wherein the memory stores computer readable instructions, and the processor executes the computer readable instructions to realize the steps of the real-time task decision method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer readable instructions, and the computer readable instructions are executed by the processor to realize the steps of the real-time task decision method according to any one of claims 1 to 7.