Unmanned vehicle optimal action sequence generation system and method based on cooperation of large model and special model
By introducing a action sequence generation system that collaborates between large models and dedicated models in the unmanned vehicle system, the problem of low accuracy and automation in the understanding and decision-making process of traditional unmanned vehicle tasks is solved, and a more efficient and more adaptable unmanned vehicle action sequence generation is achieved.
Patent Information
- Application Number
- CN202510289761.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-20
AI Technical Summary
In the process of understanding and decision-making of traditional unmanned vehicle tasks, the action plan relies on the operator's experience or fixed algorithms, resulting in low accuracy, low efficiency, poor adaptability and low degree of automation.
The optimal action sequence generation system for unmanned vehicles is adopted based on large models and dedicated models, including user interaction modules, voice processing modules, information real-time perception modules and action sequence generation modules. The system uses a vertical large model of task understanding and decision making, a task planning knowledge base, a dedicated model calling framework and an evaluation model to generate the optimal action sequence.
The accuracy, efficiency, adaptability and automation of the generation of unmanned vehicle action sequences is improved, and the occupant instructions can be more accurately understood, multiple action plans can be generated, and the final action sequence can be optimized through special models and evaluation models.
Smart Images

Figure CN120171591A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of unmanned vehicle control, and particularly to an unmanned vehicle optimal action sequence generation system and method based on the cooperation of a large model and a dedicated model. Background Art
[0002] In the traditional task understanding and decision-making process of unmanned vehicles, the action plan mostly depends on the operator to command according to experience or fixed traditional algorithms. The operator needs to use remote control means to control the unmanned vehicle and adopt a serial thinking mode for task planning. The autonomous ability of the unmanned vehicle is poor. In this command and control mode, the operator needs to pay attention to the execution of the action in real time and conduct comprehensive analysis and judgment considering factors such as time and action progress. There are many crew intervention links, the correctness of the decision-making plan is insufficient, and the generalization ability is poor, often leading to falling into local optimal solutions.
[0003] In future unmanned vehicle cluster control applications facing fast-paced and high-intensity scenarios, there is a prominent contradiction between the small number of crew members and the large number of unmanned vehicle nodes. Facing the needs of crew members for complex task understanding and efficient command and control of unmanned vehicles in future application scenarios, an unmanned vehicle optimal action sequence generation scheme with high accuracy, high efficiency, strong adaptability, and high automation is required. Summary of the Invention
[0004] The purpose of this application is to provide an unmanned vehicle optimal action sequence generation system and method based on the cooperation of a large model and a dedicated model to solve the problems of low accuracy, low efficiency, poor adaptability, and low automation in the generation of unmanned vehicle action sequences.
[0005] To achieve the above purpose, the following solutions are provided in this application.
[0006] In the first aspect, this application provides an unmanned vehicle optimal action sequence generation system based on the cooperation of a large model and a dedicated model, including: a user interaction module, a voice processing module, an information real-time perception module, and an action sequence generation module; the user interaction module is connected to the voice processing module, and both the voice processing module and the information real-time perception module are connected to the action sequence generation module; The user interaction module is used to obtain the natural voice task instructions of the crew for the tasks to be executed by the unmanned vehicle; The voice processing module is used to determine the instruction text information based on the natural voice task instructions; The information real-time perception module is used to obtain the scene state information of the unmanned vehicle; In the action sequence generation module, a task understanding and decision-making vertical large model, a task planning knowledge base, a dedicated model call framework, dedicated models of multiple action strategies, a dedicated model dynamic library of multiple dedicated models, and an evaluation model are deployed; The action sequence generation module is used for: Using the task understanding and decision-making vertical large model, based on the instruction text information and the scenario state information, retrieve from the task planning knowledge base to obtain relevant knowledge and historical cases corresponding to the instruction text information, and arrange and splice the relevant knowledge and the historical cases to generate a prompt word segment of the instruction text information; Using the task understanding and decision-making vertical large model, based on the prompt word segment, conduct task understanding on the instruction text information, decompose the to-be-executed task into multiple subtasks, conduct task planning on each subtask, generate multiple action strategies corresponding to each subtask and corresponding specific parameter call information, so as to obtain multiple types of action plans for the to-be-executed task; Determine any one of the action plans as the current plan; Using the dedicated model call framework, respectively send the specific parameter call information of the dedicated models of each action strategy in the current plan to the corresponding dedicated models, and autonomously call the corresponding dedicated models for parsing operations to obtain corresponding parsing operation results; Dynamically integrate the parsing operation results of each dedicated model to obtain the action sequence of the current plan; Using the evaluation model, respectively determine the comprehensive evaluation values of each action plan based on the action sequences of each action plan, and determine the action sequence of the action plan with the optimal comprehensive evaluation value as the optimal action sequence for executing the to-be-executed task.
[0007] Optionally, the optimal action sequence generation system for an autonomous vehicle based on the collaboration of a large model and a dedicated model further includes: an autonomous vehicle execution module, and the autonomous vehicle execution module is connected to the action sequence generation module; The autonomous vehicle execution module is used to control the autonomous vehicle to execute the to-be-executed task based on the optimal action sequence.
[0008] Optionally, the GTCRN model and the Whisper speech recognition model are deployed in the speech processing module; Determining the instruction text information based on the natural speech task instruction includes: Using the GTCRN model to filter the noise of the natural speech task instruction to obtain the processed natural speech task instruction; Using the Whisper speech recognition model to transcribe the speech and recognize the instruction of the processed natural speech task instruction to obtain the instruction text information.
[0009] Optionally, the scenario state information includes: geographical environment information, task objective information, and own side state information; The geographical environment information includes: scene information and road network information. The scene information is an urban scene or a wild scene, and the road network information is a scene with a road network or without a road network; The mission objective information includes: personnel, mobile platforms, and fixed buildings; The own-side status information includes: driving speed, current battery level, and panoramic mirror availability information. The panoramic mirror availability information is that the panoramic mirror is available or unavailable.
[0010] Optionally, the mission planning knowledge base includes: relevant knowledge and historical cases corresponding to multiple instruction text information; The relevant knowledge includes: mission scene introduction, definitions of various action strategies, basic parameters of the unmanned vehicle, and action strategies recommended under geographical environment information; The historical cases include: action plans corresponding to relevant historical instruction text information.
[0011] Optionally, the specific parameter call information includes: timestamp, action sequence ID, path planning mode, driving speed, type of dedicated model called, longitude and latitude of the starting point of the unmanned vehicle, longitude and latitude of the ending point, sensor type, reconnaissance area range, and time limit requirements.
[0012] In a second aspect, the present application provides an optimal action sequence generation method for an unmanned vehicle based on the cooperation of a large model and a dedicated model, which is implemented based on the optimal action sequence generation system for an unmanned vehicle based on the cooperation of a large model and a dedicated model according to any one of the above. The optimal action sequence generation method for an unmanned vehicle based on the cooperation of a large model and a dedicated model includes: Obtain the natural language task instruction of the mission to be executed by the crew for the unmanned vehicle; Based on the natural language task instruction, determine the instruction text information; Obtain the scene status information of the unmanned vehicle; Use the task understanding and decision-making vertical large model to retrieve from the mission planning knowledge base based on the instruction text information and the scene status information, obtain the relevant knowledge and historical cases corresponding to the instruction text information, and arrange and splice the relevant knowledge and the historical cases to generate a prompt word segment of the instruction text information; Use the task understanding and decision-making vertical large model to perform task understanding on the instruction text information based on the prompt word segment, decompose the mission to be executed into multiple subtasks, perform mission planning on each subtask, generate multiple action strategies and corresponding specific parameter call information for each subtask, so as to obtain multiple types of action plans for the mission to be executed; Determine any one of the action plans as the current plan; Using a dedicated model invocation framework, the specific parameter invocation information of the dedicated models of each action strategy in the current solution is respectively sent to the corresponding dedicated models, and the corresponding dedicated models are autonomously invoked for parsing operations to obtain the corresponding parsing operation results; Dynamically integrate the parsing operation results of each dedicated model to obtain the action sequence of the current solution; Using an evaluation model, based on the action sequences of each action plan respectively, determine the comprehensive evaluation values of each action plan, and determine the action sequence of the action plan with the optimal comprehensive evaluation value as the optimal action sequence for executing the to-be-executed task.
[0013] According to the specific embodiments provided in this application, the following technical effects are disclosed in this application: The present application discloses an optimal action sequence generation system and method for an autonomous vehicle based on the collaboration of a large model and a dedicated model. In the optimal action sequence generation system for an autonomous vehicle based on the collaboration of a large model and a dedicated model, a user interaction module obtains the natural speech task instructions of the occupant for the tasks to be executed by the autonomous vehicle; a speech processing module determines the instruction text information based on the natural speech task instructions; an information real-time perception module obtains the scene state information of the autonomous vehicle; in the action sequence generation module, a task understanding and decision-making vertical large model, a task planning knowledge base, a dedicated model calling framework, dedicated models of multiple action strategies, a dedicated model dynamic library of multiple dedicated models, and an evaluation model are deployed; the action sequence generation module is used for: using the task understanding and decision-making vertical large model, based on the instruction text information and the scene state information, retrieving from the task planning knowledge base to obtain relevant knowledge and historical cases corresponding to the instruction text information, and arranging and splicing the relevant knowledge and historical cases to generate a prompt word segment of the instruction text information; using the task understanding and decision-making vertical large model, based on the prompt word segment, performing task understanding on the instruction text information, decomposing the task to be executed into multiple subtasks, performing task planning on each subtask, generating multiple action strategies corresponding to each subtask and corresponding specific parameter call information, so as to obtain multiple types of action plans for the task to be executed; determining any one of the action plans as the current plan; using the dedicated model calling framework, respectively sending the specific parameter call information of the dedicated models of each action strategy in the current plan to the corresponding dedicated models, autonomously calling the corresponding dedicated models for parsing operations, and obtaining the corresponding parsing operation results; dynamically integrating the parsing operation results of each dedicated model to obtain the action sequence of the current plan; using the evaluation model, respectively based on the action sequences of each action plan, determining the comprehensive evaluation values of each action plan, and determining the action sequence of the action plan with the optimal comprehensive evaluation value as the optimal action sequence for executing the task to be executed. The present application uses the task understanding and decision-making vertical large model and dedicated models of multiple action strategies deployed in the action sequence generation module to generate the optimal action sequence, improving the accuracy, efficiency, adaptability, and automation degree of generating the action sequence of the autonomous vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 It is a schematic structural diagram of an optimal action sequence generation system for an autonomous vehicle based on the collaboration of a large model and a dedicated model provided by an embodiment of the present application.
[0016] Figure 2 It is a schematic diagram of the architecture for generating the optimal action sequence of an autonomous vehicle. Specific implementation manners
[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0018] The purpose of the present application is to provide an autonomous vehicle optimal action sequence generation system and method based on the cooperation of a large model and a dedicated model, aiming to improve the accuracy, efficiency, adaptability, and automation degree of generating the action sequence of an autonomous vehicle.
[0019] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0020] In an exemplary embodiment, as Figure 1 and Figure 2 shown, an autonomous vehicle optimal action sequence generation system based on the cooperation of a large model and a dedicated model is provided, including: a user interaction module, a voice processing module, an information real-time perception module, and an action sequence generation module; the user interaction module is connected to the voice processing module, and both the voice processing module and the information real-time perception module are connected to the action sequence generation module.
[0021] The user interaction module is used to obtain the natural voice task instruction of the crew member for the task to be executed by the autonomous vehicle.
[0022] The voice processing module is used to determine the instruction text information based on the natural voice task instruction.
[0023] The information real-time perception module is used to obtain the scene state information of the autonomous vehicle.
[0024] In the action sequence generation module, a task understanding and decision-making vertical large model, a task planning knowledge base, a dedicated model call framework, dedicated models of multiple action strategies, a dedicated model dynamic library of multiple dedicated models, and an evaluation model are deployed.
[0025] Specifically, the task understanding and decision-making vertical large model is a specific domain large model for autonomous vehicle task understanding and decision-making formed by training and fine-tuning on the open-source general large model base using the constructed autonomous vehicle task instruction data set, which can receive task instructions for task understanding, select appropriate action strategies, and generate reasonable action sequences.
[0026] The action sequence generation module is used for: Using a task understanding and decision-making vertical large model, retrieve from the task planning knowledge base based on the instruction text information and scenario status information to obtain relevant knowledge and historical cases corresponding to the instruction text information, and arrange and splice the relevant knowledge and historical cases to generate a prompt word segment of the instruction text information; Using a task understanding and decision-making vertical large model, perform task understanding on the instruction text information based on the prompt word segment, decompose the task to be executed into multiple subtasks, perform task planning on each subtask, generate multiple action strategies corresponding to each subtask and the corresponding specific parameter call information, so as to obtain multiple types of action plans for the task to be executed; Determine any one of the action plans as the current plan; Using a dedicated model call framework, send the specific parameter call information of the dedicated models of each action strategy in the current plan to the corresponding dedicated models respectively, and independently call the corresponding dedicated models for parsing operations to obtain the corresponding parsing operation results; Dynamically integrate the parsing operation results of each dedicated model to obtain the action sequence of the current plan; Using an evaluation model, determine the comprehensive evaluation value of each action plan based on the action sequence of each action plan respectively, and determine the action sequence of the action plan with the optimal comprehensive evaluation value as the optimal action sequence for executing the task to be executed.
[0027] Specifically, when using a task understanding and decision-making vertical large model to retrieve from the task planning knowledge base based on the instruction text information and scenario status information, a retrieval augmented generation architecture (RAG) is used to retrieve from the task planning knowledge base.
[0028] Using a dedicated model call framework, send the specific parameter call information of the dedicated models of each action strategy in the current plan to the corresponding dedicated models respectively, and independently call the corresponding dedicated models for parsing operations to obtain the corresponding parsing operation results. In the process of dynamically integrating the parsing operation results of each dedicated model to obtain the action sequence of the current plan, the dedicated model call framework will respond immediately after receiving the specific parameter call information corresponding to each action strategy in the current plan, identify the sequence of calls to each dedicated model according to the action sequence ID therein, and send the specific parameter call information corresponding to each dedicated model to each dedicated model, and independently call various dedicated model dynamic libraries such as the path planning, reconnaissance planning, and action parameter planning of the unmanned vehicle for parsing operations, and complete the information interaction between each dedicated model, and dynamically integrate the parsing operation results of each dedicated model to generate the action sequence of the current plan.
[0029] The dedicated model invocation framework is a framework that supports the flexible scheduling of dedicated models. After receiving the specific parameter invocation information, it can identify the sequence of invocation of each dedicated model based on the action sequence ID in it, and send the specific parameter invocation information corresponding to each dedicated model to each dedicated model. And there is information interaction among the dedicated models. The dedicated model invocation framework can also forward the parameter information generated by the dedicated model with a higher action sequence ID to the dedicated model with a lower action sequence ID that needs to receive this information. For example, it is identified from the specific parameter invocation information that the path planning model is called first, then the reconnaissance planning model is called, and then the path planning model is called again. Among them, it involves sending the result array point information after reconnaissance planning to the second path planning for calculation. The dedicated model invocation framework can distribute the parameter information generated by the task understanding and decision-making vertical large model to the three dedicated models, and forward the result generated by the reconnaissance planning model to the path planning model as input.
[0030] Using the evaluation model, based on the action sequences of each action plan respectively, determine the comprehensive evaluation value of each action plan. When determining the action sequence of the action plan with the optimal comprehensive evaluation value as the optimal action sequence for executing the task to be executed, a comprehensive evaluation is carried out from multiple dimensions such as time, energy consumption, completion rate, and safety. The evaluation function is: 。
[0031] Among them, is the total cost; is the weight of the task time; is the task time; is the weight of the energy consumption of the plan; is the energy consumption of the plan; is the weight of the task completion rate; is the task completion rate; is the weight of the path safety; is the path safety.
[0032] When the total cost is smaller, the generated action sequence is better.
[0033] As an alternative implementation, the optimal action sequence generation system for the unmanned vehicle based on the cooperation of the large model and the dedicated model further includes: an unmanned vehicle execution module, and the unmanned vehicle execution module is connected to the action sequence generation module.
[0034] The unmanned vehicle execution module is used to control the unmanned vehicle to execute the task to be executed based on the optimal action sequence.
[0035] As an alternative implementation, the GTCRN model and the Whisper speech recognition model are deployed in the speech processing module.
[0036] Based on the natural speech task instructions, determine the instruction text information, including: Use the GTCRN model to filter the noise of the natural speech task instructions to obtain the processed natural speech task instructions.
[0037] Specifically, the GTCRN model is a lightweight speech enhancement model. Noise filtering means filtering out mechanical noise, wind noise, and environmental echo background interference, making the processed natural speech task instructions clearer and more accurate.
[0038] Use the Whisper speech recognition model to transcribe the speech and recognize the instructions of the processed natural speech task instructions to obtain the instruction text information.
[0039] As an alternative implementation, the scenario status information includes: geographical environment information, task target information, and own status information.
[0040] The geographical environment information includes: scenario information and road network information. The scenario information is an urban scenario or a wild scenario, and the road network information is with a road network or without a road network.
[0041] The task target information includes: personnel, mobile platforms, and fixed buildings.
[0042] The own status information includes: driving speed, current battery level, and panoramic mirror availability information. The panoramic mirror availability information is that the panoramic mirror is available or the panoramic mirror is unavailable.
[0043] As an alternative implementation, the task planning knowledge base includes: relevant knowledge and historical cases corresponding to multiple instruction text information.
[0044] The relevant knowledge includes: task scenario introduction, definitions of various action strategies, basic parameters of the unmanned vehicle, and action strategies recommended under geographical environment information.
[0045] The historical cases include: action plans corresponding to relevant historical instruction text information.
[0046] As an alternative implementation, the specific parameter call information includes: timestamp, action sequence ID, path planning mode, driving speed, type of dedicated model called, longitude and latitude of the starting point of the unmanned vehicle, longitude and latitude of the ending point, sensor type, reconnaissance area range, and time limit requirements.
[0047] Specifically, the unmanned vehicle task understanding and decision-making vertical large model mainly receives multi-faceted three-dimensional information such as environmental information, target information, and platform data, understands and deconstructs the received task instructions, autonomously calls various dedicated models in the current scenario, automatically generates an action sequence, and then selects the optimal action sequence after evaluation by the evaluation model and sends it to the unmanned vehicle execution module to achieve the effective execution of the task to be executed.
[0048] By combining specific application scenarios, environmental information analysis, and task decision-making requirements, a high-quality meta-instruction dataset and a fine-tuning dataset for the unmanned vehicle task understanding and decision-making vertical large model are constructed. Based on the injection of domain expert experience, a (typical) task planning knowledge base such as a task scenario library, an action strategy library, a platform information library, and a decision-making knowledge graph is constructed, and a RAG retrieval-enhanced generation architecture is built to provide knowledge support for the unmanned vehicle's autonomous task understanding, task planning, and decision-making. Based on the fine-tuning instruction dataset, training and fine-tuning of the unmanned vehicle task understanding and decision-making vertical large model are carried out.
[0049] As an embodiment, the correspondence between tasks, action strategies, dedicated models, and response scenarios is shown in Table 1.
[0050] Table 1 Correspondence table of tasks, action strategies, dedicated models, and response scenarios
[0051] In an exemplary embodiment, a method for generating an optimal action sequence of an unmanned vehicle based on the collaboration of a large model and a dedicated model is provided, which is implemented based on the above-mentioned system for generating an optimal action sequence of an unmanned vehicle based on the collaboration of a large model and a dedicated model. The method for generating an optimal action sequence of an unmanned vehicle based on the collaboration of a large model and a dedicated model includes the following steps.
[0052] Obtain the natural speech task instruction of the occupant for the task to be executed by the unmanned vehicle.
[0053] Based on the natural speech task instruction, determine the instruction text information.
[0054] Obtain the scenario state information of the unmanned vehicle.
[0055] Using the task understanding and decision-making vertical large model, based on the instruction text information and the scenario state information, retrieve from the task planning knowledge base to obtain relevant knowledge and historical cases corresponding to the instruction text information, and arrange and splice the relevant knowledge and historical cases to generate a prompt word segment of the instruction text information.
[0056] Using the task understanding and decision-making vertical large model, based on the prompt word segment, perform task understanding on the instruction text information, decompose the task to be executed into multiple subtasks, perform task planning on each subtask, and generate multiple action strategies and corresponding specific parameter call information corresponding to each subtask, so as to obtain multiple types of action plans for the task to be executed.
[0057] Determine any one of the action plans as the current plan.
[0058] Using a dedicated model invocation framework, the specific parameter invocation information of the dedicated models for each action strategy in the current solution is sent to the corresponding dedicated models respectively, and the corresponding dedicated models are autonomously invoked for parsing and calculation to obtain the corresponding parsing and calculation results.
[0059] Dynamically integrate the parsing and calculation results of each dedicated model to obtain the action sequence of the current solution.
[0060] Using an evaluation model, based on the action sequences of each action plan respectively, determine the comprehensive evaluation value of each action plan, and determine the action sequence of the action plan with the optimal comprehensive evaluation value as the optimal action sequence for executing the task to be executed.
[0061] Using the method of the present application, the occupant issues a task to the driverless vehicle through a natural language instruction. The task understanding and decision-making vertical large model understands and plans the task information, adopts a RAG retrieval-enhanced generation architecture based on the task planning knowledge base to complete information dynamic integration, automatically generates various action decisions in the current scenario, and autonomously invokes various dedicated models such as the corresponding path planning, reconnaissance planning, and action parameter planning of the driverless vehicle for parsing and calculation. By constructing an evaluation model with multiple dimensions of safety, time, completion rate, and energy consumption, evaluate and optimize the final action sequence information generated by integrating the calculation results of the dedicated models, automatically generate the optimal action sequence of the task and send it to the driverless vehicle platform control and payload control to complete the autonomous execution of the task, realizing the full-process task chain of "occupant intelligent Q&A - task understanding, planning and decision-making - professional model calculation and optimization - optimal action sequence generation", improving the autonomous understanding and planning and decision-making ability of the driverless vehicle in complex and diverse application scenarios, enhancing the intelligent autonomous level of the driverless vehicle, and effectively improving the efficiency of one-control-multiple and one-control-cluster in future cluster application scenarios.
[0062] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0063] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0064] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present application.
Claims
1. An optimal action sequence generation system for unmanned vehicles based on the collaboration of a large model and a special model, characterized in that: The unmanned vehicle optimal action sequence generation system based on the collaboration of the large model and the special model includes: a user interaction module, a voice processing module, an information real-time perception module and an action sequence generation module; the user interaction module is connected to the voice processing module, and the voice processing module and the information real-time perception module are both connected to the action sequence generation module; The user interaction module is used to obtain the natural voice task instructions of the occupant for the task to be performed by the unmanned vehicle; The speech processing module is used to determine instruction text information based on the natural speech task instruction; The information real-time perception module is used to obtain scene status information of the unmanned vehicle; The action sequence generation module is deployed with a large vertical model of task understanding and decision-making, a task planning knowledge base, a dedicated model calling framework, dedicated models of multiple action strategies, a dedicated model dynamic library of multiple dedicated models, and an evaluation model; The action sequence generation module is used to: Using the vertical large model of task understanding and decision-making, based on the instruction text information and the scene state information, searching from the task planning knowledge base to obtain relevant knowledge and historical cases corresponding to the instruction text information, and arranging and splicing the relevant knowledge and the historical cases to generate prompt word segments of the instruction text information; Using the vertical large model of task understanding and decision-making, the instruction text information is understood based on the prompt word fragments, the task to be executed is decomposed into multiple subtasks, each subtask is task planned, and multiple action strategies and corresponding specific parameter call information corresponding to each subtask are generated, thereby obtaining multiple types of action plans for the task to be executed; Identify any one course of action as the current course of action; Using a dedicated model calling framework, specific parameter calling information of the dedicated model of each action strategy in the current solution is sent to the corresponding dedicated model respectively, and the corresponding dedicated model is autonomously called to perform analytical calculations to obtain corresponding analytical calculation results; Dynamically integrate the analytical calculation results of each dedicated model to obtain the action sequence of the current plan; The evaluation model is used to determine the comprehensive evaluation value of each action plan based on the action sequence of each action plan, and the action sequence of the action plan with the best comprehensive evaluation value is determined as the optimal action sequence for executing the task to be executed.
2. The optimal action sequence generation system for unmanned vehicles based on the collaboration of large models and special models according to claim 1 is characterized in that: The unmanned vehicle optimal action sequence generation system based on the collaboration of the large model and the special model further includes: an unmanned vehicle execution module, the unmanned vehicle execution module being connected to the action sequence generation module; The unmanned vehicle execution module is used to control the unmanned vehicle to execute the task to be executed based on the optimal action sequence.
3. The optimal action sequence generation system for unmanned vehicles based on the collaboration of large models and special models according to claim 1 is characterized in that: The speech processing module is deployed with a GTCRN model and a Whisper speech recognition model; Determining instruction text information based on the natural voice task instruction includes: Using the GTCRN model, performing noise filtering on the natural speech task instruction to obtain a processed natural speech task instruction; The Whisper speech recognition model is used to perform speech transcription and instruction recognition on the processed natural speech task instructions to obtain the instruction text information.
4. The optimal action sequence generation system for unmanned vehicles based on the collaboration of large models and special models according to claim 1 is characterized in that: The scene status information includes: geographical environment information, mission target information and own status information; The geographical environment information includes: scene information and road network information, the scene information is an urban scene or a wild scene, and the road network information is whether there is a road network or no road network; The mission target information includes: personnel, mobile platforms and fixed buildings; The own state information includes: driving speed, current battery power and surrounding mirror availability information, and the surrounding mirror availability information is that the surrounding mirror is available or the surrounding mirror is not available.
5. The optimal action sequence generation system for unmanned vehicles based on the collaboration of large models and special models according to claim 1 is characterized in that: The mission planning knowledge base includes: relevant knowledge and historical cases corresponding to a plurality of instruction text information; The relevant knowledge includes: introduction to mission scenarios, definitions of various action strategies, basic parameters of unmanned vehicles, and recommended action strategies under geographical environment information; The historical cases include: action plans corresponding to relevant historical instruction text information.
6. The optimal action sequence generation system for unmanned vehicles based on the collaboration of large models and special models according to claim 1 is characterized in that: The specific parameter calling information includes: timestamp, action sequence ID, path planning mode, driving speed, calling special model type, latitude and longitude of the unmanned vehicle starting point, longitude and longitude of the ending point, sensor type, reconnaissance area range and time limit requirements.
7. A method for generating an optimal action sequence of an unmanned vehicle based on the collaboration of a large model and a special model, which is implemented based on the optimal action sequence generation system of an unmanned vehicle based on the collaboration of a large model and a special model as claimed in any one of claims 1 to 6, characterized in that: The method for generating the optimal action sequence of an unmanned vehicle based on the collaboration of a large model and a special model includes: Obtaining natural voice task instructions from the occupant for the unmanned vehicle to perform tasks; Based on the natural voice task instruction, determining instruction text information; Obtain scene status information of the unmanned vehicle; Using the vertical large model of task understanding and decision-making, based on the instruction text information and the scene state information, searching from the task planning knowledge base to obtain relevant knowledge and historical cases corresponding to the instruction text information, and arranging and splicing the relevant knowledge and historical cases to generate prompt word segments of the instruction text information; Using the vertical large model of task understanding and decision-making, the instruction text information is understood based on the prompt word fragments, the task to be executed is decomposed into multiple subtasks, each subtask is task planned, and multiple action strategies and corresponding specific parameter call information corresponding to each subtask are generated, thereby obtaining multiple types of action plans for the task to be executed; Identify any one course of action as the current course of action; Using a dedicated model calling framework, specific parameter calling information of the dedicated model of each action strategy in the current solution is sent to the corresponding dedicated model respectively, and the corresponding dedicated model is autonomously called to perform analytical calculations to obtain corresponding analytical calculation results; Dynamically integrate the analytical calculation results of each dedicated model to obtain the action sequence of the current plan; The evaluation model is used to determine the comprehensive evaluation value of each action plan based on the action sequence of each action plan, and the action sequence of the action plan with the best comprehensive evaluation value is determined as the optimal action sequence for executing the task to be executed.
Citation Information
Cited By
Unmanned equipment cluster cooperative control method based on two-stage large model agent
CN120523114A
A collaborative control method for unmanned equipment swarm based on two-level large model agents
CN120523114B
Action plan generation method and related product
CN121278404A