A training guide control method and system based on a multi-modal large model

By integrating situational understanding and guidance decision-making through a training guidance and control method based on a multimodal large model, natural language interaction and adaptive guidance are achieved, solving the problems of low intelligence and low human-computer interaction efficiency in existing training guidance and control systems, and improving the intelligence and adaptability of the training process.

CN122173621BActive Publication Date: 2026-07-31JIANGXI LIANCHUANG PRECISION ELECTROMECHANICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGXI LIANCHUANG PRECISION ELECTROMECHANICS CO LTD
Filing Date
2026-05-12
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing training guidance and control systems rely on the human experience of trainers, resulting in rigid training processes, a lack of adaptability, one-sided data understanding, low efficiency of human-computer interaction, difficulty in knowledge updates, and delayed evaluation and feedback.

Method used

A training guidance and control method based on a multimodal large model is adopted, which integrates a situational understanding large model, a guidance and decision-making large model, and knowledge enhancement retrieval to achieve natural language interaction, intelligent situational understanding, and adaptive guidance and decision-making. The training process is monitored in real time and decision support is provided through an intelligent guidance module, an adaptive control module, and an intelligent situational display module.

Benefits of technology

It has improved the intelligence level of the training process, enhanced the relevance of training and the trainee experience, reduced the time for manual judgment, realized personalized guidance strategies and rapid adaptation to new training subjects, new equipment and new tactics, and improved the efficiency of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122173621B_ABST
    Figure CN122173621B_ABST
Patent Text Reader

Abstract

This invention discloses a training guidance and control method and system based on a multimodal large model, belonging to the field of training guidance and control technology. The method includes: loading the training scheme before training, parsing the training outline and scenario data, initializing the training terminal, training resources, and two major models—guidance decision-making and situation understanding—and knowledge enhancement retrieval; retrieving relevant training regulations, tactics, and historical cases from the guidance and control knowledge base and loading them into the model context. During training, the intelligent guidance and control module converts the natural language instructions of the guidance and control personnel into structured data, generates and issues structured instructions based on historical cases; the adaptive control module collects monitoring data in real time and issues adjustment instructions as needed; the intelligent situation display module integrates multi-source data, generates a situation description, and completes a visual display. This invention solves the problems of low intelligence, rigid training processes, incomplete data understanding, low human-computer interaction efficiency, and difficulty in knowledge updating in existing guidance and control methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of training and control technology, and in particular to a training and control method and system based on a multimodal large model. Background Technology

[0002] Typical applications of training and control mainly include three methods: traditional manual guidance, rule engine-based automatic guidance system, and traditional neural network-based auxiliary decision-making system.

[0003] Traditional manual guidance methods involve trainers manually setting training scenario parameters according to the training syllabus, relying on manual observation to judge the training situation, manually triggering guidance events, and recording the training process using paper or spreadsheets. Rule-based automated guidance systems control the state transitions of the training process through finite state machines. Trainers create training subjects, training plans, and set training scenarios on the training guidance software based on the training syllabus and historical training records. During the training implementation phase, trainers manually click training control commands and operation results on the training guidance software to advance the training process (start, pause, continue, end). Assisted decision-making systems based on traditional neural networks, building upon rule-based automated guidance systems, use CNN / RNN to extract features from training data, identify training situation categories based on classification models, and employ reinforcement learning to optimize simple decisions.

[0004] However, existing training guidance and control systems rely on the trainer's experience to create training scripts and set training scenario parameters, resulting in a rigid training process and a lack of adaptability. Furthermore, during training, only limited structured data is collected, lacking the ability to comprehensively analyze unstructured data (such as voice commands, operational details, and physiological states). Training performance evaluation depends on manual observation, which introduces subjective bias, and feedback is delayed (evaluation reports are typically only issued after training is completed). Therefore, existing guidance and control systems suffer from low levels of intelligence, rigid training processes, incomplete data understanding, low human-computer interaction efficiency, and difficulties in knowledge updating. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a training guidance and control method and system based on a multimodal large model, which aims to solve at least one technical problem existing in the background art.

[0006] The embodiments of the present invention are implemented as follows: On one hand, this invention proposes a training and control method based on a multimodal large model, applied to the business function layer. The business function layer is interconnected with the data access and interaction layer and the intelligent core layer. The business function layer includes an intelligent guidance module, an adaptive control module, and an intelligent situation display module. The data access and interaction layer includes training resources and trained terminals. The intelligent core layer includes a situation understanding large model, a guidance decision large model, and knowledge enhancement retrieval. The method includes: Load the training plan, parse the training outline and scenario data, initialize the training terminal and training resources, load the command and control decision-making model, situation understanding model and knowledge enhancement retrieval to complete the initialization of the large model, and retrieve relevant training orders, tactics and historical cases from the preset command and control knowledge base according to the loaded training plan and load them into the model context. During the training process, the intelligent instruction module receives natural language instructions from the instructor, uses the instruction decision-making big model to perform speech recognition, intent classification and entity extraction, converts natural language into structured data, retrieves historical case information from the instruction knowledge base through knowledge enhancement retrieval, and sends the obtained historical cases and structured data to the instruction decision-making big model for analysis to obtain multiple strategy options and evaluation results. Based on the constructed objective function, it generates structured instructions and sends them to each training node. The adaptive control module monitors and collects training status data and trainees' physiological status data in real time. Based on the guidance and decision-making big model, it analyzes the real-time collected data to determine whether adjustments are needed and sends the results to each training node. The intelligent situation display module integrates the collected structured data, training-related audio, training-related video, and physiological state data of trainees. It uses a situation understanding big data model to generate a natural language description of the situation and displays it in a pre-set interface with highlights and visualizations.

[0007] Furthermore, in the above-mentioned training guidance and control method based on a multimodal large model, the intelligent guidance module includes a natural language understanding unit, a guidance strategy reasoning unit, a multi-objective optimization decision-making unit, and a guidance knowledge base. During the training process, the intelligent instruction module receives natural language instructions from the instructor, uses the instruction decision-making model for speech recognition, intent classification, and entity extraction, converts the natural language into structured data, retrieves historical case information from the instruction knowledge base through knowledge-enhanced retrieval, and sends the obtained historical cases and structured data to the instruction decision-making model for analysis to obtain multiple strategy options and evaluation results. The steps of generating structured instructions based on the constructed objective function and sending them to each training node include: After receiving the natural language instructions from the director, the natural language processing unit performs speech recognition, intent classification, entity extraction, and context association, and sends the obtained instruction intent and entity information to the director's strategy reasoning unit for processing, understanding, and obtaining the director's intent. The guidance strategy reasoning unit analyzes the current situation based on the received guidance intent and entity information query, queries historical similar case information from the guidance knowledge base, and generates multiple strategy options using the guidance decision big model based on the analyzed situation information and the queried historical similar case information. It then evaluates the expected effects of the generated multiple strategy options and sends the strategy options and evaluation effects to the multi-objective optimization decision unit for processing. The multi-objective optimization decision unit constructs an objective function to select the optimal strategy based on the generated strategy scheme and evaluation results. It generates structured instruction based on constraints, performs conflict or security verification, and then issues the instruction to the relevant training nodes.

[0008] Furthermore, in the above-mentioned training-guided control method based on a multimodal large model, the objective function constructed by the multi-objective optimization decision unit is:

[0009] in: To achieve the training objective, To mitigate training safety risks, For resource consumption efficiency, For the trainees' experience, , , , These are the corresponding weight coefficients; The constraints include: equipment performance boundary constraints, training syllabus requirements constraints, safety red line constraints, and time window constraints.

[0010] Furthermore, in the above-mentioned training and control method based on a multimodal large model, the natural language processing unit uses the Whisper8 model for speech recognition and LlaMAf for intent classification and entity extraction. The inference unit for guiding strategies is fine-tuned based on ChatGLM3 or LlaMA2 to understand complex intentions and generate multiple strategy options and evaluate their effectiveness by combining the current situation and historical cases.

[0011] Furthermore, in the above-mentioned training-guided control method based on a multimodal large model, the adaptive control module includes a process state manager, an anomaly detection analyzer, and an adaptive adjustment decision-maker. The adaptive control module monitors and collects training status data and trainees' physiological status data in real time. Based on the guidance and decision-making model, it analyzes the collected data to determine whether adjustments are needed and then sends these adjustments to each training node. The steps include: After loading the training plan and training outline, the process status manager completes the training preparation work such as parsing the training outline, initializing the process status, configuring resource parameters and setting checkpoints, starts the training process, starts training status monitoring, and sends a start training command to the anomaly detection analyzer. The anomaly detection analyzer then starts monitoring the current training. During training, the process status manager collects the status data of each training node in real time, monitors the training progress, and sends the collected status data of each training node to the adaptive adjustment decision-maker for processing. The anomaly detection analyzer monitors for abnormal events. When the anomaly detection analyzer detects an anomaly in the current training process, it analyzes the anomaly and sends the anomaly and its cause to the adaptive adjustment decision-maker for processing. The adaptive adjustment decision-maker module evaluates the current training situation by integrating multi-source data based on the collected status data, current training status, abnormal data and causes of each training node, predicts the training development trend and identifies risk points. Based on the training development trend and risk points, it determines whether adjustment is needed. If adjustment is needed, it generates adjustment instructions and sends them to each training node. The process status manager completes the training process status transition according to the adjustment instructions.

[0012] Furthermore, in the above-mentioned training and control method based on a multimodal large model, the intelligent situation display module includes a multi-source data fusion unit, a situation semantic generator, an intelligent summary generator, and a visualization rendering layer. The intelligent situation display module integrates collected structured data, training-related audio, training-related video, and trainee physiological state data, uses a situation understanding big data model to generate a natural language description of the situation, and then highlights and visualizes it on a preset interface. The steps include: After receiving structured data, training-related speech, training-related video, and physiological state data of trainees, the multi-data fusion unit performs multi-source data format standardization, time alignment, outlier handling, and multi-source correlation fusion, and sends the fused data to the situational semantic generator for processing. The situation semantic generator, based on the received fused data, calls the situation understanding big model to identify key situations, determine the importance of the situations, predict development trends, and then generates a situation semantic description. The generated situation semantic information is then sent to the summary generator for processing. The intelligent summary generator filters information according to different levels based on the received situational semantic information, generates text summaries and voice briefings, prepares response content, and sends the generated situational semantic, summary and briefing information to the visualization rendering layer for rendering. The visualization rendering layer performs 2D / 3D map rendering, AR annotation overlay, chart generation, and warning prompts based on situational semantics, summary, and briefing information, and dynamically filters high-value information based on the director's focus. After the situation rendering is completed, the operator can make a voice query to the intelligent situation display module as needed, and the visual rendering layer will display the query results to the operator.

[0013] Furthermore, the above-mentioned training and control method based on a multimodal large model further includes: The training process ends when the preset goal is reached, training ends, or the termination condition is triggered. The entire training process data, guidance logs, and situation evolution records are automatically saved, and the training cases are stored in the guidance knowledge base for continuous learning and parameter fine-tuning of large models.

[0014] On the other hand, embodiments of the present invention propose a training and control system based on a multimodal large model, applied to the business function layer. The business function layer is interconnected with the data access and interaction layer and the intelligent core layer. The business function layer includes an intelligent guidance module, an adaptive control module, and an intelligent situation display module. The data access and interaction layer includes training resources and trained terminals. The intelligent core layer includes a situation understanding large model, a guidance decision large model, and knowledge enhancement retrieval. The system includes: The loading module is used to load the training plan, parse the training outline and scenario data, initialize the training terminal and training resources, load the command and control decision-making model, situation understanding model and knowledge enhancement retrieval to complete the initialization of the large model, and retrieve relevant training orders, tactics and historical cases from the preset command and control knowledge base according to the loaded training plan and load them into the model context. The sending module is used during the training process. The intelligent guidance module receives natural language instructions from the instructor, uses the guidance decision-making big model to perform speech recognition, intent classification and entity extraction, converts natural language into structured data, queries historical case information from the guidance knowledge base through knowledge enhancement retrieval, and sends the obtained historical cases and structured data to the guidance decision-making big model for analysis to obtain multiple strategy options and evaluation results. Based on the constructed objective function, it generates structured instructions and sends them to each training node. The distribution module is used by the adaptive control module to monitor and collect training status data and trainees' physiological status data in real time. Based on the guidance and decision-making big model, it analyzes the real-time collected data, determines whether adjustments are needed, and distributes the data to each training node. The display module is used by the intelligent situation display module to integrate the collected structured data, training-related voice, training-related video, and physiological state data of trainees, and use the situation understanding big model to generate a natural language description of the situation, and then highlight and visualize it on the preset interface.

[0015] In another aspect, embodiments of the present invention provide a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0016] In another aspect, embodiments of the present invention provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described above.

[0017] This invention, through the establishment of an intelligent guidance module, an adaptive control module, and an intelligent situation display module, along with a large-scale situation understanding model, a large-scale guidance decision-making model, and knowledge-enhanced retrieval, implements a training guidance and control method. This method supports interaction by the instructor using natural language, reducing operational complexity and improving human-computer interaction efficiency. The large-scale model provides intelligent auxiliary decision-making by fusing and analyzing structured data, voice, video, and other multi-source data generated during training, enabling instructors to complete complex guidance tasks with system support and reducing manual judgment time. The implemented personalized guidance strategies and training processes enhance the trainee experience and training relevance. After training, training cases and guidance experience are automatically stored in the database, forming reusable knowledge assets. The large-scale model architecture supports rapid adaptation to new training subjects, equipment, and tactics. It solves the problems of low intelligence, rigid training processes, incomplete data understanding, low human-computer interaction efficiency, and difficulty in knowledge updating in existing guidance and control methods. Attached Figure Description

[0018] Figure 1 This is a diagram illustrating the overall architecture of the training and control system in a training and control method based on a multimodal large model according to an embodiment of the present invention. Figure 2 This is a data flow diagram between training and control system modules in the training and control method based on a multimodal large model in the first embodiment of the present invention. Figure 3 This is a flowchart of the training and control method based on a multimodal large model in the first embodiment of the present invention; Figure 4 This is a schematic diagram of the composition structure of the intelligent guidance module in the training guidance and control method based on a multimodal large model in the first embodiment of the present invention; Figure 5This is a flowchart of the intelligent guidance module in the training guidance and control method based on a multimodal large model in the first embodiment of the present invention. Figure 6 This is a schematic diagram of the composition structure of the adaptive control module in the training and tuning control method based on a multimodal large model in the first embodiment of the present invention; Figure 7 This is a flowchart of the adaptive control module in the training and control method based on a multimodal large model in the first embodiment of the present invention. Figure 8 This is a schematic diagram of the composition structure of the intelligent situation display module in the training and control method based on a multimodal large model in the first embodiment of the present invention. Figure 9 This is a flowchart of the intelligent situation display module in the training and control method based on a multimodal large model in the first embodiment of the present invention. Figure 10 This is a structural block diagram of the training and control system based on a multimodal large model in the second embodiment of the present invention.

[0019] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation

[0020] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.

[0021] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0023] To address the problems of low intelligence, rigid training processes, incomplete data understanding, low human-computer interaction efficiency, and difficulty in knowledge updating in existing guidance and control systems, this invention proposes a training guidance and control method based on a large model of the guidance and control system. By integrating a large language model (LLM) as the core intelligent engine, it achieves natural language interaction, intelligent situational understanding, adaptive guidance and control decision-making, and dynamic process control.

[0024] Specifically, such as Figure 1 As shown, the training and control system based on a multimodal large model is designed with a layered architecture, including an intelligent core layer, a business function layer, and a data access and interaction layer.

[0025] The intelligent core layer, by integrating the Situational Understanding Large Model (SUM-LLM), the Direction and Decision Large Model (DDL-LLM), and Knowledge Enhancement Retrieval (RAG), is primarily responsible for data fusion processing and policy reasoning, thereby providing auxiliary decision-making for upper-layer business modules. Specifically, the Situational Understanding Large Model (SUM-LLM) is mainly responsible for multi-source data fusion, natural language understanding, and intent recognition; the Direction and Decision Large Model (DDL-LLM) is responsible for policy reasoning and process control decisions based on current training data and training scenarios; and Knowledge Enhancement Retrieval (RAG), by connecting the training doctrine library, tactical and operational method library, and historical case library, provides business knowledge to enhance retrieval functionality based on direction and decision-making and situational understanding.

[0026] The business function layer is used to implement business functions such as guidance and control, and situation recognition of the training and guidance control system based on a multimodal large model. It mainly includes the Intelligent Guidance and Control Module (IDM), the Adaptive Control Module (ACM), and the Intelligent Situation Display Module (ISDM).

[0027] The data access and interaction layer mainly realizes the external data resource interface and human-computer interaction interface of the training and control system based on the multimodal large model. It mainly includes training resource information such as training command library, tactical and operational method library and historical case library. It interacts with the trainees through the training terminal, receives training process data, and performs intelligent situation display and intelligent guidance and decision-making.

[0028] For example, such as Figure 2 As shown, the guidance and control process includes the system initialization phase, the real-time guidance and control phase, and the training completion and closed-loop optimization phase, wherein: System initialization phase: The training and control system based on the multimodal large model prepares the system environment and loads the model. The main workflow includes: loading the training scheme, parsing the training outline and scenario data, and initializing the training terminals and training resources; loading the large-scale training and control model (DDL-LLM), situation understanding model (SUM-LLM), process control model (FCL-LLM), and knowledge augmentation retrieval (RAG) to complete the initialization of the large-scale system model; and retrieving relevant training regulations, tactics, and historical cases from the training and control knowledge base according to the loaded training scheme and loading them into the model context.

[0029] Real-time guidance and control phase: During training, the intelligent guidance module, adaptive control module, and intelligent situation display module operate in parallel. The intelligent guidance and control module receives natural language speech or text commands from the instructor, uses the Directed Decision Model-LLM (DDL-LLM) for speech recognition, intent classification, and entity extraction, converts natural language into structured data, and generates multiple policy options using the DDL-LLM based on the current situation, training objective constraints, and historical cases, predicts the effects, and then outputs the optimal structured guidance command. After conflict / security verification, the command is distributed to each node. The adaptive control module monitors and collects the status data of each training node and the physiological status data of the trainees in real time, and determines the appropriate policy options based on the DDL-LLM. The sub-model process control large model (FCL-LLM) of the M) analyzes real-time data, determines whether there are any anomalies, and judges whether adjustments are needed based on the collected data and anomaly information. If adjustments are needed, it automatically generates adjustment strategies and adjustment instructions and issues them to each trained node. The intelligent situation display module integrates the collected structured data, voice, video and physiological indicators, and uses the situation understanding large model (SUM-LLM) to generate a natural language description of the situation. Based on the director's focus, it dynamically filters key information and highlights and visualizes it on 2D / 3D maps or AR interfaces.

[0030] Training End and Closed-Loop Optimization Phase: When the preset goal is reached, training ends, or the termination condition is triggered, the training process ends. The training guidance and control system based on the multimodal large model automatically saves the entire training process data, guidance logs, and situation evolution records. It also stores the training cases (including problems encountered and solutions) in the guidance knowledge base for continuous learning and parameter fine-tuning of the large model, enabling the system to self-evolve.

[0031] The following will describe in detail, with reference to specific embodiments and accompanying drawings, how to improve the existing guidance and control system, which suffers from low intelligence, rigid training process, one-sided data understanding, low human-computer interaction efficiency, and difficulty in knowledge updating.

[0032] Example 1 Please see Figure 1 The diagram illustrates a training and control method based on a multimodal large model in the first embodiment of the present invention. This method is applied to the business function layer, which is interconnected with the data access and interaction layer and the intelligent core layer. The business function layer includes an intelligent guidance module, an adaptive control module, and an intelligent situation display module. The data access and interaction layer includes training resources and trained terminals. The intelligent core layer includes a situation understanding large model, a guidance decision large model, and knowledge enhancement retrieval. The method includes steps S10 to S13.

[0033] The training and control method based on a multimodal large model is applied to the business function layer. This layer establishes bidirectional communication connections with both the data access and interaction layer and the intelligent core layer, ensuring the real-time performance and stability of data transmission. Specifically, the business function layer integrates three core functional modules: an intelligent guidance module, an adaptive control module, and an intelligent situation display module. These modules work collaboratively to complete the entire training and control process. The data access and interaction layer includes training resources and training terminals. Training resources can be configured according to actual training needs and include, but are not limited to, simulation training equipment, virtual training scenarios, and training data servers. Training terminals are the operating devices used by trainees, including computers, dedicated training terminals, and mobile terminals, used to receive guidance instructions and provide feedback on training status. The intelligent core layer serves as the intelligent support for the entire method, including a situation understanding large model, a guidance decision large model, and knowledge enhancement retrieval. These three components work together to provide intelligent decision-making and situation analysis capabilities for guidance and control. Step S10: Load the training plan, parse the training outline and scenario data, initialize the training terminal and training resources, load the command and control decision-making model, situation understanding model and knowledge enhancement retrieval to complete the initialization of the large model, and retrieve relevant training orders, tactics and historical cases from the preset command and control knowledge base according to the loaded training plan and load them into the model context.

[0034] The process begins with initial setup, loading a pre-defined training plan. This plan, prepared in advance based on factors such as training objectives, trainee size, and training duration, is stored in the local storage unit of the business function layer or on a related server. After loading, the training outline and scenario data within the plan are parsed. The training outline parsing primarily extracts core information such as training objectives, subjects, processes, and assessment criteria, clarifying the overall framework and requirements of the training. Scenario data parsing extracts setting information for the training scenario, including scenario environment parameters, configurations of both sides, and task node requirements, providing a scenario foundation for subsequent training.

[0035] Simultaneously with the completion of parsing, the training terminals and training resources in the data access interaction layer are initialized. Training terminal initialization includes starting the terminal device, loading training-related software, establishing communication connections between the terminal and the business function layer, and calibrating various terminal parameters to ensure that each training terminal can receive and provide feedback information normally. Training resource initialization includes starting the training resource device, allocating resource usage permissions, detecting resource operation status, and adjusting resource parameters to a state suitable for the training scheme to ensure that the training resources can meet the various requirements during the training process.

[0036] Subsequently, the large-scale guidance and decision-making model, situational understanding model, and knowledge enhancement retrieval component in the intelligent core layer are loaded to complete the initialization of the large model. The initialization of the large model includes loading model parameters, calibrating the model's operating environment, and debugging the interface between the model and the business function layer, to ensure that the large model can respond normally to the call requests of the business function layer.

[0037] Finally, based on the loaded training scheme, the knowledge-enhanced retrieval component retrieves training regulations, tactics, and historical cases related to this training from the pre-set command and control knowledge base. The command and control knowledge base pre-stores various standardized documents, mature tactical methods, and past training cases related to training. The knowledge-enhanced retrieval component uses keyword matching and semantic association to filter out content highly relevant to the training subjects and scenarios, and loads this retrieved content into the context of the command and control decision-making model and the situational understanding model, providing knowledge support for subsequent model decision-making and analysis.

[0038] In step S11, during the training process, the intelligent instruction module receives natural language instructions from the instructor, uses the instruction decision-making big model to perform speech recognition, intent classification and entity extraction, converts natural language into structured data, queries historical case information from the instruction knowledge base through knowledge enhancement retrieval, and sends the obtained historical cases and structured data to the instruction decision-making big model for analysis to obtain multiple strategy options and evaluation results. Based on the constructed objective function, structured instructions are generated and sent to each training node.

[0039] Specifically, during training, the intelligent guidance module, adaptive control module, and intelligent situation display module work in parallel. The intelligent guidance control module receives natural language voice or text commands from the instructor, uses the Directed Decision Model-LLM (DDL-LLM) for speech recognition, intent classification, and entity extraction, converts natural language into structured data, and generates multiple policy options using the DDL-LLM based on the current situation, training objective constraints, and historical cases, and estimates the effects. Then, it outputs the optimal structured guidance command, which is then sent to each node after conflict / security verification.

[0040] More specifically, such as Figures 4 to 5As shown, the intelligent instruction module includes a natural language understanding unit, an instruction strategy reasoning unit, a multi-objective optimization decision-making unit, and an instruction knowledge base. Each unit works together to achieve efficient processing of instruction commands.

[0041] The specific process is as follows: The Natural Language Processing (NLP) unit receives natural language instructions from the instructor in real time. The instructor can issue instructions via voice or text input. Upon receiving the instructions, the NLP unit first performs speech recognition, converting the spoken instructions into text. Then, it categorizes the text instructions by intent, clarifying the instructor's core needs, such as adjusting training difficulty, issuing tactical orders, checking training status, or terminating training. Simultaneously, it performs entity extraction, extracting key information from the instructions, including training nodes, trainees, equipment names, and time requirements. Furthermore, it performs contextual analysis, combining current training progress and previously issued instructions to ensure accurate understanding of the instructions. Finally, it sends the obtained instruction intent and entity information to the instructor's strategy reasoning unit for further processing, ultimately achieving a precise grasp of the instructor's intent.

[0042] After receiving the guidance intent and entity information from the natural language processing unit, the guidance strategy inference unit first analyzes the current training situation. This current situation includes information on various aspects such as the training progress, training status, resource usage, and trainee performance of each training node. This information is collected in real-time by the adaptive control module and fed back to the guidance strategy inference unit. Next, the guidance strategy inference unit queries the guidance knowledge base for historical cases similar to the current guidance intent and training situation. By comparing the processing methods and effects of historical cases, it provides a reference for the current guidance decision. Then, the guidance strategy inference unit inputs the analyzed situation information and the retrieved historical similar case information into the guidance decision-making model. The model performs in-depth analysis and inference based on this information, generating multiple guidance strategy options, each with a different implementation path and expected effect. Afterward, the guidance strategy inference unit evaluates the expected effects of the generated strategy options, including the probability of achieving the training objective, training safety risks, resource consumption, and trainee experience, forming detailed evaluation results. Finally, all strategy options and their corresponding evaluation results are sent to the multi-objective optimization decision-making unit for processing.

[0043] After receiving the strategy proposals and evaluation results from the guidance strategy inference unit, the multi-objective optimization decision unit constructs an objective function based on the core objective of this training. This objective function is then used to filter all strategy proposals, selecting the optimal guidance strategy. After selection, the optimal strategy is verified against preset constraints to ensure its feasibility and security. Upon successful verification, standardized structured guidance instructions are generated. These instructions use a unified data format, clearly specifying the recipient, execution content, execution time, and execution requirements. Finally, these structured guidance instructions are distributed to the relevant training nodes to ensure that each node can accurately understand and execute the instructions.

[0044] For example, the objective function constructed by the multi-objective optimization decision unit is:

[0045] in: To achieve the training objective, To mitigate training safety risks, For resource consumption efficiency, For the trainees' experience, , , , These are the corresponding weight coefficients; The constraints include: equipment performance boundary constraints, training syllabus requirements constraints, safety red line constraints, and time window constraints.

[0046] To further improve the accuracy and efficiency of directing instruction processing, this embodiment specifies the implementation methods of the natural language processing unit and the directing strategy reasoning unit. The natural language processing unit utilizes the Whisper8 model for speech recognition. The Whisper8 model possesses efficient speech-to-text capabilities, accurately recognizing the director's voice instructions. Even in environments with slight noise, it maintains high accuracy and supports the recognition of multiple languages ​​and accents, adapting to the usage habits of different directors. After completing speech recognition, the natural language processing unit performs intent classification and entity extraction based on the LlaMAf model. The LlaMAf model, after targeted fine-tuning, accurately understands the professional terminology and instruction logic of the directing domain, effectively distinguishes different types of directing intents, and accurately extracts key entity information from the instructions, providing reliable support for subsequent directing decisions. The inference unit for directing strategies is fine-tuned based on the ChatGLM3 or LlaMA2 model. The two models can be selected according to the actual application scenario. After fine-tuning, the model can fully understand the complex intentions in the directing domain, accurately analyze the correlation between the current training situation and historical cases, generate a variety of reasonable strategy schemes in combination with the actual situation of the current training, and accurately evaluate the expected effect of each strategy scheme to ensure the diversity and feasibility of the strategy schemes.

[0047] In step S12, the adaptive control module monitors and collects training status data and trainees' physiological status data in real time. Based on the guidance and decision-making big model, it analyzes the real-time collected data, determines whether adjustments are needed, and sends the results to each training node.

[0048] This module includes a process status manager, an anomaly detection analyzer, and an adaptive adjustment decision-maker. These components work together to achieve adaptive control of the training process. The specific implementation steps are as follows: The process status manager first completes pre-training preparations by loading the training plan and training outline. This includes parsing the training outline to clarify core information such as training subjects, processes, nodes, and assessment criteria; initializing the process status by setting the training process to its initial state and clarifying the sequence and execution requirements of each training node; configuring resource parameters by allocating corresponding training resources to each training node according to the requirements of the training plan and setting resource usage parameters; and setting checkpoints at key nodes in the training process to record the training status for later backtracking and recovery in case of anomalies. After completing these preparations, the process status manager starts the training process and simultaneously activates the training status monitoring function, collecting various data in real time during the training process and sending a start training command to the anomaly detection analyzer. Upon receiving the command, the anomaly detection analyzer activates its own monitoring function to begin real-time monitoring of the current training process.

[0049] During training, the process status manager continuously collects status data from each training node in real time. This status data includes operational data, training progress data, and equipment operation data of the training nodes. Simultaneously, it monitors the progress of the entire training process in real time to determine whether the training is proceeding according to the preset syllabus. After preliminary processing and standardization of the collected status data from each training node, the process status manager sends it to the adaptive adjustment decision-maker for further processing, providing data support for adaptive adjustment decisions.

[0050] The anomaly detection and analysis system continuously monitors for abnormal events during training. These events include training terminal malfunctions, abnormal training resources, trainee operational violations, significant delays in training progress, and security risks. When the anomaly detection and analysis system detects an anomaly in the current training process, it immediately conducts an in-depth analysis to determine the type, severity, scope of impact, and cause of the anomaly. For example, a training terminal malfunction might be caused by communication interruptions, hardware failures, or software crashes. After investigating relevant data and identifying the specific cause, the anomaly detection and analysis system sends detailed information about the anomaly and the identified cause to the adaptive adjustment decision-maker for further processing.

[0051] The adaptive adjustment decision-maker module receives status data and current training status from each training node sent by the process status manager, as well as abnormal data and their causes sent by the anomaly detection and analysis unit. It then comprehensively analyzes this multi-source data to fully assess the current training situation. During the assessment, the adaptive adjustment decision-maker combines the training objectives and preset evaluation criteria to predict the training trend and identify potential risks, such as training delays leading to failure to meet training objectives on time, or safety hazards causing training accidents. Based on the predicted training trend and identified risks, the adaptive adjustment decision-maker determines whether adjustments to the training process are necessary. If no adjustments are needed, the current training state is maintained, and the training progress continues to be monitored. If adjustments are required, corresponding adjustment instructions are generated based on the specific circumstances. These instructions include adjustments to training progress, resource allocation, training difficulty, and fault handling. After generating the adjustment instructions, they are sent to the relevant training nodes to ensure timely execution.

[0052] After receiving the adjustment instructions from the adaptive adjustment decision-maker, the process state manager completes the state transition of the training process according to the requirements of the adjustment instructions. For example, it adjusts the execution order of training nodes, modifies the training schedule, and reallocates training resources. At the same time, it updates the checkpoint information of the training process to ensure that the training process can proceed smoothly according to the adjusted plan until the training is completed.

[0053] In step S13, the intelligent situation display module integrates the collected structured data, training-related voice, training-related video, and physiological state data of the trainees, uses the situation understanding big model to generate a natural language description of the situation, and highlights and visualizes it on the preset interface.

[0054] The intelligent situation display module is responsible for fusing and visualizing various types of data during training, enabling instructors to monitor the training situation in real time. This module includes a multi-source data fusion unit, a situation semantic generator, an intelligent summary generator, and a visualization rendering layer. These components work collaboratively to achieve efficient display of situation information. The specific implementation steps are as follows: The multi-data fusion unit receives various types of data from different modules in real time, including structured data generated by the business function layer, relevant voice data and video data during training, as well as physiological status data of trainees. The structured data includes standardized data such as instructor instructions, training progress, and resource usage; training-related voice data includes instructor instructions, trainee feedback, and equipment operation sounds; training-related video data includes operation screens of each training node and training scene screens; and trainee physiological status data includes heart rate, blood pressure, and fatigue level, collected and transmitted by physiological monitoring devices worn by the trainees. After receiving this data, the multi-data fusion unit first performs multi-source data format standardization processing, converting data of different formats into a unified standard format to avoid processing anomalies caused by inconsistent data formats. Then, it performs time alignment processing, aligning data from different sources and at different collection times according to the timeline to ensure data temporal consistency. Next, it performs outlier handling, identifying and removing outliers in the data through preset outlier judgment rules, such as abnormal fluctuations in physiological state data or blurry or damaged frames in video data, to ensure data accuracy. Finally, it performs multi-source correlation fusion, performing correlation analysis on different types of data to uncover the inherent connections between data and form unified fused data. The fused data is then sent to the situational semantic generator for processing.

[0055] After receiving fused data from the multi-data fusion unit, the situation semantic generator invokes the situation understanding big model in the intelligent core layer to perform in-depth analysis of the fused data. Based on the fused data, the situation understanding big model identifies key situations in the current training process, such as whether the training progress is normal, whether there are abnormal events, and whether the trainees are in good condition; it judges the importance of each situation, distinguishes between core and secondary situations, and prioritizes situations that have a greater impact on the training objective; it predicts the development trend of the situations, such as the spread trend of abnormal events and the changing trend of training progress. Based on the above analysis results, the situation semantic generator generates an easy-to-understand natural language description of the situation, transforming the complex fused data into textual situation information that clearly reflects the overall situation of the current training. The generated situation semantic information is then sent to the intelligent summarizing generator for processing.

[0056] After receiving the situational semantic information from the situational semantic generator, the intelligent summarization generator filters the information according to preset hierarchical requirements. These hierarchies are divided into core, important, and general levels. The core level contains situational information that has the greatest impact on training decisions; the important level contains situational information that requires focused attention; and the general level contains routine situational information. Based on the filtering results, the intelligent summarization generator generates a text summary and a voice briefing. The text summary concisely summarizes the core situational and key information of the current training, while the voice briefing converts the text summary into audio format, allowing the instructor to understand the situational context even when reading the text. Simultaneously, it prepares response content to address subsequent queries from the instructor. After completing the above processing, the intelligent summarization generator sends the generated situational semantics, text summary, and voice briefing information to the visualization rendering layer for rendering and display.

[0057] After receiving situational semantics, summaries, and briefing information from the intelligent summary generator, the visualization rendering layer performs multi-form visualization rendering according to preset display rules and the instructor's usage habits. Specifically, this includes 2D / 3D map rendering, displaying training scenarios, training node locations, resource distribution, and other information in map form, intuitively reflecting the spatial distribution of training; AR annotation overlay, overlaying key situational information and anomaly alerts onto video footage or maps to highlight important information; chart generation, displaying data such as training progress, resource consumption, and trainee physiological status in line charts, bar charts, pie charts, etc., facilitating quick analysis of data trends by the instructor; and warning alerts, highlighting abnormal events and risk points to remind the instructor to pay attention and handle them promptly. Simultaneously, the visualization rendering layer can dynamically filter high-value information based on the instructor's focus. Instructors can set focus points through clicks, voice commands, etc., and the visualization rendering layer will prioritize displaying situational information related to those focus points, improving the instructor's work efficiency.

[0058] After the situational awareness rendering is complete, the instructor can make voice queries to the intelligent situational awareness display module according to actual needs. Query content includes training progress, trainee status, details of abnormal events, resource usage, etc. Upon receiving the instructor's query command, the natural language processing unit of the intelligent situational awareness display module performs speech recognition and intent parsing, converting the query command into a structured query request, which is then sent to the intelligent summary generator. The intelligent summary generator extracts the corresponding results from the processed situational information based on the query request and sends them to the visualization rendering layer. The visualization rendering layer displays the query results to the instructor in the form of text, charts, and audio, ensuring that the instructor can quickly obtain the information they need.

[0059] In addition, in some optional embodiments of the present invention, the training process is automatically terminated when the training reaches a preset goal, the training ends normally, or a termination condition is triggered.

[0060] The preset goals are the training completion standards set in advance in the training plan, such as trainees completing all training subjects or reaching the preset assessment scores; the termination conditions include emergency situations such as major safety accidents during training, failure of core training resources that cannot be recovered, and sudden health problems of trainees.

[0061] After the training process is completed, the business function layer automatically saves all training process data, command logs, and situation evolution records. The complete training data includes all collected data, command data, and feedback data during the training process. The command logs include all commands issued by the instructor and records of handled anomalies. The situation evolution records include changes in the situation during the training process and situation information at key points. This data will be stored in a designated database for easy retrieval, review, and analysis later.

[0062] Meanwhile, the training cases will be compiled and stored in the command and control knowledge base. The case content includes the training plan, process, problems encountered and their handling, training results, etc. These cases will serve as sample data for the subsequent large model to continuously learn and fine-tune parameters, helping the command and control decision-making large model and situational understanding large model to continuously optimize performance, improve the accuracy and adaptability of command and control decisions, and achieve iterative upgrades of the models.

[0063] In summary, the training and control method based on a multimodal large model in the above embodiments of the present invention, by setting up an intelligent guidance module, an adaptive control module, and an intelligent situation display module, as well as a situation understanding large model, a guidance decision large model, and knowledge enhancement retrieval, supports the interaction of the instructor using natural language, reducing operational complexity and improving human-computer interaction efficiency. The large model provides intelligent auxiliary decision-making by fusing and analyzing structured data, voice, video, and other multi-source data generated during training, enabling the instructor to complete complex guidance tasks with system support and reducing manual judgment time. The implemented personalized guidance strategies and training processes can improve the trainee experience and training relevance. After training, training cases and guidance experience are automatically stored in the database, forming reusable knowledge assets. The large model architecture supports rapid adaptation to new training subjects, new equipment, and new tactics. It solves the problems of low intelligence, rigid training processes, partial data understanding, low human-computer interaction efficiency, and difficulty in knowledge updating in existing guidance and control methods.

[0064] Example 2 Please see Figure 10 The figure shows a training and control system based on a multimodal large model proposed in the second embodiment of the present invention. This system is applied to the business function layer, which is interconnected with the data access and interaction layer and the intelligent core layer. The business function layer includes an intelligent guidance module, an adaptive control module, and an intelligent situation display module. The data access and interaction layer includes training resources and trained terminals. The intelligent core layer includes a situation understanding large model, a guidance decision large model, and knowledge enhancement retrieval. The system includes: The loading module 100 is used to load the training plan, parse the training outline and scenario data, initialize the training terminal and training resources, load the command and control decision-making model, situation understanding model and knowledge enhancement retrieval to complete the initialization of the large model, and retrieve relevant training orders, tactics and historical cases from the preset command and control knowledge base according to the loaded training plan and load them into the model context. The sending module 200 is used during the training process. The intelligent guidance module receives natural language instructions from the instructor, uses the guidance decision-making big model to perform speech recognition, intent classification and entity extraction, converts natural language into structured data, queries historical case information from the guidance knowledge base through knowledge enhancement retrieval, and sends the obtained historical cases and structured data to the guidance decision-making big model for analysis to obtain multiple strategy options and evaluation results. Based on the constructed objective function, it generates structured instructions and sends them to each training node. The distribution module 300 is used by the adaptive control module to monitor and collect training status data and trainees' physiological status data in real time. Based on the guidance and decision-making big model, it analyzes the real-time collected data to determine whether adjustments are needed and distributes the data to each training node. The display module 400 is used by the intelligent situation display module to integrate the collected structured data, training-related voice, training-related video and physiological state data of trainees, use the situation understanding big model to generate a situation natural language description, and highlight and visualize it on the preset interface.

[0065] The functions or operation steps implemented by the above modules are largely the same as those in the above method embodiments, and will not be repeated here.

[0066] Example 3 In another aspect, the present invention provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in Embodiment 1 above.

[0067] Example 4 In another aspect, the present invention provides an electronic device, the electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described in Embodiment 1 above.

[0068] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0069] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequential list of executable instructions for implementing logical functions, and can be embodied in any computer-readable storage medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable storage medium" can mean any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0070] More specific examples (a non-exhaustive list) of computer-readable storage media include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable storage media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0071] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0072] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0073] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A training guide control method based on a multi-modal large model, characterized in that, The method is applied to the business function layer, which is interconnected with the data access and interaction layer and the intelligent core layer. The business function layer includes an intelligent guidance module, an adaptive control module, and an intelligent situation display module. The data access and interaction layer includes training resources and trained terminals. The intelligent core layer includes a large-scale situation understanding model, a large-scale guidance and decision-making model, and knowledge enhancement retrieval. The method includes: Load the training plan, parse the training outline and scenario data, initialize the training terminal and training resources, load the command and control decision-making model, situation understanding model and knowledge enhancement retrieval to complete the initialization of the large model, and retrieve relevant training orders, tactics and historical cases from the preset command and control knowledge base according to the loaded training plan and load them into the model context. During the training process, the intelligent instruction module receives natural language instructions from the instructor, uses the instruction decision-making big model to perform speech recognition, intent classification and entity extraction, converts natural language into structured data, retrieves historical case information from the instruction knowledge base through knowledge enhancement retrieval, and sends the obtained historical cases and structured data to the instruction decision-making big model for analysis to obtain multiple strategy options and evaluation results. Based on the constructed objective function, it generates structured instructions and sends them to each training node. The adaptive control module monitors and collects training status data and trainees' physiological status data in real time. Based on the guidance and decision-making big model, it analyzes the real-time collected data to determine whether adjustments are needed and sends the results to each training node. The intelligent situation display module integrates the collected structured data, training-related audio, training-related video, and physiological state data of trainees, uses the situation understanding big model to generate a natural language description of the situation, and highlights and visualizes it on the preset interface accordingly. When the preset goal is reached, training ends, or the termination condition is triggered, the training process ends. The entire training process data, guidance log, and situation evolution record are automatically saved, and the training cases are stored in the guidance knowledge base for continuous learning and parameter fine-tuning of large models. The intelligent guidance module includes a natural language understanding unit, a guidance strategy reasoning unit, a multi-objective optimization decision-making unit, and a guidance knowledge base; The multi-objective optimization decision unit constructs an objective function to select the optimal strategy based on the generated strategy scheme and evaluation results. It generates structured instruction based on constraints, performs conflict or security verification, and then issues the instruction to the relevant training nodes. The objective function constructed by the multi-objective optimization decision unit is: in: To achieve the training objective, To mitigate training safety risks, For resource consumption efficiency, For the trainees' experience, , , , These are the corresponding weight coefficients; The constraints include: equipment performance boundary constraints, training syllabus requirements constraints, safety red line constraints, and time window constraints; The Natural Language Processing Unit uses the Whisper8 model for speech recognition and LlaMAf for intent classification and entity extraction. The inference unit for guiding strategies is fine-tuned based on ChatGLM3 or LlaMA2 to understand complex intentions and generate multiple strategy options and evaluate their effectiveness by combining the current situation and historical cases.

2. The training and control method based on a multimodal large model according to claim 1, characterized in that, During the training process, the intelligent instruction module receives natural language instructions from the instructor, uses the instruction decision-making model for speech recognition, intent classification, and entity extraction, converts the natural language into structured data, retrieves historical case information from the instruction knowledge base through knowledge-enhanced retrieval, and sends the obtained historical cases and structured data to the instruction decision-making model for analysis to obtain multiple strategy options and evaluation results. The steps of generating structured instructions based on the constructed objective function and sending them to each training node include: After receiving the natural language instructions from the director, the natural language processing unit performs speech recognition, intent classification, entity extraction, and context association, and sends the obtained instruction intent and entity information to the director's strategy reasoning unit for processing, understanding, and obtaining the director's intent. The guidance strategy reasoning unit analyzes the current situation based on the received guidance intent and entity information query, queries historical similar case information from the guidance knowledge base, and generates multiple strategy options using the guidance decision big model based on the analyzed situation information and the queried historical similar case information. It then evaluates the expected effects of the generated multiple strategy options and sends the strategy options and evaluation results to the multi-objective optimization decision unit for processing.

3. The training steering control method based on a multi-modal large model according to claim 1, characterized in that, The adaptive control module includes a process status manager, an anomaly detection analyzer, and an adaptive adjustment decision-maker. The adaptive control module monitors and collects training status data and trainees' physiological status data in real time. Based on the guidance and decision-making model, it analyzes the collected data to determine whether adjustments are needed and then sends these adjustments to each training node. The steps include: After loading the training plan and training outline, the process status manager completes the training preparation work such as parsing the training outline, initializing the process status, configuring resource parameters and setting checkpoints, starts the training process, starts training status monitoring, and sends a start training command to the anomaly detection analyzer. The anomaly detection analyzer then starts monitoring the current training. During training, the process status manager collects the status data of each training node in real time, monitors the training progress, and sends the collected status data of each training node to the adaptive adjustment decision-maker for processing. The anomaly detection analyzer monitors for abnormal events. When the anomaly detection analyzer detects an anomaly in the current training process, it analyzes the anomaly and sends the anomaly and its cause to the adaptive adjustment decision-maker for processing. The adaptive adjustment decision-maker module evaluates the current training situation by integrating multi-source data based on the collected status data, current training status, abnormal data and causes of each training node, predicts the training development trend and identifies risk points. Based on the training development trend and risk points, it determines whether adjustment is needed. If adjustment is needed, it generates adjustment instructions and sends them to each training node. The process status manager completes the training process status transition according to the adjustment instructions.

4. The training and tuning control method based on a multi-modal large model according to claim 1, characterized in that, The intelligent situation display module includes a multi-source data fusion unit, a situation semantic generator, an intelligent summary generator, and a visualization rendering layer. The intelligent situation display module integrates collected structured data, training-related audio, training-related video, and trainee physiological state data, uses a situation understanding big data model to generate a natural language description of the situation, and then highlights and visualizes it on a preset interface. The steps include: After receiving structured data, training-related speech, training-related video, and physiological state data of trainees, the multi-data fusion unit performs multi-source data format standardization, time alignment, outlier handling, and multi-source correlation fusion, and sends the fused data to the situational semantic generator for processing. The situation semantic generator, based on the received fused data, calls the situation understanding big model to identify key situations, determine the importance of the situations, predict development trends, and then generates a situation semantic description. The generated situation semantic information is then sent to the summary generator for processing. The intelligent summary generator filters information according to different levels based on the received situational semantic information, generates text summaries and voice briefings, prepares response content, and sends the generated situational semantic, summary and briefing information to the visualization rendering layer for rendering. The visualization rendering layer performs 2D / 3D map rendering, AR annotation overlay, chart generation, and warning prompts based on situational semantics, summary, and briefing information, and dynamically filters high-value information based on the director's focus. After the situation rendering is completed, the operator can make a voice query to the intelligent situation display module as needed, and the visual rendering layer will display the query results to the operator.

5. A training guide control system based on a multi-modal large model, characterized by, The system is used to implement the training and control method based on a multimodal large model as described in any one of claims 1 to 4, and is applied to the business function layer. The business function layer is interconnected with the data access and interaction layer and the intelligent core layer. The business function layer includes an intelligent guidance module, an adaptive control module, and an intelligent situation display module. The data access and interaction layer includes training resources and trained terminals. The intelligent core layer includes a situation understanding large model, a guidance decision large model, and knowledge enhancement retrieval. The system comprises: The loading module is used to load the training plan, parse the training outline and scenario data, initialize the training terminal and training resources, load the command and control decision-making model, situation understanding model and knowledge enhancement retrieval to complete the initialization of the large model, and retrieve relevant training orders, tactics and historical cases from the preset command and control knowledge base according to the loaded training plan and load them into the model context. The sending module is used during the training process. The intelligent guidance module receives natural language instructions from the instructor, uses the guidance decision-making big model to perform speech recognition, intent classification and entity extraction, converts natural language into structured data, queries historical case information from the guidance knowledge base through knowledge enhancement retrieval, and sends the obtained historical cases and structured data to the guidance decision-making big model for analysis to obtain multiple strategy options and evaluation results. Based on the constructed objective function, it generates structured instructions and sends them to each training node. The distribution module is used by the adaptive control module to monitor and collect training status data and trainees' physiological status data in real time. Based on the guidance and decision-making big model, it analyzes the real-time collected data, determines whether adjustments are needed, and distributes the data to each training node. The display module is used by the intelligent situation display module to integrate the collected structured data, training-related voice, training-related video, and physiological state data of trainees, and use the situation understanding big model to generate a natural language description of the situation, and then highlight and visualize it on the preset interface.

6. A readable storage medium, having stored thereon a computer program, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1 to 4.

7. An electronic device, comprising: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the steps of the method as described in any one of claims 1 to 4.