Model training method and device for multi-modal large model and storage medium

By using a fusion training method combining interactive instruction sets and domain knowledge graphs in a multimodal large model, the problem of insufficient professionalism in the construction safety large model is solved, and the understanding of industry standards in the field of physical engineering and the accurate analysis of construction safety are realized.

CN121998096APending Publication Date: 2026-05-08ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG DAHUA TECH CO LTD
Filing Date
2026-01-28
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

The large-scale construction safety model lacks professionalism in construction safety scenarios, resulting in insufficient understanding of industry standards, low accuracy in locating safety risks, and misleading opinions generated.

Method used

The multimodal large model is trained using an interactive instruction set based on domain knowledge graphs. By retrieving domain knowledge content corresponding to the interactive instructions from the target knowledge base and integrating it with the initial prompts, the model is trained to ensure that it can understand industry standards in the field of physical engineering and output results that meet construction safety requirements.

Benefits of technology

It enhances the professionalism of multimodal large models in construction safety scenarios, enabling them to effectively learn and understand industry standards in the field of physical engineering, output accurate analysis results and rectification suggestions applicable to construction safety, and reduce false alarms and missed alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998096A_ABST
    Figure CN121998096A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device for a multi-modal large model and a storage medium, and the method comprises the steps: retrieving domain knowledge content corresponding to each interaction instruction in an interaction instruction set from a target knowledge base under the condition of carrying out model training on the multi-modal large model through an interaction instruction set; fusing the domain knowledge content corresponding to each interaction instruction and the initial prompt content of each interaction instruction to obtain fused prompt content of each interaction instruction; and performing model training on the multi-modal large model by using the fusion prompt content of each interaction instruction and the expected output content of each interaction instruction to obtain a trained modal large model. Through the method and the device, the problem of insufficient professionality of a construction safety large model in a construction safety scene in related technologies is solved, and the effect of improving the professionality of the large model is further achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to a method, apparatus, and storage medium for training a multimodal large model. Background Technology

[0002] With the continuous advancement of urbanization and investment in infrastructure construction, the construction industry continues to expand. However, with the frequent occurrence of construction activities, the situation regarding construction safety remains severe. Traditional manual inspection methods suffer from prominent problems such as low efficiency, strong subjectivity, incomplete coverage, and delayed response, making it difficult to meet the needs of modern smart construction site management that requires refined and intelligent processes.

[0003] In related technologies, intelligent construction site detection systems based on large-scale artificial intelligence models can automate the identification, location, and early warning of common safety risks, improving the efficiency of safety hazard discovery and the speed of management closure. However, these systems typically employ general-purpose multimodal large-scale models, which suffer from insufficient understanding of industry standards, low accuracy in safety risk location, and the generation of misleading opinions in specialized engineering scenarios. Therefore, it is evident that large-scale construction safety models in related technologies lack sufficient professionalism for specific construction safety scenarios. Summary of the Invention

[0004] This application provides a method, apparatus, and storage medium for training a multimodal large model, in order to at least address the problem of insufficient professionalism in construction safety scenarios in related technologies.

[0005] According to one aspect of the embodiments of this application, a model training method for a multimodal large model is provided, comprising: when training the multimodal large model using an interaction instruction set, retrieving domain knowledge content corresponding to each interaction instruction in the interaction instruction set from a target knowledge base, wherein the target knowledge base is constructed based on a domain knowledge graph, the domain knowledge graph is used to record domain knowledge content in the field of physical engineering, and each interaction instruction includes initial prompt content and expected output content; fusing the domain knowledge content corresponding to each interaction instruction and the initial prompt content of each interaction instruction to obtain a fused prompt content for each interaction instruction; and using the fused prompt content of each interaction instruction and the expected output content of each interaction instruction to train the multimodal large model to obtain the trained multimodal large model.

[0006] According to another aspect of the embodiments of this application, a model training apparatus for a multimodal large model is also provided, comprising: a retrieval unit, configured to retrieve domain knowledge content corresponding to each interactive instruction in the interactive instruction set from a target knowledge base when training the multimodal large model using an interactive instruction set, wherein the target knowledge base is constructed based on a domain knowledge graph, the domain knowledge graph being used to record domain knowledge content in the field of physical engineering, and each interactive instruction including initial prompt content and expected output content; a fusion unit, configured to fuse the domain knowledge content corresponding to each interactive instruction and the initial prompt content of each interactive instruction to obtain fused prompt content for each interactive instruction; and a training unit, configured to train the multimodal large model using the fused prompt content of each interactive instruction and the expected output content of each interactive instruction to obtain the trained multimodal large model.

[0007] In one exemplary embodiment, the apparatus further includes: a construction unit, configured to construct the domain knowledge graph based on domain knowledge source data of the physical engineering domain before retrieving domain knowledge content corresponding to each interaction instruction in the interaction instruction set from the target knowledge base, wherein the domain knowledge source data is data collected from domain knowledge sources to describe domain knowledge in the physical engineering domain, and the domain knowledge source data includes at least one of the following: at least one level of standard text; standard design diagrams; abnormal state record data; and experience description data.

[0008] In an exemplary embodiment, the construction unit includes: an extraction module, configured to extract a set of domain knowledge triples from the domain knowledge source data, wherein each domain knowledge triple in the set of domain knowledge triples is a triple with an abnormal state type, a state baseline entry, and a state correction scheme as entities; and a construction module, configured to construct a knowledge graph using each entity in each domain knowledge triple as a node to obtain the domain knowledge graph.

[0009] In one exemplary embodiment, the apparatus further includes: an interaction unit, configured to interact with an interactive language model using each job scene image in a set of job scene images and a step-by-step prompt template before fusing the domain knowledge content corresponding to each interaction instruction and the initial prompt content of each interaction instruction, to obtain a multi-turn interaction instruction set; and a generation unit, configured to generate the interaction instruction set based on the multi-turn interaction instruction set, wherein each interaction instruction is at least a portion of an interaction instruction in a multi-turn interaction instruction set; wherein each job scene image is a real job scene image of an engineering operation in a physical engineering field, or a synthetic job scene image of an engineering operation in a physical engineering field; wherein the step-by-step prompt template is used to interact with the interactive language model step-by-step according to a target prompt chain, and the prompt nodes in the target prompt chain include at least one of the following: job scene description information, abnormal state location indication information, abnormal state type, abnormal level, state baseline entry, and state correction scheme.

[0010] In one exemplary embodiment, the apparatus further includes: a removal unit, configured to remove multi-turn interaction instructions from the multi-turn interaction instruction set that satisfy at least one of the following screening conditions before generating the interaction instruction set based on the multi-turn interaction instruction set, thereby obtaining an updated multi-turn interaction instruction set: the included interaction content does not match the prompt node in the target prompt chain; the semantic similarity between the included state correction scheme and the corresponding state baseline entry is lower than a similarity threshold; or it was sampled during the sampling review process and the sampling review result was that the review failed.

[0011] In an exemplary embodiment, the retrieval unit includes: a filtering module, configured to filter M domain knowledge contents from the target knowledge base for each interactive instruction according to the order of high to low vector similarity between the encoding vector of each domain knowledge content in the target knowledge base and the encoding vector of the initial prompt content of each interactive instruction, to obtain domain knowledge content corresponding to each interactive instruction; wherein, the encoding vector of each domain knowledge content is a vector obtained by encoding each domain knowledge content, the encoding vector of the initial prompt content of each interactive instruction is a vector obtained by encoding the initial prompt content of each interactive instruction, and M is a positive integer greater than or equal to 2; wherein, among the M domain knowledge contents filtered for each interactive instruction, the first to Nth domain knowledge contents are domain knowledge contents used in the first model training stage of the multimodal large model, the (N+1)th domain knowledge content is domain knowledge contents used in the second model training stage of the multimodal large model, N is a positive integer greater than or equal to 1 and less than M, and the first model training stage is earlier than the second model training stage.

[0012] In an exemplary embodiment, the fusion unit includes: an adding module, configured to add target prompt content and domain knowledge content corresponding to each interaction instruction to the initial prompt content of each interaction instruction, to obtain fused prompt content with each interaction instruction, wherein the target prompt content is used to prompt the multimodal big model to generate output content corresponding to the initial prompt content of each interaction instruction based on the domain knowledge content corresponding to each interaction instruction.

[0013] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to perform the steps in any of the above method embodiments when executed by a processor.

[0014] According to another aspect of the embodiments of this application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the steps in any of the method embodiments described above.

[0015] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to perform the steps of any of the above method embodiments through the computer program.

[0016] This application describes a method for training a multimodal large-scale model using an interaction instruction set. The method involves retrieving domain knowledge content corresponding to each interaction instruction from a target knowledge base (built on a domain knowledge graph that records domain knowledge content in the field of physical engineering). Each interaction instruction includes initial prompts and expected outputs. The domain knowledge content corresponding to each interaction instruction and the initial prompts are then fused to obtain a fused prompt for each instruction. The fused prompts and expected outputs of each interaction instruction are then used to train the multimodal large-scale model, resulting in a trained model. Because the multimodal large-scale model is trained using domain knowledge content from a domain knowledge graph that records domain knowledge in the field of physical engineering, the trained model can effectively learn and understand industry standards in the field of physical engineering and output construction safety results applicable to the field of physical engineering based on the requirements of the interaction instructions. Therefore, this method addresses the problem of insufficient professionalism in construction safety scenarios in related technologies, thereby improving the professionalism of large-scale models. Attached Figure Description

[0017] Figure 1 This is a schematic diagram illustrating an application scenario of a multimodal large model training method according to an embodiment of this application;

[0018] Figure 2 This is a flowchart illustrating an optional multimodal large model training method according to an embodiment of this application;

[0019] Figure 3 This is a schematic diagram of an optional multimodal large model training method according to an embodiment of this application;

[0020] Figure 4 This is a flowchart illustrating another optional multimodal large model training method according to an embodiment of this application;

[0021] Figure 5 This is a structural block diagram of an optional multimodal large model training device according to an embodiment of this application;

[0022] Figure 6 This is a computer system architecture block diagram of an optional electronic device according to an embodiment of this application. Detailed Implementation

[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0025] According to one aspect of the embodiments of this application, a method for training a multimodal large model is provided. Optionally, in this embodiment, the above-described method for training a multimodal large model can be applied, but is not limited to, to applications such as... Figure 1 The hardware environment shown includes terminal device 102 and server 104. Server 104 can be connected to terminal device 102 via a network and can be used to provide services (e.g., application services, etc.) to terminal device 102 or clients installed on terminal device 102. A database can be set up on server 104 or independently of server 104 to provide data storage services for server 104.

[0026] The aforementioned network may include, but is not limited to, at least one of the following: wired network and wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network (WAN), metropolitan area network (MAN), and local area network (LAN). The aforementioned wireless network may include, but is not limited to, at least one of the following: Wireless Fidelity (WIFI) and Bluetooth. Terminal device 102 may be, but is not limited to, a personal computer (PC), mobile phone, tablet computer, etc. Server 104 may be, but is not limited to, a cloud server, server cluster, or other server types.

[0027] The multimodal large model training method of this application embodiment can be executed by server 104, terminal device 102, or jointly by server 104 and terminal device 102. Alternatively, the multimodal large model training method of this application embodiment can be executed by a client installed on the terminal device 102.

[0028] Taking the multimodal large model training method of this embodiment executed by terminal device 102 as an example, Figure 2 This is a flowchart illustrating an optional multimodal large model training method according to an embodiment of this application, as shown below. Figure 2 As shown, the process of this method may include the following steps:

[0029] Step S202: When training a multimodal large model using an interactive instruction set, retrieve the domain knowledge content corresponding to each interactive instruction in the interactive instruction set from the target knowledge base. The target knowledge base is constructed based on a domain knowledge graph, which is used to record domain knowledge content in the field of physical engineering. Each interactive instruction includes initial prompt content and expected output content.

[0030] Step S204: Merge the domain knowledge content corresponding to each interactive instruction and the initial prompt content of each interactive instruction to obtain the merged prompt content of each interactive instruction;

[0031] Step S206: Use the fused prompt content of each interactive instruction and the expected output content of each interactive instruction to train the multimodal large model, and obtain the trained modal large model.

[0032] The multimodal large model training method in this embodiment can be applied to the field of computer technology, specifically to the scenario of training multimodal large models in the field of physics and engineering.

[0033] With the continuous advancement of urbanization and investment in infrastructure construction, the construction industry continues to expand. However, with the frequent occurrence of construction activities, the situation regarding construction safety remains severe. Traditional manual inspection methods suffer from prominent problems such as low efficiency, strong subjectivity, incomplete coverage, and delayed response, making it difficult to meet the needs of modern smart construction site management that requires refined and intelligent processes.

[0034] In related technologies, intelligent construction site detection systems based on large-scale AI models can automatically identify, locate, and issue early warnings for common safety risks (such as "unprotected edges," "illegally stacked materials," "unpaved or damaged road surfaces," and "personnel not wearing safety protective equipment"), improving the efficiency of safety hazard discovery and the speed of management closure. However, these systems typically employ general-purpose multimodal models. While these models possess excellent overall generalization and reasoning capabilities, they suffer from insufficient understanding of industry standards, low accuracy in safety risk location, and the generation of misleading opinions when directly applied to specialized engineering scenarios. Therefore, it is evident that large-scale construction safety models in related technologies lack sufficient professionalism for specific construction safety scenarios.

[0035] To at least partially solve the above-mentioned technical problems, in this embodiment, an interactive instruction set is used to train a multimodal large model based on the domain knowledge content of a domain knowledge graph. The domain knowledge graph is used to record the domain knowledge content of the physical engineering field. The trained multimodal large model can effectively learn and understand the industry standards of the physical engineering field, and can output construction safety results applicable to the physical engineering field based on the requirements of the interactive instructions. In this way, a multimodal large model with sufficient professionalism in the physical engineering field can be trained, thereby solving the above-mentioned technical problems.

[0036] In this embodiment, when training a multimodal large model using an interaction instruction set, domain knowledge content corresponding to each interaction instruction in the interaction instruction set is retrieved from the target knowledge base. The target knowledge base is constructed based on a domain knowledge graph, which records domain knowledge content in the field of physical engineering. Each interaction instruction includes initial prompts and expected outputs. Optionally, the domain knowledge graph can include information such as basic specifications, standard requirements, accident cases, and expert experience in the physical engineering industry. Furthermore, it can establish the inherent connections between these information through relationships between entities, forming a structured and queryable knowledge system.

[0037] The interactive instructions include initial prompts and expected outputs. The prompts may include a problem scenario and a question. For example, an initial prompt may include a real or generated construction image or video in the field of physical engineering and questions related to the construction image (such as whether there are safety risks in the scenario, and suggestions for rectifying safety risks). The expected outputs may be the standard answers corresponding to the initial prompts. The expected outputs may be manually annotated for the initial prompts or generated by a model and then reviewed. In this embodiment, there are no restrictions on this.

[0038] The domain knowledge content corresponding to each interaction command and the initial prompt content of each interaction command are fused to obtain the fused prompt content for each interaction command. Here, fusing the domain knowledge content corresponding to the interaction command and the initial prompt content can guide the model to output analysis results that are more in line with the standards of the physical engineering field. Optionally, the fusion can adopt methods such as direct splicing, feature embedding, or self-attention mechanisms to ensure that the multimodal model can refer to the domain knowledge content to make corresponding outputs for the interaction commands. For example, when the initial prompt content is related to the command of "rebar tying", the fused prompt content will include not only image descriptions but also information related to "construction safety requirements" and "rebar tying specifications" to ensure that the multimodal model follows industry standards during analysis and clearly indicates whether it complies with the specifications, the reasons for violations, and recommended rectification measures in the output results.

[0039] The multimodal large model is trained using the fused prompts and expected outputs of each interaction command. Here, the expected outputs are based on the specifications recorded in the domain knowledge graph, containing accurate safety analysis results, hazard identification, and specific remedial measures recommendations, providing a clear learning objective for the multimodal model. Optionally, the training process can be implemented using iterative optimization algorithms (such as gradient descent). In each training round, the multimodal model predicts the output based on the fused prompts, and then adjusts the model parameters to reduce the difference between the predicted results and the expected outputs. Through iterative learning, the multimodal model can gradually learn the specifications of the physical engineering domain and, when outputting, strictly adhere to the guidance of the domain knowledge graph, providing professional and accurate safety analysis results and remedial recommendations.

[0040] Optionally, the underlying model of the multimodal large model can be the Qwen3 Vision-Language Model (Qwen3-VL), such as... Figure 3As shown, Qwen3-VL adopts a three-module architecture, consisting of a visual encoder, an MLP-based visual-language fusion unit, and a large language model core. The visual encoder uses a Sigmoid Loss for Image-Pre-training (SigLIP-2) architecture as its visual backbone, enabling efficient handling of dynamic input resolutions. It integrates 2D Rotary Position Embedding (2D-RoPE) technology and follows the Coordinated Multiple Points Transmission and Reception (CoMP) method to interpolate the absolute position embeddings based on the input size (following the CoMP method). The visual-language fusion unit can be designed using a two-layer multilayer perceptron (MLP). Its core function is to compress each 2x2 spatial feature block output by the visual encoder into a single visual label and project it onto the large language model. The Model (LLM) module aligns the dimensions of the hidden layers. In addition, it can be equipped with a dedicated fusion unit to support the DeepStack mechanism within the LLM, enabling deep fusion of multimodal information at different depths of the model. The large language model can be built on the Qwen3 series backbone network and can provide variants with various parameter scales, such as dense versions and expert hybrid versions.

[0041] In addition, to improve the professional capabilities of the underlying model (Qwen3-VL mentioned above) in the field of physical engineering, the fine-tuning process of Qwen3-VL can include a supervised fine-tuning stage and a reinforcement learning fine-tuning stage.

[0042] The supervised fine-tuning phase utilizes high-quality domain instruction data for supervised training of the model. First, input processing is performed: the image is encoded into visual features by a visual encoder, then compressed and projected into visual labels aligned with the LLM hidden dimensions by an MLP fusion processor. This visual label sequence is concatenated with the label sequence of the instruction text and fed into a Qwen3-LLM decoder with a DeepStack mechanism. Training optimization then proceeds, with the model performing autoregressive generation based on the concatenated unified label sequence. Its output is compared with the standard answer in the data. The loss function can be cross-entropy loss, calculating only the model's prediction loss for the answer text portion. Backpropagation optimizes the model parameters, enabling the model to learn to follow instructions and output content consistent with expectations in the physical engineering domain. Training can be performed in batches and iteratively, with an early stopping strategy that can terminate training prematurely if the loss function does not significantly decrease after multiple consecutive training steps to prevent overfitting.

[0043] The goal of the reinforcement learning fine-tuning phase is to further improve the model's output quality. This phase can include: the model receiving multimodal inputs in the same format and generating multiple candidate outputs; the reward rule scoring these candidate outputs based on dimensions such as relevance, safety, and domain expertise; and adjusting the model's policy through optimization algorithms (such as Group Relative Policy Optimization (GRPO)) to generate text that yields higher rewards. This phase aims to refine the model's behavior and improve output quality.

[0044] Furthermore, training configuration is required before training. Parameters are initialized based on the pre-trained Qwen3-VL model weights. The number of training steps (ranging from thousands to tens of thousands) can be flexibly set according to computational resources and data scale, and an appropriate learning rate scheduler and optimizer can be employed. For blurry or low-resolution images, the fine-tuning process can preserve and rely on the visual encoder's native resolution processing and adaptive positional encoding capabilities to retain input information to the greatest extent possible. After training, performance validation can be performed using domain-relevant multimodal benchmarks (such as visual question answering and image caption generation tasks) to technically evaluate the fine-tuned model and verify its performance improvement.

[0045] The embodiments provided in this application, when training a multimodal large model using an interaction instruction set, retrieve domain knowledge content corresponding to each interaction instruction in the interaction instruction set from a target knowledge base. The target knowledge base is constructed based on a domain knowledge graph, which records domain knowledge content in the field of physical engineering. Each interaction instruction includes initial prompt content and expected output content. The domain knowledge content corresponding to each interaction instruction and the initial prompt content of each interaction instruction are fused to obtain the fused prompt content of each interaction instruction. The multimodal large model is trained using the fused prompt content and the expected output content of each interaction instruction to obtain a trained multimodal large model. Because the training of the multimodal large model uses domain knowledge content from an interaction instruction set based on a domain knowledge graph, and the domain knowledge graph records domain knowledge content in the field of physical engineering, the trained multimodal large model can effectively learn and understand industry standards in the field of physical engineering, and can output results applicable to construction safety in the field of physical engineering based on the requirements of the interaction instructions. This solves the problem of insufficient professionalism in construction safety scenarios in related technologies and improves the professionalism of large models.

[0046] In an exemplary embodiment, before retrieving domain knowledge content corresponding to each interaction instruction in the interaction instruction set from the target knowledge base, the method further includes: constructing a domain knowledge graph based on domain knowledge source data in the field of physical engineering, wherein the domain knowledge source data is data collected from domain knowledge sources to describe domain knowledge in the field of physical engineering, and the domain knowledge source data includes at least one of the following: standard text at at least one level; standard design diagrams; abnormal state record data; and experience description data.

[0047] To address the issue of insufficient industry-specific knowledge in the application of multimodal large models in the field of physical engineering, leading to inadequate understanding of professional standards and low accuracy in identifying safety hazards, this embodiment constructs a professional domain knowledge graph by collecting domain knowledge source data from the physical engineering field. This provides accurate and comprehensive physical engineering domain knowledge support for multimodal model training, thereby improving the model's industry adaptability and professionalism.

[0048] In this embodiment, a domain knowledge graph is constructed based on domain knowledge source data in the field of physical engineering. The domain knowledge source data refers to data collected from domain knowledge sources that describes domain knowledge in the field of physical engineering. This domain knowledge source data includes at least one of the following: standard text at at least one level; standard design diagrams; abnormal state record data; and experience description data. Optionally, web crawler systems such as Scrapy and Selenium can be used to collect domain knowledge data. For example, national standards, industry standards, and local standards in the field of physical engineering (i.e., standard text at least one level) can be collected from official platforms, or specification atlases (standard design diagrams), typical accident case reports (abnormal state record data), and expert experience documents (experience description data) can be obtained from partner organizations.

[0049] The collected domain knowledge source data needs to be parsed and structured. For example, for scanned PDF files, PaddleOCR + LayoutParser technology can be used to extract the text and image layout information, and the Bidirectional Encoder Representations from Transformers-Conditional Random Fields (BERT-CRF) model can be used to identify key entities in the security clauses. After processing the domain knowledge source data, a professional domain knowledge graph can be constructed. Optionally, Neo4j can be used to store the domain knowledge graph to achieve semantic association and cross-standard references.

[0050] Optionally, after the domain knowledge graph is constructed, it can be dynamically updated. New domain knowledge source data can be collected at specified intervals, and a semantic difference detection module based on "Sentence-BERT" can be designed to encode the new and old domain knowledge source data into vectors, calculate cosine similarity, identify added, modified or deleted terms, and trigger incremental updates to the corresponding nodes or relationships in the domain knowledge graph. At the same time, change logs are recorded to ensure that the domain knowledge graph is synchronized with the latest knowledge.

[0051] This embodiment constructs a comprehensive and structured domain knowledge graph of the physical engineering field to guide the training of a multimodal large model. This significantly improves the multimodal large model's understanding of physical engineering specifications, as well as its accuracy and professionalism in identifying and describing safety hazards. As a result, the output results are more in line with industry standards, reducing false alarms and missed alarms in safety hazard identification and enhancing the feasibility of professional rectification suggestions.

[0052] In an exemplary embodiment, a domain knowledge graph is constructed based on domain knowledge source data in the field of physical engineering, including: extracting a set of domain knowledge triples from the domain knowledge source data, wherein each domain knowledge triple in the set of domain knowledge triples is a triple with an abnormal state type, a state baseline entry, and a state correction scheme as entities; and constructing a knowledge graph with each entity in each domain knowledge triple as a node to obtain the domain knowledge graph.

[0053] To address the lack of in-depth understanding and accurate correction suggestions for anomalous states in specific domains within multimodal large models in the field of physical engineering, this embodiment proposes a method to extract domain knowledge triples from domain knowledge source data and construct a domain knowledge graph. This method can improve the ability of multimodal large models to identify anomalous states in physical engineering and provide effective state correction solutions.

[0054] In this embodiment, a set of domain knowledge triples is extracted from the domain knowledge source data. Each domain knowledge triple in the set is a triple consisting of an anomaly state type, a state baseline entry, and a state correction scheme as entities. Optionally, natural language processing and machine learning techniques, such as the BERT-CRF entity recognition model and dependency parsing, are used to extract a set of domain knowledge triples related to anomalies. Each triple consists of three entities: an anomaly state type, a state baseline entry, and a state correction scheme. For example, "Anomaly State Type: Incomplete Edge Protection", "State Baseline Entry: Item X of X", and "State Correction Scheme: Install a 1-2 meter high guardrail".

[0055] A domain knowledge graph is constructed by using each entity in each domain knowledge triple as a node. Here, the domain knowledge graph can be structured into a chain-like structure with nodes representing abnormal state types, state baseline entries, and state correction schemes, and edges representing the relationships between nodes. This facilitates the retrieval and invocation of relevant domain knowledge content by multimodal models during training and inference.

[0056] In this embodiment, a structured and queryable knowledge graph of the physical engineering domain is constructed. This graph uses the triple of "abnormal state type - state baseline entry - state correction scheme" as its core knowledge structure, which provides professional and accurate domain knowledge support for the training of multimodal large models and improves the professionalism of multimodal large models.

[0057] In an exemplary embodiment, before fusing the domain knowledge content corresponding to each interaction instruction and the initial prompt content of each interaction instruction, the method further includes: interacting with the interactive language model using each job scene image in the job scene image set and a step-by-step prompt template to obtain a multi-turn interaction instruction set; generating an interaction instruction set based on the multi-turn interaction instruction set, wherein each interaction instruction is at least a portion of an interaction instruction in a multi-turn interaction instruction set; wherein each job scene image is a real job scene image of an engineering operation in a physical engineering field, or a synthetic job scene image of an engineering operation in a physical engineering field; wherein the step-by-step prompt template is used to interact with the interactive language model step-by-step according to the target prompt chain, and the prompt nodes in the target prompt chain include at least one of the following: job scene description information, abnormal state location indication information, abnormal state type, abnormal level, state baseline entry, and state correction scheme.

[0058] To address the issue of the lack of detailed understanding and description capabilities for specific operational scenarios when applying multimodal large models in the field of physical engineering, this embodiment uses real or synthetic operational scenario images combined with step-by-step prompt templates for multi-round interactions to generate an interactive instruction set containing detailed scenario descriptions and specific questions. This can improve the model's understanding and responsiveness in complex physical engineering scenarios.

[0059] In this embodiment, each task scene image in the task scene image set and a step-by-step prompt template are used to interact with the interactive language model to obtain a multi-turn interactive instruction set. Here, the step-by-step prompt template is used to interact with the interactive language model step by step according to the target prompt chain. Optionally, a step-by-step prompt template based on Chain-of-Thought (CoT) can be designed to force the multimodal large model to reason sequentially. "Few-shot" examples can also be introduced to guide the model to generate standard format output. The prompt nodes in the target prompt chain include at least one of the following: job scenario description information, abnormal state location indication information, abnormal state type, abnormality level, state baseline entry, and state correction scheme. For example, the process of a multimodal large model interacting with an interactive language model step by step according to the target prompt chain can be as follows: After inputting a job scenario image and a step-by-step prompt template into the multimodal large model, the multimodal large model is first instructed to output job scenario description information, then instructed to locate abnormal states and generate abnormal state location indication information, determine the abnormal state type and abnormality level, provide the state baseline entry corresponding to the abnormal state, and finally output the state correction scheme corresponding to the abnormal state. During the interaction, the multimodal large model can learn step by step how to describe image content, locate abnormal states, identify abnormal types and their severity, reference relevant standard entries, and design reasonable correction schemes based on the prompts. The results of these interactions can be summarized into a multi-round interaction instruction set, containing detailed analysis and suggestions from the multimodal large model for different job scenario images.

[0060] Each work scene image is either a real work scene image of a physical engineering project in the field of physical engineering, or a composite work scene image of a physical engineering project in the field of physical engineering. Here, the collection of work scene images can cover a variety of engineering scenarios such as water supply, drainage, roads, bridges, tunnels, gas supply, and power supply, covering different construction stages, weather conditions, and shooting angles, ensuring data diversity, and ensuring that it can cover most scenarios in the field of physical engineering.

[0061] Based on a multi-turn interaction instruction set, an interaction instruction set is generated, wherein each interaction instruction is at least a portion of a multi-turn interaction instruction set within the multi-turn interaction instruction set. Here, a single-turn interaction instruction may include a questioning process that includes the aforementioned job scenario description information, abnormal state location indication information, abnormal state type, abnormal level, state baseline entries, and state correction scheme. A multi-turn interaction instruction can be obtained by continuing to input prompts to ask questions to the multimodal large model based on the output of a single-turn interaction instruction.

[0062] This embodiment enhances the scene understanding and abnormal state recognition capabilities of the multimodal large model by generating an interactive instruction set through multiple rounds of interaction with real or synthetic work scene images, thereby improving the scene adaptability of the multimodal large model and the accuracy and professionalism of abnormal state recognition.

[0063] In an exemplary embodiment, before generating the interaction instruction set based on the multi-turn interaction instruction set, the method further includes: removing multi-turn interaction instructions from the multi-turn interaction instruction set that meet at least one of the following screening conditions to obtain an updated multi-turn interaction instruction set: the included interaction content does not match the prompt node in the target prompt chain; the semantic similarity between the included state correction scheme and the corresponding state baseline entry is lower than the similarity threshold; or it was sampled during the sampling review process and the sampling review result was that the review failed.

[0064] To improve the quality of the interaction instruction set, multiple rounds of interaction instructions can be screened before generating the interaction instruction set to ensure that each interaction instruction is accurate, professional, and conforms to industry standards, thereby ensuring that the multimodal large model can learn high-quality information.

[0065] In this embodiment, a combination of machine screening and manual sampling can be used to remove multi-turn interaction instructions from the multi-turn interaction instruction set that meet at least one of the following removal conditions, resulting in an updated multi-turn interaction instruction set. The removal conditions for machine screening include: the included interaction content does not match the prompt nodes in the target prompt chain; the semantic similarity between the included state correction scheme and the corresponding state baseline entry is lower than a similarity threshold. The removal conditions for manual sampling include: being sampled during the sampling review process, and the sampling review result being "review failed".

[0066] Optionally, detecting whether the included interactive content does not match the prompt nodes in the target prompt chain can be done by checking whether it contains standard numbers (regular expressions), bounding box formats, risk level enumerations, etc.; detecting whether the semantic similarity between the included state correction scheme and the corresponding state baseline item is lower than the similarity threshold can be done by using "SimCSE" to calculate the semantic similarity between the generated state correction scheme and the corresponding state baseline item, and then comparing it with the similarity threshold; manual sampling can extract a certain amount (e.g., 5%) of random samples from the multi-round interactive instructions after machine screening, and remove the multi-round interactive instructions that failed the review of the sampling review.

[0067] This embodiment, through a data quality control process for multi-round interactive instructions, can ensure the professionalism and compliance of the dataset, and improve the training efficiency and training effect of the model.

[0068] In an exemplary embodiment, the process of obtaining domain knowledge content corresponding to each interactive instruction in the interactive instruction set from the target knowledge base includes: selecting M domain knowledge contents from the target knowledge base for each interactive instruction in descending order of vector similarity between the encoded vector of each domain knowledge content in the target knowledge base and the encoded vector of the initial prompt content of each interactive instruction, thereby obtaining domain knowledge content corresponding to each interactive instruction; wherein, the encoded vector of each domain knowledge content is a vector obtained by encoding each domain knowledge content, and the encoded vector of the initial prompt content of each interactive instruction is a vector obtained by encoding the initial prompt content of each interactive instruction, and M is a positive integer greater than or equal to 2; wherein, among the M domain knowledge contents selected for each interactive instruction, the first to Nth domain knowledge contents are domain knowledge contents used in the first model training stage of the multimodal large model, and the (N+1)th domain knowledge content is domain knowledge contents used in the second model training stage of the multimodal large model, and N is a positive integer greater than or equal to 1 and less than M, and the first model training stage is earlier than the second model training stage.

[0069] In order to efficiently and accurately retrieve the most relevant knowledge content for each interaction command from the target knowledge base built on a large domain knowledge graph, and to use this knowledge content to fine-tune the model in order to gradually improve the model's industry adaptability and professionalism, in this embodiment, a vector similarity-based filtering method can be used to retrieve the domain knowledge content corresponding to the interaction command, and different training methods can be used in stages to improve the model's adaptability in various situations.

[0070] In this embodiment, M domain knowledge contents are selected from the target knowledge base for each interactive instruction based on the descending order of vector similarity between the encoded vector of each domain knowledge content and the encoded vector of the initial prompt content of each interactive instruction, thus obtaining the domain knowledge content corresponding to each interactive instruction. Here, the encoded vector of each domain knowledge content is a vector obtained by encoding each domain knowledge content, and the encoded vector of the initial prompt content of each interactive instruction is a vector obtained by encoding the initial prompt content of each interactive instruction. Optionally, a pre-trained semantic encoding model, such as Simple Contrastive Learning of Sentence Embeddings (SimCSE) or Bidirectional Encoder Representation from Transformers (BERT), can be used to transform each node in the target knowledge base into a fixed-dimensional encoded vector, constructing an efficient knowledge index based on a vector database; the initial prompt content of each interactive instruction in the interactive instruction set can also be transformed into an encoded vector using the same semantic encoding model.

[0071] After obtaining the encoding vectors of each domain knowledge content and the initial prompt content of each interactive instruction, the vector similarity (e.g., cosine similarity) between the encoding vectors of each domain knowledge content in the target knowledge base and the encoding vectors of the initial prompt content of each interactive instruction can be calculated. Based on the similarity level, M domain knowledge content items are selected from the target knowledge base for each interactive instruction, where M is a positive integer greater than or equal to 2. These knowledge content items can be used for fine-tuning training of multimodal large models.

[0072] In the fine-tuning training of multimodal large models, a phased training approach can be adopted. In the initial training phase (i.e., the first model training phase), the first to Nth domain knowledge contents selected from M domain knowledge contents for each interaction command are used for training, and high-confidence retrieval results are used for training. In the later training phase (i.e., the second model training phase), the N+1th domain knowledge contents selected from M domain knowledge contents for each interaction command are used for training, and lower-confidence retrieval results are introduced for training, thereby improving the tolerance of multimodal large models to imperfect retrieval.

[0073] This embodiment, through a vector similarity-based retrieval mechanism, ensures that the model can acquire domain knowledge content most relevant to the current interaction command during training, thereby improving the relevance and efficiency of model learning. The phased model training method enables the model to learn progressively from basic to advanced levels, gradually enhancing its professionalism in the field of physical engineering.

[0074] In one exemplary embodiment, the domain knowledge content corresponding to each interaction instruction and the initial prompt content of each interaction instruction are fused to obtain the fused prompt content of each interaction instruction. This includes adding target prompt content and domain knowledge content corresponding to each interaction instruction to the initial prompt content of each interaction instruction to obtain the fused prompt content of each interaction instruction. The target prompt content is used to prompt the multimodal large model to generate output content corresponding to the initial prompt content of each interaction instruction based on the domain knowledge content corresponding to each interaction instruction.

[0075] To address the potential issues of insufficient generalization and lack of expertise in multimodal large models when dealing with complex physical engineering problems, and to ensure that the model can make decisions and recommendations based on industry standards and professional knowledge, this embodiment integrates domain knowledge content with the initial prompts of interactive commands to generate more accurate and professional output content.

[0076] In this embodiment, target prompt content and domain knowledge content corresponding to each interaction instruction are added to the initial prompt content of each interaction instruction to obtain fused prompt content for each interaction instruction. The target prompt content is used to prompt the multimodal large model to generate output content corresponding to the initial prompt content of each interaction instruction based on the domain knowledge content corresponding to each interaction instruction. Here, the role of the target prompt content is to guide the model to focus on and apply domain knowledge content closely related to each interaction instruction when processing it, ensuring that the model's output content is both professional and meets industry requirements. For example, the fused prompt content of an interaction instruction including the target prompt content could be: "[System] You are a senior construction safety expert. Please answer the questions according to the following standard clauses: {Retrieved state baseline entries}; {Retrieved similar cases, state correction schemes, etc.}". During the model inference stage, the user's interaction with the model can be: "[User] {Original image description} {Specific question}; [Answer] {Expected output}". Here, the expected output is the output result that satisfies the template form of the above fused prompt content in response to the user's question.

[0077] This embodiment, by introducing target prompts when integrating domain knowledge content corresponding to interaction commands and interaction commands, ensures that the multimodal large model answers in a standardized format, thus enhancing the professionalism of the multimodal large model.

[0078] The model training method for multimodal large models in the embodiments of this application will be explained below with reference to optional examples. Figure 4 This is a flowchart illustrating the model training method for a multimodal large model in this optional example, such as... Figure 4 As shown, the process of training this multimodal large model may include the following steps:

[0079] The system collects domain knowledge source data from multiple sources and constructs a domain knowledge graph based on the collected data. This graph is then input into the retrieval and fusion module. Optionally, the retrieval and fusion module may include a Retrieval-Augmented Generation (RAG) system, which enables the multimodal large model to retrieve knowledge base (i.e., domain knowledge graph) data and has a dynamic update function. This allows the model to improve the quality of responses by updating the knowledge base without retraining the model.

[0080] Interactive instruction sets are generated through a data engine. Optionally, the data engine may include visual semantic models such as Qwen3-VL, GLM4.5 Multilingual Vision Understanding Model (GLM4.5 for short), and ByteDance SAIL Vision-Language Model 2 (SAIL-VL 2 for short), which can be used to generate interactive instruction sets from images of the work scene.

[0081] Training the base model using reinforcement learning modules, retrieval and fusion modules, and an interaction instruction set yields a trained multimodal large model. Optionally, the reinforcement learning module can include reinforcement learning (RL) combined with learning from other agents' rewards (LoAR), enabling the multimodal large model to learn the optimal policy through trial and error and adjust its behavior by observing reward signals from other agents. The base model, which serves as the underlying model of the multimodal large model, can be Qwen3-VL, possessing strong generalization capabilities and basic modal understanding. Through training with a domain knowledge graph, reinforcement learning modules, and an interaction instruction set, it can be gradually transformed into a specialized multimodal large model for the physics and engineering domain.

[0082] This optional example demonstrates how to collect domain knowledge source data, construct a structured domain knowledge base based on a knowledge graph, generate an interactive instruction set through a data engine, and train a basic model to obtain a professional multimodal large model in the field of physics and engineering. It covers the entire process of professional knowledge collection, multimodal instruction dataset generation, and online fine-tuning training, which can effectively inject domain knowledge into the multimodal large model and improve its professionalism.

[0083] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0084] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / random access memory (RAM), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0085] According to another aspect of the embodiments of this application, a model training apparatus for a multimodal large model is also provided. This apparatus can be used to implement the model training method for a multimodal large model provided in the above embodiments, and will not be repeated hereafter. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0086] Figure 5 This is a structural block diagram of an optional multimodal large model training device according to an embodiment of this application, such as... Figure 5 As shown, the model training device for this multimodal large model includes:

[0087] The retrieval unit 502 is used to retrieve domain knowledge content corresponding to each interactive instruction in the interactive instruction set from the target knowledge base when training a multimodal large model using an interactive instruction set. The target knowledge base is constructed based on a domain knowledge graph, which is used to record domain knowledge content in the field of physical engineering. Each interactive instruction includes initial prompt content and expected output content.

[0088] The fusion unit 504 is used to fuse the domain knowledge content corresponding to each interaction command and the initial prompt content of each interaction command to obtain the fused prompt content of each interaction command.

[0089] Training unit 506 is used to train a multimodal large model using the fused prompt content of each interactive instruction and the expected output content of each interactive instruction, so as to obtain a trained multimodal large model.

[0090] It should be noted that the retrieval unit 502 in this embodiment can be used to perform the above step S202, the fusion unit 504 in this embodiment can be used to perform the above step S204, and the training unit 506 in this embodiment can be used to perform the above step S206.

[0091] The embodiments provided in this application, when training a multimodal large model using an interaction instruction set, retrieve domain knowledge content corresponding to each interaction instruction in the interaction instruction set from a target knowledge base. The target knowledge base is constructed based on a domain knowledge graph, which records domain knowledge content in the field of physical engineering. Each interaction instruction includes initial prompt content and expected output content. The domain knowledge content corresponding to each interaction instruction and the initial prompt content of each interaction instruction are fused to obtain the fused prompt content of each interaction instruction. The multimodal large model is trained using the fused prompt content and the expected output content of each interaction instruction to obtain a trained multimodal large model. Because the training of the multimodal large model uses domain knowledge content from an interaction instruction set based on a domain knowledge graph, and the domain knowledge graph records domain knowledge content in the field of physical engineering, the trained multimodal large model can effectively learn and understand industry standards in the field of physical engineering, and can output results applicable to construction safety in the field of physical engineering based on the requirements of the interaction instructions. This solves the problem of insufficient professionalism in construction safety scenarios in related technologies and improves the professionalism of large models.

[0092] In one exemplary embodiment, the apparatus further includes: a construction unit, configured to construct a domain knowledge graph based on domain knowledge source data in the field of physical engineering before retrieving domain knowledge content corresponding to each interaction instruction in the interaction instruction set from the target knowledge base, wherein the domain knowledge source data is data collected from domain knowledge sources to describe domain knowledge in the field of physical engineering, and the domain knowledge source data includes at least one of the following: at least one level of standard text; standard design diagrams; abnormal state record data; and experience description data.

[0093] In an exemplary embodiment, the construction unit includes: an extraction module, configured to extract a set of domain knowledge triples from the domain knowledge source data, wherein each domain knowledge triple in the set of domain knowledge triples is a triple with an abnormal state type, a state baseline entry, and a state correction scheme as entities; and a construction module, configured to construct a knowledge graph with each entity in each domain knowledge triple as a node, thereby obtaining a domain knowledge graph.

[0094] In an exemplary embodiment, the apparatus further includes: an interaction unit, configured to interact with an interactive language model using each work scenario image in the work scenario image set and a step-by-step prompt template before fusing the domain knowledge content corresponding to each interaction instruction and the initial prompt content of each interaction instruction, to obtain a multi-round interaction instruction set; and a generation unit, configured to generate an interaction instruction set based on the multi-round interaction instruction set, wherein each interaction instruction is at least a portion of an interaction instruction in a multi-round interaction instruction set; wherein each work scenario image is a real work scenario image of a physical engineering operation in the field of physical engineering, or a synthetic work scenario image of a physical engineering operation in the field of physical engineering; wherein the step-by-step prompt template is used to interact with the interactive language model step-by-step according to a target prompt chain, and the prompt nodes in the target prompt chain include at least one of the following: work scenario description information, abnormal state location indication information, abnormal state type, abnormal level, state baseline entry, and state correction scheme.

[0095] In an exemplary embodiment, the apparatus further includes: a removal unit, configured to remove multi-turn interaction instructions from the multi-turn interaction instruction set that meet at least one of the following screening conditions before generating an interaction instruction set based on the multi-turn interaction instruction set, thereby obtaining an updated multi-turn interaction instruction set: the included interaction content does not match the prompt node in the target prompt chain; the semantic similarity between the included state correction scheme and the corresponding state baseline entry is lower than a similarity threshold; or it was sampled during the sampling review process and the sampling review result is that the review failed.

[0096] In an exemplary embodiment, the retrieval unit includes: a filtering module, configured to filter M domain knowledge contents from the target knowledge base for each interactive instruction according to the order of vector similarity between the encoded vector of each domain knowledge content in the target knowledge base and the encoded vector of the initial prompt content of each interactive instruction, from high to low, to obtain the domain knowledge content corresponding to each interactive instruction; wherein, the encoded vector of each domain knowledge content is a vector obtained by encoding each domain knowledge content, and the encoded vector of the initial prompt content of each interactive instruction is a vector obtained by encoding the initial prompt content of each interactive instruction, and M is a positive integer greater than or equal to 2; wherein, among the M domain knowledge contents filtered for each interactive instruction, the first to Nth domain knowledge contents are domain knowledge contents used in the first model training stage of the multimodal large model, and the (N+1)th domain knowledge content is domain knowledge contents used in the second model training stage of the multimodal large model, and N is a positive integer greater than or equal to 1 and less than M, and the first model training stage is earlier than the second model training stage.

[0097] In an exemplary embodiment, the fusion unit includes: an adding module, configured to add target prompt content and domain knowledge content corresponding to each interaction instruction to the initial prompt content of each interaction instruction, thereby obtaining fused prompt content for each interaction instruction, wherein the target prompt content is used to prompt the multimodal big model to generate output content corresponding to the initial prompt content of each interaction instruction based on the domain knowledge content corresponding to each interaction instruction.

[0098] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.

[0099] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, the computer-readable storage medium including a stored program, wherein the program executes the steps in any of the above method embodiments when it is run.

[0100] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, ROMs, RAMs, portable hard drives, magnetic disks, or optical disks.

[0101] According to another aspect of the embodiments of this application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor is configured to perform the steps of any of the method embodiments described above via the computer program. In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0102] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0103] According to another aspect of the embodiments of this application, a computer program product is also provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via communication section 609, and / or installed from removable medium 611. When the computer program is executed by central processing unit 601, it performs various functions provided in the embodiments of this application. The sequence numbers of the embodiments of this application above are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0104] Figure 6 A schematic block diagram of a computer system architecture for implementing embodiments of the present application is shown. Figure 6 As shown, the computer system 600 includes a Central Processing Unit (CPU) 601, which performs various appropriate actions and processes based on programs stored in ROM 602 or loaded into RAM 603 from storage section 608. Random access memory 603 also stores various programs and data required for system operation. The CPU 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0105] The following components are connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), and speakers, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card, such as a local area network card or modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 610 as needed so that computer programs read from it can be installed into storage section 608 as needed.

[0106] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 609, and / or installed from removable medium 611. When the computer program is executed by central processing unit 601, it performs various functions defined in the system of this application.

[0107] It should be noted that, Figure 6 The computer system 600 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0108] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0109] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. A method for training a multimodal large model, characterized in that, include: When training a multimodal large model using an interactive instruction set, domain knowledge content corresponding to each interactive instruction in the interactive instruction set is retrieved from the target knowledge base. The target knowledge base is constructed based on a domain knowledge graph, which is used to record domain knowledge content in the field of physical engineering. Each interactive instruction includes initial prompt content and expected output content. The domain knowledge content corresponding to each interactive instruction and the initial prompt content of each interactive instruction are merged to obtain the merged prompt content of each interactive instruction; The multimodal large model is trained using the fused prompt content of each interaction instruction and the expected output content of each interaction instruction to obtain the trained multimodal large model.

2. The method according to claim 1, characterized in that, Before retrieving domain knowledge content corresponding to each interaction instruction in the interaction instruction set from the target knowledge base, the method further includes: Based on the domain knowledge source data of the physical engineering field, the domain knowledge graph is constructed. The domain knowledge source data is data collected from domain knowledge sources to describe the domain knowledge of the physical engineering field. The domain knowledge source data includes at least one of the following: standard text at at least one level; standard design diagrams; abnormal state record data; and experience description data.

3. The method according to claim 2, characterized in that, The construction of the domain knowledge graph based on the domain knowledge source data of the physical engineering field includes: Extract a set of domain knowledge triples from the domain knowledge source data, wherein each domain knowledge triple in the set of domain knowledge triples is a triple with an abnormal state type, a state baseline entry, and a state correction scheme as entities; A knowledge graph is constructed using each entity in each domain knowledge triple as a node, resulting in the domain knowledge graph.

4. The method according to claim 1, characterized in that, Before fusing the domain knowledge content corresponding to each interaction instruction and the initial prompt content of each interaction instruction, the method further includes: Each task scenario image and step-by-step prompt template in the task scenario image set are used to interact with the interactive language model to obtain a multi-round interactive instruction set; Based on the multi-turn interaction instruction set, the interaction instruction set is generated, wherein each interaction instruction is at least a portion of an interaction instruction in a multi-turn interaction instruction set; Each of the work scene images is either a real work scene image of a physical engineering operation in the field of physical engineering, or a composite work scene image of a physical engineering operation in the field of physical engineering. The step-by-step prompt template is used to interact with the interactive language model step by step according to the target prompt chain. The prompt nodes in the target prompt chain include at least one of the following: job scenario description information, abnormal state location indication information, abnormal state type, abnormal level, state baseline entry, and state correction scheme.

5. The method according to claim 4, characterized in that, Before generating the interaction instruction set based on the multi-round interaction instruction set, the method further includes: The multi-turn interaction instruction set is updated by removing multi-turn interaction instructions that satisfy at least one of the following filtering conditions from the multi-turn interaction instruction set: The included interactive content does not match the prompt nodes in the target prompt chain; The semantic similarity between the included state correction scheme and the corresponding state baseline item is lower than the similarity threshold; It was sampled during the sampling audit process, and the sampling audit result was that the audit failed.

6. The method according to claim 1, characterized in that, The step of retrieving domain knowledge content from the target knowledge base corresponding to each interaction instruction in the interaction instruction set includes: Based on the order of high to low vector similarity between the encoding vector of each domain knowledge content in the target knowledge base and the encoding vector of the initial prompt content of each interactive instruction, M domain knowledge contents are selected from the target knowledge base for each interactive instruction to obtain the domain knowledge content corresponding to each interactive instruction. Wherein, the encoding vector of each domain knowledge content is a vector obtained by encoding each domain knowledge content, the encoding vector of the initial prompt content of each interactive instruction is a vector obtained by encoding the initial prompt content of each interactive instruction, and M is a positive integer greater than or equal to 2; Among the M domain knowledge contents selected for each interactive instruction, the first to Nth domain knowledge contents are domain knowledge contents used in the first model training stage of the multimodal large model, and the (N+1)th domain knowledge contents are domain knowledge contents used in the second model training stage of the multimodal large model. N is a positive integer greater than or equal to 1 and less than M. The first model training stage is earlier than the second model training stage.

7. The method according to any one of claims 1 to 6, characterized in that, The process of fusing the domain knowledge content corresponding to each interaction command and the initial prompt content of each interaction command to obtain the fused prompt content of each interaction command includes: Add target prompt content and domain knowledge content corresponding to each interaction instruction to the initial prompt content of each interaction instruction to obtain fused prompt content with each interaction instruction. The target prompt content is used to prompt the multimodal big model to generate output content corresponding to the initial prompt content of each interaction instruction based on the domain knowledge content corresponding to each interaction instruction.

8. A model training device for a multimodal large model, characterized in that, include: The retrieval unit is used to retrieve domain knowledge content corresponding to each interactive instruction in the interactive instruction set from the target knowledge base when training a multimodal large model using an interactive instruction set. The target knowledge base is constructed based on a domain knowledge graph, which is used to record domain knowledge content in the field of physical engineering. Each interactive instruction includes initial prompt content and expected output content. The fusion unit is used to fuse the domain knowledge content corresponding to each interaction instruction and the initial prompt content of each interaction instruction to obtain the fused prompt content of each interaction instruction. The training unit is used to train the multimodal large model using the fused prompt content of each interaction instruction and the expected output content of each interaction instruction, so as to obtain the trained multimodal large model.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.