Multimodal emergency information generation methods, devices, electronic equipment, and media
By combining multimodal data filtering, knowledge distillation of thought chains, and large language models, a multimodal emergency vertical domain model is constructed, which solves the problems of applicability and relevance of emergency information generation methods and realizes efficient and accurate information generation in complex emergency scenarios.
Patent Information
- Application Number
- CN202510093796.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-01-21
AI Technical Summary
Existing emergency information generation methods are poorly applicable and targeted in multimodal applications, making it difficult to meet real-time and personalized needs and effectively cope with complex emergency scenarios.
By acquiring emergency-related graphic and textual data, multimodal data filtering and expansion are performed to construct graphic and textual instruction fine-tuning data based on knowledge distillation of thought chains. The initial multimodal emergency model is then fine-tuned a second time to form a multimodal emergency vertical domain model. Combined with large language models and retrieval enhancement technology, emergency-related response information is generated.
It significantly improves the accuracy, practicality, and diversity of emergency information generation, enabling the rapid generation of accurate and real-time effective emergency-related responses, and is suitable for complex emergency scenarios.
Smart Images

Figure CN119938858B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device, and medium for generating emergency information based on multimodality. Background Technology
[0002] In recent years, with the increasing demands on emergency response capabilities, efficiently generating science-based information for handling emergencies has become a key application area for artificial intelligence technology. In emergency scenarios such as fires, construction accidents, and medical emergencies, timely provision of scientific and accurate emergency knowledge and personalized response plans to the public can not only effectively reduce disaster losses and improve the utilization rate of public resources, but also enhance public understanding and participation in the entire emergency response process.
[0003] With the rapid development of artificial intelligence, multimodal application scenarios have gradually become a research hotspot in the field of emergency information generation. Multimodal technology, by integrating various data formats such as images and text, can more comprehensively perceive and understand complex scenarios. In emergency scenarios such as fire prevention and medical emergency response, images are often key information carriers. By identifying the specific environment and potential risk points in images, combined with professional guidance on text generation, users can be provided with more accurate and intuitive solutions. In practice, multimodal applications not only enhance the richness and interest of information expression but also provide strong support for personalized and dynamic generation of emergency information.
[0004] However, existing emergency information generation methods still have significant shortcomings in multimodal applications. For example, the commonly used fixed template method can only handle preset scenarios and has poor applicability. Another example is the manual writing method, which, while offering some flexibility, lacks more efficient analytical capabilities and is less targeted. Furthermore, traditional methods remain limited in terms of information update frequency and content accuracy, making it difficult to meet the demands for real-time and personalized information. Summary of the Invention
[0005] This invention provides a multimodal emergency information generation method, device, electronic device, and medium to solve or partially solve the problems of poor applicability and targeting of current emergency information generation methods, limitations in information update frequency and content accuracy, and difficulty in meeting real-time and personalized needs.
[0006] This invention provides a multimodal emergency information generation method, the method comprising:
[0007] Acquire emergency-related graphic and textual data and general task datasets, and perform multimodal data filtering and expansion on the emergency-related graphic and textual data to obtain graphic and textual question-and-answer pair data;
[0008] Based on the aforementioned text-image question-and-answer pair data, construct text-image instruction fine-tuning data based on knowledge distillation of the thought chain;
[0009] The pre-built initial multimodal emergency model is fine-tuned a second time using the graphic and textual instruction fine-tuning data and the general task dataset to obtain a multimodal emergency vertical domain model.
[0010] The multimodal emergency vertical domain model is used to perform image recognition analysis based on retrieval enhancement on emergency-related query information input by users, and to generate emergency-related response information.
[0011] Optionally, the emergency-related graphic data includes multiple valid graphic-text pairs, each of which includes a valid image description text and a valid emergency image; the multimodal data filtering and expansion of the emergency-related graphic data to obtain graphic-text question-and-answer pair data includes:
[0012] By integrating various refined and multimodal emergency tasks, instruction texts are constructed based on task-specific instruction requirements, task category listings, and specific examples.
[0013] Based on the instruction text, all the effective emergency images are divided into tasks using a pre-trained multimodal task partitioning model. Based on the task partitioning results, emergency-irrelevant images are removed, and all the remaining effective emergency images are labeled to obtain the corresponding target emergency images.
[0014] The image description text corresponding to the target emergency image is used as the target image description text, and the target emergency image and the target image description text are used as an image-text question-and-answer pair;
[0015] All the aforementioned text-image question-and-answer pairs will be integrated into text-image question-and-answer pair data.
[0016] Optionally, the step of constructing image-text instruction fine-tuning data based on the image-text question-and-answer pair data and knowledge distillation of the thought chain includes:
[0017] An initial image-text description model is constructed based on a multimodal large language model framework;
[0018] Using the aforementioned text-image question-and-answer pair data, the initial text-image description model is trained by text-image expansion generation based on knowledge distillation of thought chains, and text-image instruction fine-tuning data is generated based on the trained text-image description model.
[0019] Optionally, the step of using the text-image question-and-answer pair data to train the initial text-image description model through text-image expansion generation based on knowledge distillation of thought chains, and generating text-image instruction fine-tuning data based on the trained text-image description model, includes:
[0020] For each of the image-text question-and-answer pairs, based on the task division results corresponding to the target emergency image, the target emergency image is described using the initial image-text description model, and the image description model output text is generated.
[0021] The output text of the image description model is compared with the target image description text. Based on the differences in the content comparison, the output text of the image description model is corrected to obtain the corrected image description text.
[0022] When the image descriptions and text corrections for all the image-text question-and-answer pairs are completed, the trained image-text description model is obtained.
[0023] Based on the trained image-text description model, emergency vertical image-text pairs are generated for each of the image description correction texts, and all the emergency vertical image-text pairs are integrated into image-text instruction fine-tuning data.
[0024] Optionally, the step of performing secondary instruction fine-tuning on the pre-constructed initial multimodal emergency model based on the image and text instruction fine-tuning data and the general task dataset to obtain a multimodal emergency vertical domain model includes:
[0025] An initial multimodal emergency response model is constructed based on a multimodal large language model framework, and a LoRA model fine-tuning module is introduced into the initial multimodal emergency response model;
[0026] Set the fine-tuning training parameters of the LoRA model fine-tuning module;
[0027] The fine-tuned data of the graphic and text instructions is mixed with the general task dataset to obtain an emergency graphic and text hybrid dataset;
[0028] The initial multimodal emergency model is trained using the aforementioned emergency image and text hybrid dataset. During the training process, the initial multimodal emergency model is fine-tuned using secondary instructions based on the fine-tuning training parameters to obtain a multimodal emergency vertical domain model.
[0029] Optionally, the step of performing retrieval-enhanced image recognition analysis on the emergency-related query information input by the user through the multimodal emergency vertical domain model, and generating emergency-related response information, includes:
[0030] Obtain emergency-related query information input by the user, which includes the image to be analyzed and the text information to be answered;
[0031] The image to be analyzed is input into a pre-trained multimodal image annotation model for image annotation to obtain an overview of the image content corresponding to the image to be analyzed.
[0032] Using the image content overview and the text information to be answered as key information for retrieval, emergency-related data is retrieved from the vector database.
[0033] The emergency-related data, the image to be analyzed, and the text information to be answered are input into the multimodal emergency vertical domain model;
[0034] The multimodal emergency vertical domain model is used to perform image recognition analysis on the image to be analyzed, focusing on the text information to be answered. During the image recognition analysis, the emergency-related data is used to enhance the context of the text information to be answered, generating emergency-related response information.
[0035] Optionally, the process of collecting the emergency-related graphic and textual data includes:
[0036] Build a keyword thesaurus related to emergency scenarios;
[0037] Obtain an open-source image and text dataset from the internet, which contains multiple open-source images, each of which corresponds to a descriptive text for an image.
[0038] Based on the keyword lexicon, keyword matching and filtering are performed on each image description text;
[0039] When the number of occurrences of the keywords in the image description text exceeds a preset threshold, the image description text is determined to be valid image description text, and the open-source image corresponding to the valid image description text is a valid emergency image.
[0040] The effective image description text and the effective emergency image are used as a valid image-text pair;
[0041] All the aforementioned valid image and text pairs will be integrated into emergency-related image and text data.
[0042] The present invention also provides a multimodal emergency information generation device, comprising:
[0043] The multimodal data filtering and expansion unit is used to acquire emergency-related graphic and textual data and general task datasets, and to perform multimodal data filtering and expansion on the emergency-related graphic and textual data to obtain graphic and textual question-and-answer pair data;
[0044] The image and text instruction fine-tuning data construction unit is used to construct image and text instruction fine-tuning data based on the knowledge distillation of the thought chain according to the image and text question and answer pair data;
[0045] The secondary instruction fine-tuning unit is used to perform secondary instruction fine-tuning on the pre-constructed initial multimodal emergency model based on the graphic instruction fine-tuning data and the general task dataset to obtain a multimodal emergency vertical domain model.
[0046] The image recognition and analysis unit is used to perform retrieval-enhanced image recognition analysis on the emergency-related query information input by the user through the multimodal emergency vertical domain model, and generate emergency-related response information.
[0047] The present invention also provides an electronic device, the device comprising a processor and a memory:
[0048] The memory is used to store program code and transmit the program code to the processor;
[0049] The processor is used to execute the multimodal emergency information generation method as described above, according to the instructions in the program code.
[0050] The present invention also provides a computer-readable storage medium for storing program code for executing the multimodal emergency information generation method as described in any of the preceding claims.
[0051] As can be seen from the above technical solutions, the present invention has the following advantages:
[0052] This paper presents a multimodal emergency information generation method. First, it acquires emergency-related text and image data and a general task dataset, and then performs multimodal data filtering and expansion on the emergency-related text and image data to obtain text-image question-and-answer pair data. Next, based on the text-image question-and-answer pair data, it constructs text-image instruction refinement data based on thought chain knowledge distillation. Then, based on the text-image instruction refinement data and the general task dataset, it performs secondary instruction refinement on the pre-constructed initial multimodal emergency model to obtain a multimodal emergency vertical domain model. Finally, through the multimodal emergency vertical domain model, it performs image recognition analysis based on retrieval enhancement on the emergency-related query information input by the user, and generates emergency-related response information. By combining a large language model, multimodal technology, retrieval enhancement, and thought chain knowledge distillation, a thought chain data-enhanced multimodal emergency vertical domain model with image recognition and reasoning capabilities in emergency scenarios is constructed. In practical applications, based on the multimodal emergency vertical domain model, accurate and real-time effective emergency-related response information can be quickly generated for user-input emergency-related query information. This not only effectively solves the adaptability problem of traditional methods in complex emergency scenarios and significantly improves the efficiency of emergency information generation, but also enhances the accuracy, practicality, and diversity of emergency information generation. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 This is a flowchart illustrating the steps of a multimodal emergency information generation method.
[0055] Figure 2 This is a schematic diagram of the overall process of a multimodal emergency information generation method;
[0056] Figure 3 This is a structural block diagram of an emergency information generation device based on multimodality. Detailed Implementation
[0057] This invention provides a multimodal emergency information generation method, device, electronic device, and medium to solve or partially solve the problems of poor applicability and targeting of current emergency information generation methods, limitations in information update frequency and content accuracy, and difficulty in meeting real-time and personalized needs.
[0058] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0059] As an example, with the rapid development of artificial intelligence, multimodal application scenarios have gradually become a research hotspot in the field of emergency information generation. Multimodal technology, by integrating various data formats such as images and text, can more comprehensively perceive and understand complex scenarios. In emergency scenarios such as fire prevention and medical emergency response, images are often key information carriers. By identifying the specific environment and potential risk points in images, combined with professional guidance on text generation, users can be provided with more accurate and intuitive solutions. In practice, multimodal applications not only enhance the richness and interest of information expression but also provide strong support for personalized and dynamic generation of emergency information.
[0060] However, existing emergency information generation methods still have significant shortcomings in multimodal applications. For example, the commonly used fixed template method can only handle preset scenarios and has poor applicability. Another example is the manual writing method, which, while offering some flexibility, lacks more efficient analytical capabilities and is less targeted. Furthermore, traditional methods remain limited in terms of information update frequency and content accuracy, making it difficult to meet the demands for real-time and personalized information.
[0061] Further analysis reveals that the fixed template method can only handle preset scenarios and cannot extract information from images, nor can it cope with diverse data inputs in complex environments. While manual writing offers some flexibility, it lacks efficient image analysis capabilities and cannot dynamically combine visual information to generate targeted emergency guidance.
[0062] These shortcomings prevent traditional methods from fully leveraging the potential of multimodal technologies in emergency information generation, limiting their application value in complex emergency scenarios. In this context, introducing an AI system for image recognition and situation analysis in emergency scenarios is particularly necessary. Achieving image recognition and situation analysis in emergency scenarios means that by combining image and text information, it is possible to more comprehensively perceive and analyze complex data in real-world situations. This would not only improve the accuracy of emergency information but also significantly enhance the practicality and response efficiency of emergency information by dynamically generating personalized solutions tailored to different user needs.
[0063] Therefore, one of the core inventive points of this invention is to propose a multimodal emergency information generation method to address the limitations of traditional emergency information generation methods in multimodal applications. By combining a large language model, multimodal technology, retrieval enhancement, and knowledge distillation based on thought chains, a multimodal emergency science popularization large language model with the ability to interpret images and deduce events in emergency scenarios is constructed. In practical applications, based on the multimodal emergency science popularization large language model, accurate and real-time emergency-related response information can be quickly generated for user-input emergency-related queries. The technical solution provided by this invention can effectively solve the adaptability problem of traditional methods in complex emergency scenarios, significantly improve the efficiency of emergency information generation, and enhance the accuracy, practicality, and diversity of emergency information generation.
[0064] Reference Figure 1 The diagram illustrates a flowchart of a multimodal emergency information generation method provided by an embodiment of the present invention, which may specifically include the following steps:
[0065] Step 101: Obtain emergency-related graphic and textual data and general task dataset, and perform multimodal data filtering and expansion on the emergency-related graphic and textual data to obtain graphic and textual question-and-answer pair data;
[0066] This step mainly involves collecting emergency-related graphic and textual data and general task datasets, as well as using multimodal large language models and mind chain technology to achieve data filtering (screening) and data expansion.
[0067] The general task dataset refers to the training data downloaded from the official InternVL2 large language model repository. This data is sourced from the official GitHub repository. This portion of the data is used to preserve the general answering capabilities of the large language model, thereby maintaining model performance.
[0068] For the collection of emergency-related graphic data, this embodiment of the invention mainly adopts a keyword matching strategy. By analyzing the content of the image description, it filters out graphic pairs related to the current vertical domain (i.e., the emergency vertical domain).
[0069] In some embodiments, the collection of emergency-related graphic data can be achieved by executing the following sub-steps S01 to S05:
[0070] Step S01: Construct a keyword thesaurus related to emergency scenarios;
[0071] A keyword lexicon related to emergency scenarios is established using a multi-keyword matching method. Based on this keyword lexicon, during subsequent filtering, the frequency of keyword occurrences in the image description text can be counted. If the frequency of keyword occurrences exceeds a preset threshold, the current image-text pair is recorded as a valid image-text pair.
[0072] Step S02: Obtain the open-source image and text dataset from the Internet. The open-source image and text dataset contains multiple open-source images, and each open-source image corresponds to one image description text.
[0073] This invention primarily filters and selects emergency-related images from open-source online image and text datasets. In these datasets, each data entry contains an image and its corresponding description (i.e., image description text). Therefore, it is possible to first collect open-source online image and text datasets, such as the "Noah-Wukong" Chinese image and text dataset (a dataset specifically designed for multimodal learning and research).
[0074] Step S03: Perform keyword matching and filtering on each image description text based on the keyword thesaurus;
[0075] Next, a multi-keyword matching strategy can be used to selectively filter image and text data that are highly relevant to emergency response from open-source image and text datasets on the internet.
[0076] Step S04: When the number of times the keywords in the image description text appear is greater than the preset threshold, the image description text is determined to be valid image description text, and the open source image corresponding to the valid image description text is a valid emergency image.
[0077] Step S05: Combine the valid image description text with the valid emergency image as a valid image-text pair, and integrate all valid image-text pairs into emergency-related image-text data.
[0078] Through data integration, the emergency-related graphic data in this embodiment of the invention can include multiple valid graphic pairs, each of which contains a valid image description text and a valid emergency image.
[0079] For example, taking fire-related image and text data as an example, by selecting keywords ["fire", "flammable", "explosive", "fire scene", "fire", "firefighting", "fire extinguishing", "fire prevention", "burning"], and filtering by keyword matching, fire-related image and text pairs can be obtained. Image types can include images of flammable and explosive materials, images of fire hazards, images of fire scenes, fire safety education images, etc.
[0080] Once the relevant emergency-related image and text data is obtained, various refined multimodal emergency tasks can be designed for users, and multimodal large language models can be used to classify the images and assign tasks.
[0081] In some embodiments, the process of performing multimodal data filtering and expansion on emergency-related graphic data to obtain graphic question-and-answer pair data can be achieved by executing the following sub-steps S11 to S13:
[0082] Step S11: Integrate various refined multimodal emergency tasks, and construct instruction text based on task division instruction requirements, task division category list, and specific examples;
[0083] The instruction text designed in this embodiment of the invention conforms to the format of "instruction requirements for task division + list of task division categories + specific examples". Taking fire-related image and text pairs as an example, the designed instruction text example is: "Which of the following points does the content of this image best fit? Options are as follows: flammable material analysis and fire hazard analysis; fire accident in progress; fire accident ruins; fire safety education; completely unrelated to fire. Please only answer the content of the options. Example answer: fire accident ruins."
[0084] Step S12: Based on the instruction text, perform task division on all valid emergency images using a pre-trained multimodal task division model, remove emergency-irrelevant images based on the task division results, and annotate all the remaining valid emergency images to obtain the corresponding target emergency images.
[0085] In this step, the pre-trained multimodal large language model InternVL2-76B is used as a multimodal task partitioning model to filter valid emergency images and complete task partitioning. Specifically, the instruction text and each valid emergency image are input into the multimodal task partitioning model InternVL2-76B. The multimodal task partitioning model InternVL2-76B can then perform task partitioning on these input images and filter out emergency-irrelevant images.
[0086] Taking fire-related image-text pairs as an example, the task categories can mainly include five types: "Flammable material analysis and fire hazard analysis; During a fire accident; Fire accident ruins; Fire safety education; Completely unrelated to fire." Images classified as "Completely unrelated to fire" are considered invalid (i.e., emergency-irrelevant images) and are discarded. All remaining valid emergency images are then annotated to perform data augmentation and obtain the corresponding target emergency images.
[0087] Step S13: Use the image description text corresponding to the target emergency image as the target image description text, use the target emergency image and the target image description text as an image-text question-and-answer pair, and integrate all image-text question-and-answer pairs into image-text question-and-answer pair data.
[0088] Step 102: Based on the text-image question-and-answer pair data, construct text-image instruction fine-tuning data based on knowledge distillation of the mind chain;
[0089] This step mainly involves using the collected image-text question-and-answer pair data to construct image-text instruction fine-tuning data for training a multimodal emergency vertical domain large model.
[0090] To construct fine-tuned data for graphic and textual instructions, this embodiment of the invention mainly combines multimodal large language models and mind chain technology to expand and generate emergency-related graphic and textual data, thereby enhancing the knowledge of the large language model.
[0091] In some embodiments, the process of constructing image-text instruction fine-tuning data based on mind chain knowledge distillation according to image-text question-and-answer pair data can be achieved by executing the following sub-steps S21 to S22:
[0092] Step S21: Construct an initial graph-text description model based on the multimodal large language model framework;
[0093] In this embodiment of the invention, the multimodal large language model InternVL2-76B is selected as the framework to construct an initial graph-text description model.
[0094] Step S22: Using image-text question-and-answer pairs, train the initial image-text description model by image-text expansion generation based on knowledge distillation of the thought chain, and generate image-text instruction fine-tuning data based on the trained image-text description model.
[0095] Step S22 mainly realizes the generation of emergency vertical domain image-text pair data (i.e., image-text instruction fine-tuning data) based on the image-text question-and-answer data and the image content and image description text of each image. The extended generation of image-text data can be divided into three steps: first, the image content is described by the initial image-text description model; then, the initial image-text description model compares the text content and corrects the image description content; finally, the trained image-text description model generates emergency vertical domain image-text pair data related to the specified task based on the corrected image description content.
[0096] Specifically, the process of using text-image question-and-answer pairs of data to train the initial text-image description model through text-image expansion generation based on knowledge distillation of thought chains, and then generating text-image instructions to fine-tune the data based on the trained text-image description model, can be achieved by executing the following sub-steps S221 to S224:
[0097] Step S221: For each image-text question-and-answer pair, based on the task division results corresponding to the target emergency image, perform image description on the target emergency image through the initial image-text description model, and generate the image description model output text;
[0098] The initial image description model is used to describe the target emergency image. The specific task prompt is: "Please describe the content of the image in detail; this image is related to {TASK}." Here, TASK refers to the task division result from step S12, such as "fire safety education." Next, the output of the image description model is recorded, denoted as the image description content. That is, the image description model outputs text.
[0099] Step S222: Compare the content of the output text of the image description model with the target image description text, and based on the content comparison difference, correct the text of the output text of the image description model to obtain the corrected image description text;
[0100] The image description model's output text is compared with the target image description text, and the model's output image description content is corrected. The specific task prompt for this process is: "The content of this image is related to {CAPTION}. Your previous description of this image is {MLLM_CAPTION}. Please analyze further and correct your description." Here, CAPTION refers to the original image description text (i.e., the target image description text), and MLLM_CAPTION refers to the image description model's description of the image. Next, the output of the image description model is recorded, denoted as the corrected image description text.
[0101] Step S223: When the image descriptions and text corrections for all image-text question-answer pairs are completed, the trained image-text description model is obtained;
[0102] Step S224: Based on the trained image-text description model, generate emergency vertical image-text pairs corresponding to each image description correction text, and integrate all emergency vertical image-text pairs into image-text instruction fine-tuning data.
[0103] Step S224 mainly involves designing different task prompts based on different task division pairs, so that the trained text and image description model can generate more refined emergency vertical domain question and answer pairs.
[0104] For example, assuming the task classification result of the target emergency image is "Flammable Material Analysis and Fire Hazard Analysis", the corresponding task Prompt could be: "Based on this image, and its content being related to {MLLM_REFINE_CAPTION}, generate a question-and-answer pair related to flammable material analysis and fire hazard analysis. The answer should be longer than 200 characters. The format is similar to the example below: {Question: "What fire hazards are there in the image?", Answer: "The battery is too close to the stove."}
[0105] For example, assuming the task classification result of the target emergency image is "fire safety education," its task prompt could be: "Based on this image, which is related to {MLLM_REFINE_CAPTION}, generate a question-and-answer pair related to fire safety education. The answer should be longer than 200 characters. The format is similar to the example below: {Question: "What basic escape steps should be followed to ensure personal safety in the event of a fire at home?", Answer: "In the event of a fire at home, it is crucial to remain calm and quickly take the correct escape steps. Once a fire is discovered, immediately call the local fire department (such as 119 in China), clearly stating the specific address of the fire, the size of the fire, and whether anyone is trapped."}
[0106] Here, MLLM_REFINE_CAPTION refers to the image description correction text. Ultimately, from the output of the trained image description model InternVL2-76B, all emergency vertical domain image-text pairs generated around the image content and task category can be obtained, i.e., image-text instruction fine-tuning data.
[0107] Step 103: Based on the graphic instruction fine-tuning data and the general task dataset, perform secondary instruction fine-tuning on the pre-constructed initial multimodal emergency model to obtain a multimodal emergency vertical domain model;
[0108] In some embodiments, the process of performing secondary instruction fine-tuning on the pre-built initial multimodal emergency model based on the graphic instruction fine-tuning data and the general task dataset to obtain a multimodal emergency vertical domain model can be achieved by executing the following sub-steps S31 to S34:
[0109] Step S31: Construct an initial multimodal emergency response model based on the multimodal large language model framework. The initial multimodal emergency response model introduces the LoRA model fine-tuning module.
[0110] An initial multimodal emergency response model was constructed based on the multimodal large language model InternVL2-8B, and a LoRA (Low-Rank Adaptation) model fine-tuning module was introduced into the initial multimodal emergency response model.
[0111] In other words, the model requiring fine-tuning in this embodiment of the invention is the initial multimodal emergency response model InternVL2-8B. Specifically, the initial multimodal emergency response model can be fine-tuned by combining text and graphical instruction fine-tuning data with a general task dataset.
[0112] It should be noted that the initial multimodal emergency model can also be constructed using other sizes of multimodal large language models, such as InternVL2-4B. The purpose of choosing InternVL2-8B in this invention is that, compared to other sizes, the 8B model offers higher inference cost-effectiveness, can perform inference on a single consumer-grade graphics card, and achieves better inference performance and lower loss than the largest model (InternVL2-76B). It is understood that this invention does not impose any limitations on this.
[0113] In specific training, this invention introduces a LoRA module into the initial multimodal emergency response model. The LoRA module is a technique for fine-tuning large pre-trained models. Unlike traditional fine-tuning methods, the LoRA module achieves efficient adaptive adjustment by updating only a small subset of parameters in the model. By introducing the LoRA module, computational resources and time costs can be significantly reduced while maintaining the model's competitive performance.
[0114] Step S32: Set the fine-tuning training parameters for the LoRA model fine-tuning module;
[0115] Set the fine-tuning training parameters for the LoRA model fine-tuning module according to the recommended parameters in the official InternVL2 training documentation.
[0116] Step S33: Mix the fine-tuned image and text instruction data with the general task dataset to obtain an emergency image and text hybrid dataset;
[0117] To maintain InternVL2's general text and image question answering capabilities, during model training, it is necessary to mix the text and image instruction fine-tuning data (which can be understood as a vertical task dataset) with the general task dataset and then use it for fine-tuning training.
[0118] Specifically, first, all fine-tuning data of InternVL2 (i.e., fine-tuning data of text and image instructions) is retained, and the general task dataset is modified into the training data format of InternVL2. The two types of data are mixed, and finally the new data is organized into the JSON file of the original data for subsequent fine-tuning training.
[0119] Step S34: Train the initial multimodal emergency model using the emergency image and text hybrid dataset, and fine-tune the initial multimodal emergency model with secondary instructions based on the fine-tuning of training parameters during the training process to obtain the multimodal emergency vertical domain model.
[0120] Therefore, in the knowledge enhancement stage, this embodiment of the invention combines a multimodal large language model with mind chain knowledge enhancement technology to filter out invalid information and expand mind chain image annotations in emergency-related text and image data. The data obtained in the knowledge enhancement stage is then used for knowledge injection during the fine-tuning training of the multimodal large language model. This allows the fine-tuned multimodal large language model to be combined with retrieval enhancement generation technology to build a more reliable multimodal emergency knowledge popularization platform.
[0121] Step 104: Using the multimodal emergency vertical domain model, perform image recognition analysis based on retrieval enhancement on the emergency-related query information input by the user, and generate emergency-related response information.
[0122] In practical applications, this invention introduces a small-scale multimodal large language model, InternVL2-1B, for image annotation to provide a content summary of user-input images. In subsequent processing, the image annotations and user-input questions can be used as retrieval queries to retrieve emergency-related information from a vector database. The retrieved keys can be stored for source tracing output (i.e., indicating which information the model's current response is related to). These stored keys can serve as quick search keywords for subsequent source tracing. The emergency-related information data can then be used to enhance the context of the user-input questions, thereby improving retrieval and assisting the model in emergency-related question-and-answering, generating emergency-related response information, and ultimately improving the relevance of the multimodal large language model's responses.
[0123] Based on the foregoing discussion, in some embodiments, the process of performing retrieval-enhanced image recognition analysis on user-input emergency-related query information and generating emergency-related response information using a multimodal emergency vertical domain model can be achieved by executing the following sub-steps S41 to S45:
[0124] Step S41: Obtain emergency-related query information input by the user. The emergency-related query information includes the image to be analyzed and the text information to be answered.
[0125] Step S42: Input the image to be analyzed into a pre-trained multimodal image annotation model to perform image annotation and obtain an overview of the image content corresponding to the image to be analyzed;
[0126] Step S43: Using the image content overview and the text information to be answered as key information for retrieval, retrieve emergency-related data from the vector database;
[0127] Step S44: Input emergency-related data, images to be analyzed, and text information to be answered into the multimodal emergency vertical domain model;
[0128] Step S45: The multimodal emergency vertical domain model is used to perform image recognition analysis of the image to be analyzed, focusing on the text information to be answered. During the image recognition analysis, emergency-related data is used to enhance the context of the text information to be answered, generating emergency-related response information.
[0129] During the content generation phase, by comprehensively analyzing user-input images and text information, and combining retrieval-enhanced generation technology, relevant professional knowledge is quickly acquired to generate logically clear, visually appealing, and easily understandable emergency-related content, providing reliable generation basis. The generated results include, but are not limited to, knowledge-based Q&As and targeted processing steps for emergency scenarios. All outputs are presented in an intuitive and concise manner to help users quickly grasp key information and take effective action. By combining image understanding and language generation technologies, a deep understanding of emergency scenarios and accurate knowledge generation are achieved. This provides efficient and intelligent solutions for complex and ever-changing emergency needs, and is widely applicable to various emergency scenarios such as fire fighting, construction accidents, and medical emergencies.
[0130] From an overall technical perspective, this invention designs two key personalized multimodal tasks. One is emergency scenario hazard identification (such as fire hazard identification), and the other is multimodal emergency knowledge dissemination. The former, based on image understanding and natural language reasoning, can provide users with accurate hazard detection services. The latter, through optimized information presentation, enables users to obtain immediate, relevant, and easily understandable assistance, thereby improving the overall efficiency of emergency response.
[0131] Taking fire-related emergency scenarios as an example, the technical solution provided in this invention, through an advanced large language model combined with pixel-level image understanding technology, supports user input image queries and intelligently identifies potential fire risk points in the images, such as improper storage of flammable materials or aging electrical wiring. Simultaneously, the model can generate professional prevention suggestions based on the identification results, helping non-professionals to conduct self-checks and make corrections, extending fire prevention from the professional field to daily life and making it an important part of every family. Furthermore, popularizing fire knowledge through a combination of text and images transforms emergency information into more intuitive, interesting, and professionally detailed learning content. By integrating text descriptions with multimedia elements, the public can easily learn fire prevention knowledge, improving memory and practicality, thereby comprehensively enhancing the public's emergency response capabilities and self-protection awareness.
[0132] In this invention, addressing the limitations of traditional emergency information generation methods in multimodal applications, a multimodal emergency information generation method is proposed. By combining a large language model, multimodal technology, retrieval enhancement, and thought chain knowledge distillation, a multimodal emergency science popularization large language model based on thought chain data enhancement is constructed, enabling image recognition and situational analysis in emergency scenarios. In practical applications, based on user-input emergency-related queries, the multimodal emergency science popularization large language model can quickly generate accurate and real-time effective emergency-related responses. By adopting the technical solution provided by this invention, the adaptability problem of traditional methods in complex emergency scenarios can be effectively solved, significantly improving the efficiency of emergency information generation and enhancing the accuracy, practicality, and diversity of emergency information generation.
[0133] For better illustration, refer to Figure 2 Taking fire-related graphic data as an example, this paper illustrates the overall process of a multimodal emergency information generation method provided by an embodiment of the present invention. It should be noted that this embodiment only provides a brief description of the general process of multimodal emergency information generation. The specific implementation process of each step can be understood by referring to the relevant content in the foregoing embodiments, and will not be elaborated here. It is understood that the present invention does not impose any limitations on this.
[0134] S1: Collect emergency-related graphic data for fine-tuning.
[0135] First, we obtained an open-source image and text dataset from the internet, and then filtered it by frequency based on text keywords ["fire", "flammable", "explosive", "fire scene", "fire", "firefighting", "fire extinguishing", "fire prevention", "burning"] to obtain highly relevant emergency image and text data.
[0136] Next, based on the pre-trained multimodal task partitioning model InternVL2-76B, the obtained emergency highly relevant image and text data is subjected to multimodal data filtering and expansion (image classification and task partitioning) to obtain image and text question-answering pair data.
[0137] S2: Construct image-text instruction fine-tuning data based on knowledge distillation of thought chains to train a multimodal vertical domain large model.
[0138] We selected the multimodal large language model InternVL2-76B as the framework to construct an initial graph-text description model.
[0139] Using image-text question-and-answer pairs, the initial image-text description model InternVL2-76B was trained by knowledge distillation based on thought chain (image description and corrected description content). Based on the trained image-text description model, emergency-related question-and-answer pairs (emergency vertical domain image-text pairs) were generated to obtain the vertical domain task dataset (i.e. image-text instruction fine-tuning data).
[0140] The pre-acquired general task dataset and the constructed vertical domain task dataset are mixed together. The mixed dataset is used to train the initial multimodal emergency model InternVL2-8B, which incorporates the LoRA model fine-tuning module. During the training process, the initial multimodal emergency model InternVL2-8B is fine-tuned based on the fine-tuned training parameters to obtain the multimodal emergency vertical domain model InternVL2-8B.
[0141] S3: Combining retrieval enhancement generation technology with a multimodal large language model, an emergency information question-and-answer system will be built.
[0142] Obtain emergency-related query information input by the user, which includes images to be analyzed and text information to be answered;
[0143] The image to be analyzed is input into the pre-trained multimodal image annotation model InternVL2-1B for image annotation, and an image content overview corresponding to the image to be analyzed is obtained.
[0144] Using the image content overview and the text information to be answered as key information for retrieval, emergency-related data was retrieved from the vector database.
[0145] Input emergency-related data, images to be analyzed, and text information to be answered into the multimodal emergency vertical domain model InternVL2-8B;
[0146] The multimodal emergency vertical domain model InternVL2-8B is used to perform image recognition analysis on the image to be analyzed, focusing on the text information to be answered. During the image recognition analysis, emergency-related data is used to enhance the context of the text information to be answered, generating emergency-related response information.
[0147] Reference Figure 3 The diagram illustrates a structural block diagram of a multimodal emergency information generation device provided by an embodiment of the present invention, which may specifically include:
[0148] The multimodal data filtering and expansion unit 301 is used to acquire emergency-related graphic and textual data and general task datasets, and to perform multimodal data filtering and expansion on the emergency-related graphic and textual data to obtain graphic and textual question-and-answer pair data.
[0149] The image and text instruction fine-tuning data construction unit 302 is used to construct image and text instruction fine-tuning data based on the knowledge distillation of the thought chain according to the image and text question and answer pair data;
[0150] The secondary instruction fine-tuning unit 303 is used to perform secondary instruction fine-tuning on the pre-constructed initial multimodal emergency model based on the graphic instruction fine-tuning data and the general task dataset to obtain a multimodal emergency vertical domain model.
[0151] The image recognition and analysis unit 304 is used to perform image recognition analysis based on retrieval enhancement on the emergency-related query information input by the user through the multimodal emergency vertical domain model, and generate emergency-related response information.
[0152] In one optional embodiment, the emergency-related graphic data includes multiple valid graphic-text pairs, each valid graphic-text pair including a valid image description text and a valid emergency image; the multimodal data filtering and expansion unit 301 includes:
[0153] The instruction text construction unit is used to integrate various refined multimodal emergency tasks and construct instruction text based on task division instruction requirements, task division category listing, and specific examples.
[0154] The target emergency image generation unit is used to divide all the effective emergency images into tasks based on the instruction text using a pre-trained multimodal task partitioning model, remove emergency-irrelevant images based on the task partitioning results, and annotate all the remaining effective emergency images to obtain the corresponding target emergency image.
[0155] The image-text question-and-answer pair determination unit is used to take the image description text corresponding to the target emergency image as the target image description text, and to take the target emergency image and the target image description text as an image-text question-and-answer pair;
[0156] The image-text question-and-answer pair data integration unit is used to integrate all the image-text question-and-answer pairs into image-text question-and-answer pair data.
[0157] In one optional embodiment, the graphic instruction fine-tuning data construction unit 302 includes:
[0158] The initial text-image description model building unit is used to build an initial text-image description model based on the multimodal large language model framework.
[0159] The image-text instruction fine-tuning data generation unit is used to train the initial image-text description model by using the image-text question-and-answer pair data and performing image-text expansion generation training based on the knowledge distillation of the thought chain, and to generate image-text instruction fine-tuning data based on the trained image-text description model.
[0160] In one optional embodiment, the graphic instruction fine-tuning data generation unit includes:
[0161] The image description unit is used to describe the target emergency image for each image-text question-and-answer pair based on the task division result corresponding to the target emergency image, and generate the image description model output text through the initial image-text description model.
[0162] The text correction unit is used to compare the content of the text output by the image description model with the target image description text, and correct the text of the text output by the image description model based on the content comparison difference to obtain the corrected image description text.
[0163] The output unit of the trained image-text description model is used to obtain the trained image-text description model when the image description and text correction of all the image-text question-answer pairs are completed.
[0164] The emergency vertical image-text pair generation unit is used to generate emergency vertical image-text pairs corresponding to each of the image description correction texts based on the trained image-text description model, and to integrate all the emergency vertical image-text pairs into image-text instruction fine-tuning data.
[0165] In one optional embodiment, the secondary instruction fine-tuning unit 303 includes:
[0166] An initial multimodal emergency model construction unit is used to construct an initial multimodal emergency model based on a multimodal large language model framework. The initial multimodal emergency model incorporates a LoRA model fine-tuning module.
[0167] The fine-tuning training parameter setting unit is used to set the fine-tuning training parameters of the LoRA model fine-tuning module;
[0168] The data mixing unit is used to mix the fine-tuning data of the graphic instructions with the general task dataset to obtain an emergency graphic mixed dataset;
[0169] The model training and fine-tuning unit is used to train the initial multimodal emergency model using the emergency image and text hybrid dataset, and during the training process, to perform secondary fine-tuning of the initial multimodal emergency model based on the fine-tuning training parameters to obtain a multimodal emergency vertical domain model.
[0170] In one optional embodiment, the image recognition and analysis unit 304 includes:
[0171] An emergency-related query information acquisition unit is used to acquire emergency-related query information input by the user, wherein the emergency-related query information includes an image to be analyzed and text information to be answered;
[0172] The image content summary generation unit is used to input the image to be analyzed into a pre-trained multimodal image annotation model for image annotation and obtain the image content summary corresponding to the image to be analyzed.
[0173] An emergency-related data retrieval unit is used to retrieve emergency-related data from a vector database by using the image content overview and the text information to be answered as key retrieval information.
[0174] The data input unit is used to input the emergency-related data, the image to be analyzed, and the text information to be answered into the multimodal emergency vertical domain model.
[0175] The image recognition and analysis subunit is used to perform image recognition and analysis on the image to be analyzed around the text information to be answered using the multimodal emergency vertical domain model. During the image recognition and analysis process, the emergency-related data is used to enhance the context of the text information to be answered, and emergency-related response information is generated.
[0176] In one optional embodiment, the emergency information generation device further includes:
[0177] The keyword thesaurus building unit is used to build a keyword thesaurus related to emergency scenarios;
[0178] The network open source image and text dataset acquisition unit is used to acquire the network open source image and text dataset, which contains multiple open source images, and each open source image corresponds to an image description text.
[0179] The keyword matching and filtering unit is used to perform keyword matching and filtering on each image description text based on the keyword thesaurus.
[0180] The effective image description text determination unit is used to determine that the image description text is effective image description text when the number of times the keywords in the image description text appear is greater than a preset threshold, and the open source image corresponding to the effective image description text is a valid emergency image.
[0181] The effective image-text pair determination unit is used to determine the effective image description text and the effective emergency image as a valid image-text pair;
[0182] The emergency-related graphic data integration unit is used to integrate all the aforementioned valid graphic pairs into emergency-related graphic data.
[0183] As the device embodiment is basically similar to the method embodiment, it is described in a relatively simple way. For relevant details, please refer to the description of the method embodiment above.
[0184] This invention also provides an electronic device, which includes a processor and a memory:
[0185] The memory is used to store program code and transfer the program code to the processor;
[0186] The processor is used to execute the multimodal emergency information generation method of any embodiment of the present invention according to the instructions in the program code.
[0187] This invention also provides a computer-readable storage medium for storing program code for executing the multimodal emergency information generation method of any embodiment of this invention.
[0188] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0189] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units through some interfaces, and may be electrical, mechanical, or other forms.
[0190] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0191] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0192] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0193] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating emergency information based on multimodal approaches, characterized in that, include: Acquire emergency-related graphic and textual data and general task datasets, and perform multimodal data filtering and expansion on the emergency-related graphic and textual data to obtain graphic and textual question-and-answer pair data; Based on the aforementioned text-image question-and-answer pair data, construct text-image instruction fine-tuning data based on knowledge distillation of the thought chain; The pre-built initial multimodal emergency model is fine-tuned a second time using the graphic and textual instruction fine-tuning data and the general task dataset to obtain a multimodal emergency vertical domain model. The multimodal emergency vertical domain model is used to perform image recognition analysis based on retrieval enhancement on emergency-related query information input by users, and to generate emergency-related response information. The step of constructing image-text instruction fine-tuning data based on the image-text question-and-answer pair data and the thought chain knowledge distillation includes: An initial image-text description model is constructed based on a multimodal large language model framework; Using the aforementioned text-image question-and-answer pair data, the initial text-image description model is trained by text-image expansion generation based on knowledge distillation of the thought chain, and text-image instruction fine-tuning data is generated based on the trained text-image description model. The process involves using the multimodal emergency vertical domain model to perform enhanced image recognition analysis on the emergency-related query information input by the user, and generating emergency-related response information, including: Obtain emergency-related query information input by the user, which includes the image to be analyzed and the text information to be answered; The image to be analyzed is input into a pre-trained multimodal image annotation model for image annotation to obtain an overview of the image content corresponding to the image to be analyzed. Using the image content overview and the text information to be answered as key information for retrieval, emergency-related data is retrieved from the vector database. The emergency-related data, the image to be analyzed, and the text information to be answered are input into the multimodal emergency vertical domain model; The multimodal emergency vertical domain model is used to perform image recognition analysis on the image to be analyzed, focusing on the text information to be answered. During the image recognition analysis, the emergency-related data is used to enhance the context of the text information to be answered, generating emergency-related response information.
2. The emergency information generation method based on multimodal processing according to claim 1, characterized in that, The emergency-related graphic data includes multiple valid graphic-text pairs, each of which contains a valid image description text and a valid emergency image; the multimodal data filtering and expansion of the emergency-related graphic data to obtain graphic-text question-and-answer pair data includes: By integrating various refined and multimodal emergency tasks, instruction texts are constructed based on task-specific instruction requirements, task category listings, and specific examples. Based on the instruction text, all the effective emergency images are divided into tasks using a pre-trained multimodal task partitioning model. Based on the task partitioning results, emergency-irrelevant images are removed, and all the remaining effective emergency images are labeled to obtain the corresponding target emergency images. The image description text corresponding to the target emergency image is used as the target image description text, and the target emergency image and the target image description text are used as an image-text question-and-answer pair; All the aforementioned text-image question-and-answer pairs will be integrated into text-image question-and-answer pair data.
3. The emergency information generation method based on multimodality according to claim 2, characterized in that, The step of using the image-text question-and-answer pair data to train the initial image-text description model through image-text expansion generation based on knowledge distillation of thought chains, and generating image-text instruction fine-tuning data based on the trained image-text description model, includes: For each of the image-text question-and-answer pairs, based on the task division results corresponding to the target emergency image, the target emergency image is described using the initial image-text description model, and the image description model output text is generated. The output text of the image description model is compared with the target image description text. Based on the differences in the content comparison, the output text of the image description model is corrected to obtain the corrected image description text. When the image descriptions and text corrections for all the image-text question-and-answer pairs are completed, the trained image-text description model is obtained. Based on the trained image-text description model, emergency vertical image-text pairs are generated for each of the image description correction texts, and all the emergency vertical image-text pairs are integrated into image-text instruction fine-tuning data.
4. The emergency information generation method based on multimodal processing according to claim 1, characterized in that, The step of performing secondary instruction fine-tuning on the pre-constructed initial multimodal emergency model based on the image and text instruction fine-tuning data and the general task dataset to obtain a multimodal emergency vertical domain model includes: An initial multimodal emergency response model is constructed based on a multimodal large language model framework, and a LoRA model fine-tuning module is introduced into the initial multimodal emergency response model; Set the fine-tuning training parameters of the LoRA model fine-tuning module; The fine-tuned data of the graphic and text instructions is mixed with the general task dataset to obtain an emergency graphic and text hybrid dataset; The initial multimodal emergency model is trained using the aforementioned emergency image and text hybrid dataset. During the training process, the initial multimodal emergency model is fine-tuned using secondary instructions based on the fine-tuning training parameters to obtain a multimodal emergency vertical domain model.
5. The emergency information generation method based on multimodality according to any one of claims 1 to 4, characterized in that, The process of collecting emergency-related graphic and textual data includes: Build a keyword thesaurus related to emergency scenarios; Obtain an open-source image and text dataset from the internet, which contains multiple open-source images, each of which corresponds to a descriptive text for an image. Based on the keyword lexicon, keyword matching and filtering are performed on each image description text; When the number of occurrences of the keywords in the image description text exceeds a preset threshold, the image description text is determined to be valid image description text, and the open-source image corresponding to the valid image description text is a valid emergency image. The effective image description text and the effective emergency image are used as a valid image-text pair; All the aforementioned valid image and text pairs will be integrated into emergency-related image and text data.
6. A multimodal emergency information generation device, characterized in that, include: The multimodal data filtering and expansion unit is used to acquire emergency-related graphic and textual data and general task datasets, and to perform multimodal data filtering and expansion on the emergency-related graphic and textual data to obtain graphic and textual question-and-answer pair data; The image and text instruction fine-tuning data construction unit is used to construct image and text instruction fine-tuning data based on the knowledge distillation of the thought chain according to the image and text question and answer pair data; The secondary instruction fine-tuning unit is used to perform secondary instruction fine-tuning on the pre-constructed initial multimodal emergency model based on the graphic instruction fine-tuning data and the general task dataset to obtain a multimodal emergency vertical domain model. The image recognition and analysis unit is used to perform image recognition and analysis based on retrieval enhancement on the emergency-related query information input by the user through the multimodal emergency vertical domain model, and generate emergency-related response information; The image and text command fine-tuning data construction unit includes: The initial text-image description model building unit is used to build an initial text-image description model based on the multimodal large language model framework. The image and text instruction fine-tuning data generation unit is used to use the image and text question-and-answer pair data to train the initial image and text description model based on the knowledge distillation of the thinking chain, and to generate image and text instruction fine-tuning data based on the trained image and text description model. The image recognition and analysis unit includes: An emergency-related query information acquisition unit is used to acquire emergency-related query information input by the user, wherein the emergency-related query information includes an image to be analyzed and text information to be answered; The image content summary generation unit is used to input the image to be analyzed into a pre-trained multimodal image annotation model for image annotation and obtain the image content summary corresponding to the image to be analyzed. An emergency-related data retrieval unit is used to retrieve emergency-related data from a vector database by using the image content overview and the text information to be answered as key retrieval information. The data input unit is used to input the emergency-related data, the image to be analyzed, and the text information to be answered into the multimodal emergency vertical domain model. The image recognition and analysis subunit is used to perform image recognition and analysis on the image to be analyzed around the text information to be answered using the multimodal emergency vertical domain model. During the image recognition and analysis process, the emergency-related data is used to enhance the context of the text information to be answered, and emergency-related response information is generated.
7. An electronic device, characterized in that, The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the multimodal emergency information generation method according to any one of the claims 1-5 according to the instructions in the program code.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program code for executing the multimodal emergency information generation method according to any one of claims 1-5.
Citation Information
Patent Citations
Chinese and cross-language query expansion method based on retrieval enhancement and knowledge distillation
CN118673095A
Label calculation method based on hybrid expert LLM
CN118760736A