Emergency information generation method and device based on multiple modes, electronic equipment and medium
Through multimodal data screening, thinking chain knowledge distillation and secondary instruction adjustment of multimodal emergency model, the problem of insufficient applicability and targeting of existing emergency information generation methods in multimodal applications is solved, and efficient, accurate and personalized emergency information generation is achieved.
Patent Information
- Application Number
- CN202510093796.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The existing emergency information generation methods are poor in multimodal applications, and it is difficult to meet real-time and personalized needs.
By obtaining emergency-related graphic data and general task data sets, multimodal data screening and expansion, building graphic instruction fine-tuning data based on thinking chain knowledge distillation, and performing secondary command fine-tuning of the initial multimodal emergency model to obtain a multimodal emergency vertical domain model, which is used to perform search-enhanced graph recognition analysis based on search-enhanced query information input by users and generate emergency-related response information.
It significantly improves the efficiency of emergency information generation, enhances the accuracy, practicality and diversity of emergency information generation, and can quickly generate accurate and real-time and effective emergency response information.
Smart Images

Figure CN119938858A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multimodal-based emergency information generation method, device, electronic equipment and medium. Background Art
[0002] In recent years, as society's requirements for emergency response capabilities continue to increase, how to efficiently generate popular science information for emergency situations has become one of the important application directions of artificial intelligence technology. In emergency scenarios such as fires, building accidents, and medical emergencies, timely providing the public with scientific and accurate emergency knowledge and personalized response plans can not only effectively reduce disaster losses and improve the utilization rate of social public resources, but also enhance the public's understanding and participation in the entire emergency response process.
[0003] With the rapid development of artificial intelligence, multimodal application scenarios have gradually become a research hotspot in the field of emergency information generation. Multimodal technology can perceive and understand complex scenes more comprehensively by integrating multiple data forms such as images and text. In emergency scenarios such as fire prevention and medical first aid, images are often key information carriers. By identifying the specific environment and potential risk points in the picture, and combining text to generate professional guidance, users can be provided with more accurate and intuitive solutions. In actual situations, multimodal applications not only enhance the richness and fun of information expression, but also provide strong support for personalized and dynamic generation of emergency information.
[0004] However, the existing emergency information generation methods still have significant deficiencies in multimodal applications. For example, the currently commonly used fixed template method can only handle preset scenarios and has poor applicability. Another example is the manual writing method. Although it has a certain degree of flexibility, it lacks more efficient analysis capabilities and has poor pertinence. At the same time, traditional methods are still limited in the frequency of information updates and content accuracy, and it is difficult to meet the needs of real-time and personalization. Summary of the invention
[0005] The present invention provides a multimodal-based emergency information generation method, device, electronic device and medium, which are used to solve or partially solve the problems that the current emergency information generation method has poor applicability and pertinence, is limited in information update frequency and content accuracy, and is difficult to meet real-time and personalized needs.
[0006] The present invention provides a multi-modal emergency information generation method, the method comprising:
[0007] Acquire emergency-related graphic and text data and general task data sets, and perform multimodal data screening and expansion on the emergency-related graphic and text data to obtain graphic and text question-answer pair data;
[0008] According to the picture-text question-answer pair data, construct picture-text instruction fine-tuning data based on thought chain knowledge distillation;
[0009] Performing secondary instruction fine-tuning on the pre-built initial multimodal emergency model according to the graphic instruction fine-tuning data and the general task data set to obtain a multimodal emergency vertical domain model;
[0010] Through the multimodal emergency vertical domain model, the emergency-related query information input by the user is subjected to retrieval-enhanced graph analysis, and emergency-related response information is generated.
[0011] Optionally, the emergency-related image-text data includes a plurality of valid image-text pairs, each of which includes a valid image description text and a valid emergency image; the multimodal data screening and expansion of the emergency-related image-text data to obtain image-text question-answer pair data includes:
[0012] Integrate a variety of refined multi-modal emergency tasks, and construct instruction texts based on task division instruction requirements, task division category enumeration, and specific examples;
[0013] Based on the instruction text, all the valid emergency images are divided into tasks through a pre-trained multimodal task division model, emergency irrelevant images are eliminated based on the task division results, and all the retained valid emergency images are respectively annotated to obtain corresponding target emergency images;
[0014] The image description text corresponding to the target emergency image is used as the target image description text, and the target emergency image and the target image description text are used as an image-text question-answer pair;
[0015] All of the picture-text question-answer pairs are integrated into picture-text question-answer pair data.
[0016] Optionally, constructing the image-text instruction fine-tuning data based on the thought chain knowledge distillation according to the image-text question-answer pair data includes:
[0017] Build an initial picture and text description model based on the multimodal large language model framework;
[0018] The image-text question-answer pair data is used to perform image-text extension generation training based on thought chain knowledge distillation on the initial image-text description model, and image-text instruction fine-tuning data is generated based on the trained image-text description model.
[0019] Optionally, the adopting the picture-text question-answer pair data to perform picture-text extension generation training based on thought chain knowledge distillation on the initial picture-text description model, and generating picture-text instruction fine-tuning data based on the trained picture-text description model, includes:
[0020] For each of the image-text question-answer pairs, based on the task division result corresponding to the target emergency image, the target emergency image is described by the initial image-text description model to generate an image description model output text;
[0021] Comparing the content of the image description model output text with the target image description text, and based on the content comparison difference, performing text correction on the image description model output text to obtain an image description correction text;
[0022] When the image description and text correction of all the image-text question-answer pairs are completed, a trained image-text description model is obtained;
[0023] Based on the trained image-text description model, emergency vertical domain image-text pairs corresponding to each image description correction text are generated respectively, and all the emergency vertical domain image-text pairs are integrated into image-text instruction fine-tuning data.
[0024] Optionally, performing secondary instruction fine-tuning on the pre-built initial multimodal emergency model according to the graphic instruction fine-tuning data and the general task data set to obtain a multimodal emergency vertical domain model includes:
[0025] An initial multimodal emergency model is constructed based on a multimodal large language model framework, wherein a LoRA model fine-tuning module is introduced into the initial multimodal emergency model;
[0026] Set the fine-tuning training parameters of the LoRA model fine-tuning module;
[0027] Mixing the graphic-text instruction fine-tuning data with the general task data set to obtain an emergency graphic-text mixed data set;
[0028] The emergency graphic-text mixed data set is used to train the initial multimodal emergency model, and during the training process, the initial multimodal emergency model is fine-tuned by secondary instructions based on the fine-tuning training parameters to obtain a multimodal emergency vertical domain model.
[0029] Optionally, the multimodal emergency vertical domain model is used to perform retrieval-enhanced image analysis on the emergency-related query information input by the user, and generate emergency-related reply information, including:
[0030] Acquire emergency-related query information input by a user, wherein the emergency-related query information includes an image to be analyzed and text information to be answered;
[0031] Inputting the image to be analyzed into a pre-trained multimodal image annotation model for image annotation to obtain an image content overview corresponding to the image to be analyzed;
[0032] Retrieving emergency-related data from a vector database using the image content overview and the text information to be answered as retrieval key information;
[0033] Inputting the emergency-related data, the image to be analyzed, and the text information to be answered into the multimodal emergency vertical domain model;
[0034] The multimodal emergency vertical domain model is used to perform image recognition analysis on the image to be analyzed around the text information to be answered, and during the image recognition analysis process, the emergency-related data is used to perform context enhancement on the text information to be answered to generate emergency-related reply information.
[0035] Optionally, the process of collecting emergency-related graphic and text data includes:
[0036] Build a keyword lexicon related to emergency scenarios;
[0037] Obtaining an open-source image and text dataset on the Internet, wherein the open-source image and text dataset on the Internet includes a plurality of open-source images, and each of the open-source images corresponds to an image description text;
[0038] Perform keyword matching screening on each of the image description texts based on the keyword word library;
[0039] When the number of occurrences of the keyword in the image description text is greater than a preset number threshold, the image description text is determined to be a valid image description text, and the open source image corresponding to the valid image description text is a valid emergency image;
[0040] The valid image description text and the valid emergency image are used as a valid image-text pair;
[0041] All of the valid image-text pairs are integrated into emergency-related image-text data.
[0042] The present invention also provides a multi-modal emergency information generation device, comprising:
[0043] A multimodal data screening and expansion unit, used to obtain emergency-related graphic data and general task data sets, and perform multimodal data screening and expansion on the emergency-related graphic data to obtain graphic-text question-answer pair data;
[0044] A graphic-text instruction fine-tuning data construction unit, used to construct graphic-text instruction fine-tuning data based on thought chain knowledge distillation according to the graphic-text question-answer pair data;
[0045] A secondary instruction fine-tuning unit, used for performing secondary instruction fine-tuning on the pre-built initial multimodal emergency model according to the graphic instruction fine-tuning data and the general task data set to obtain a multimodal emergency vertical domain model;
[0046] The image recognition and analysis unit is used to perform retrieval-enhanced image recognition analysis on the emergency-related query information input by the user through the multimodal emergency vertical domain model, and generate emergency-related response information.
[0047] The present invention also provides an electronic device, the device comprising a processor and a memory:
[0048] The memory is used to store program code and transmit the program code to the processor;
[0049] The processor is used to execute the multi-modal based emergency information generation method as described in any one of the above items according to the instructions in the program code.
[0050] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to execute the multi-modal-based emergency information generation method as described in any one of the above items.
[0051] It can be seen from the above technical solutions that the present invention has the following advantages:
[0052] A method for generating emergency information based on multimodality is provided. First, emergency-related graphic data and general task data sets are obtained, and multimodal data screening and expansion are performed on the emergency-related graphic data to obtain graphic-text question-answer pair data; then, according to the graphic-text question-answer pair data, graphic-text instruction fine-tuning data based on thought chain knowledge distillation is constructed; then, according to the graphic-text instruction fine-tuning data and the general task data set, the pre-built initial multimodal emergency model is fine-tuned for the second time to obtain a multimodal emergency vertical domain model; finally, through the multimodal emergency vertical domain model, the emergency-related query information input by the user is analyzed based on retrieval enhancement, and emergency-related reply information is generated. By combining a large language model, multimodal technology, retrieval enhancement, and thought chain knowledge distillation, a multimodal emergency vertical domain model based on thought chain data enhancement with the ability to recognize images and distinguish things in emergency scenarios is constructed, and in practical applications, for the emergency-related query information input by the user, based on the multimodal emergency vertical domain model, accurate and real-time effective emergency-related reply information can be quickly generated. This can not only effectively solve the adaptability problem of traditional methods in complex emergency scenarios and significantly improve the efficiency of emergency information generation, but also enhance the accuracy, practicality and diversity of emergency information generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0054] Figure 1 is a flowchart of the steps of a multimodal-based emergency information generation method;
[0055] Figure 2 It is a schematic diagram of the overall process of a multi-modal-based emergency information generation method;
[0056] Figure 3 The structure block diagram of a multi-modal based emergency information generation device. DETAILED DESCRIPTION
[0057] The embodiments of the present invention provide a multimodal-based emergency information generation method, device, electronic device and medium, which are used to solve or partially solve the problem that the current emergency information generation method has poor applicability and pertinence, is limited in information update frequency and content accuracy, and is difficult to meet real-time and personalized needs.
[0058] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0059] As an example, with the rapid development of artificial intelligence, multimodal application scenarios have gradually become a research hotspot in the field of emergency information generation. Multimodal technology can perceive and understand complex scenes more comprehensively by integrating multiple data forms such as images and text. In emergency scenarios such as fire prevention and medical first aid, images are often the key information carriers. By identifying the specific environment and potential risk points in the picture, and combining text to generate professional guidance, users can be provided with more accurate and intuitive solutions. In actual situations, multimodal applications not only enhance the richness and fun of information expression, but also provide strong support for personalized and dynamic generation of emergency information.
[0060] However, the existing emergency information generation methods still have significant deficiencies in multimodal applications. For example, the currently commonly used fixed template method can only handle preset scenarios and has poor applicability. Another example is the manual writing method. Although it has a certain degree of flexibility, it lacks more efficient analysis capabilities and has poor pertinence. At the same time, traditional methods are still limited in the frequency of information updates and content accuracy, and it is difficult to meet the needs of real-time and personalization.
[0061] Through further analysis, the present invention found that the fixed template method can only process preset scenes, and cannot extract information from images, and is also difficult to cope with diverse data input in complex environments. Although manual writing has certain flexibility, it lacks efficient image analysis capabilities and cannot dynamically combine visual information to generate targeted emergency guidance.
[0062] These shortcomings make it impossible for traditional methods to fully tap the potential of multimodal technology in emergency information generation, limiting its application value in complex emergency scenarios. In this case, it is particularly necessary to introduce an artificial intelligence system for image recognition and event identification in emergency scenarios. If image recognition and event identification can be achieved in emergency scenarios, it means that by combining image and text information, complex data in actual scenarios can be more comprehensively perceived and analyzed, which can not only improve the accuracy of emergency information, but also significantly improve the practicality and response efficiency of emergency information by dynamically generating personalized solutions suitable for different user needs.
[0063] Therefore, one of the core inventive points of the embodiment of the present invention is: in view of the limitations of the traditional emergency information generation method in multimodal applications, a multimodal-based emergency information generation method is proposed. By combining a large language model, multimodal technology, retrieval enhancement and thought chain knowledge distillation, a multimodal emergency science popularization large language model based on thought chain data enhancement with the ability to recognize pictures and distinguish things in emergency scenarios is constructed, and in actual applications, based on the multimodal emergency science popularization large language model, accurate and real-time effective emergency-related response information can be quickly generated for the emergency-related query information input by the user. The technical solution provided by the present invention can effectively solve the adaptability problems of traditional methods in complex emergency scenarios, significantly improve the generation efficiency of emergency information, and enhance the accuracy, practicality and diversity of emergency information generation.
[0064] Reference Figure 1 , shows a flowchart of a method for generating emergency information based on multimodality provided by an embodiment of the present invention, which may specifically include the following steps:
[0065] Step 101, obtaining emergency-related graphic data and a general task data set, and performing multimodal data screening and expansion on the emergency-related graphic data to obtain graphic-text question-answer pair data;
[0066] This step mainly realizes the collection of emergency-related graphic data and general task data sets, and uses multimodal large language models and thinking chain technology to achieve data filtering (screening) and data expansion.
[0067] The general task dataset refers to the training data downloaded from the InternVL2 large language model official repository. The data source is the official Github repository. This part of the data is used to retain the general answering ability of the large language model to maintain the model performance.
[0068] For the collection of emergency-related graphic and text data, the embodiment of the present invention mainly adopts a keyword matching strategy, and filters out graphic and text pairs related to the current vertical domain (i.e., the emergency vertical domain) by analyzing the content of the image description.
[0069] In some embodiments, the collection process of emergency-related graphic and text data can be implemented by executing the following sub-steps S01 to S05:
[0070] Step S01: construct a keyword database related to emergency scenarios;
[0071] A keyword database related to emergency scenarios is established by using a multi-keyword matching method. Based on this keyword database, the number of times the keyword appears in the image description text can be counted during subsequent filtering and screening. If the number of keyword appearances is greater than a preset threshold, the current image-text pair is recorded as a valid image-text pair.
[0072] Step S02: obtaining an open-source image and text dataset from the Internet, where the open-source image and text dataset includes multiple open-source images, and each open-source image corresponds to an image description text;
[0073] The embodiment of the present invention mainly filters and screens emergency-related images from open-source image and text datasets on the Internet. In these datasets, each piece of data contains an image and its corresponding content description (i.e., image description text). Therefore, open-source image and text datasets on the Internet can be collected first. For example, the "Noah-Wukong" Chinese image and text dataset (a data set specially designed for multimodal learning and research).
[0074] Step S03: performing keyword matching screening on each image description text based on the keyword word library;
[0075] Then, a multi-keyword matching strategy can be used to selectively screen image and text data that are highly relevant to emergency situations from open source image and text datasets on the Internet.
[0076] Step S04: when the number of occurrences of the keyword in the image description text is greater than a preset number threshold, the image description text is determined to be a valid image description text, and the open source image corresponding to the valid image description text is a valid emergency image;
[0077] Step S05: taking the valid image description text and the valid emergency image as a valid image-text pair, and integrating all valid image-text pairs into emergency-related image-text data.
[0078] Through data integration, the emergency-related image-text data in the embodiment of the present invention may include a plurality of valid image-text pairs, each valid image-text pair including a valid image description text and a valid emergency image.
[0079] For example, take fire-related image and text data as an example. Select the keywords ["fire", "flammable", "explosive", "fire scene", "fire", "fire fighting", "fire extinguishing", "fire prevention", "burning"], and obtain fire-related image and text pairs through keyword matching and screening. The image types can include images of flammable and explosive materials, images of fire hazards, images of fire scenes, images of fire safety education, etc.
[0080] After obtaining emergency-related graphic and text data, we can then design a variety of refined multimodal emergency tasks for users, and use a multimodal large language model to classify images and divide tasks.
[0081] In some embodiments, the process of performing multimodal data screening and expansion on emergency-related graphic data to obtain graphic-text question-answer pair data can be achieved by executing the following sub-steps S11 to S13:
[0082] Step S11: Integrate a variety of refined multi-modal emergency tasks, and construct instruction text based on task division instruction requirements, task division category enumeration, and specific examples;
[0083] The instruction text designed in the embodiment of the present invention conforms to the format of "instruction requirements for task division + category enumeration of task division + specific examples". Still taking the fire-related picture and text pair as an example, the designed instruction text example is: "Which of the following points is more consistent with the content of this picture? The options are as follows: flammable material analysis and fire hazard analysis; fire accident in progress; fire accident ruins; fire safety education; completely unrelated to fire. Please only answer the content of the option. Answer example: fire accident ruins."
[0084] Step S12: Based on the instruction text, all valid emergency images are divided into tasks through the pre-trained multimodal task division model, emergency-irrelevant images are eliminated based on the task division results, and all retained valid emergency images are annotated to obtain corresponding target emergency images;
[0085] In this step, the pre-trained multimodal large language model InternVL2-76B is mainly used as a multimodal task division model to screen effective emergency images and complete task division. Specifically, the instruction text and each effective emergency image are input into the multimodal task division model InternVL2-76B, and the multimodal task division model InternVL2-76B can be used to divide the input images and screen out emergency-irrelevant images.
[0086] Still taking the fire-related image-text pair as an example, the task classification categories can mainly include: "Analysis of flammable materials and fire hazard analysis; fire accident in progress; fire accident ruins; fire safety education; completely unrelated to fire", a total of 5 categories. Among them, images classified as "completely unrelated to fire" will be regarded as invalid images (i.e., emergency-irrelevant images) and will be removed. All the retained valid emergency images are annotated separately to achieve data expansion for image data and obtain the corresponding target emergency images.
[0087] Step S13: The image description text corresponding to the target emergency image is used as the target image description text, the target emergency image and the target image description text are used as an image-text question-answer pair, and all image-text question-answer pairs are integrated into image-text question-answer pair data.
[0088] Step 102, constructing image-text instruction fine-tuning data based on thought chain knowledge distillation according to the image-text question-answer pair data;
[0089] This step mainly uses the collected picture-text question-answer data to build picture-text instruction fine-tuning data for training the multimodal emergency vertical domain large model.
[0090] In order to construct the fine-tuning data of graphic instructions, the embodiment of the present invention mainly combines the multimodal large language model with the thinking chain technology to expand and generate emergency-related graphic data to achieve knowledge enhancement of the large language model.
[0091] In some embodiments, the process of constructing the image-text instruction fine-tuning data based on the thought chain knowledge distillation according to the image-text question-answer pair data can be implemented by executing the following sub-steps S21 to S22:
[0092] Step S21: constructing an initial image-text description model based on the multimodal large language model framework;
[0093] In the embodiment of the present invention, the multimodal large language model InternVL2-76B is selected as a framework to construct an initial graphic description model.
[0094] Step S22: Using the picture-text question-answer pair data, the initial picture-text description model is trained for picture-text extension generation based on thought chain knowledge distillation, and picture-text instruction fine-tuning data is generated based on the trained picture-text description model.
[0095] Step S22 mainly realizes the generation of emergency vertical domain image-text pair data (i.e., image-text instruction fine-tuning data) based on the image content and image description text of each image in the image-text question-answer data. The extended generation of image-text data can be mainly divided into three steps: first, the image content is described by the initial image-text description model; then the initial image-text description model performs text content comparison and corrects the image description content; finally, the trained image-text description model generates emergency vertical domain image-text pair data related to the specified task based on the corrected image description content.
[0096] Specifically, the process of using the picture-text question-answer pair data to perform picture-text extension generation training based on thought chain knowledge distillation on the initial picture-text description model, and generating picture-text instruction fine-tuning data based on the trained picture-text description model can be achieved by executing the following sub-steps S221 to S224:
[0097] Step S221: for each image-text question-answer pair, based on the task division result corresponding to the target emergency image, the target emergency image is described by the initial image-text description model to generate an image description model output text;
[0098] The target emergency image is described by the initial image-text description model. The specific task prompt (prompt word or instruction) is: "Please describe the content of the image in detail. The image is related to {TASK}." Among them, TASK refers to the task division result of the previous step S12, such as "fire safety education". Then the output of the image-text description model is recorded as the description content of the image by the image-text description model. That is, the image description model outputs text.
[0099] Step S222: performing a content comparison between the text output by the image description model and the target image description text, and based on the content comparison difference, performing text correction on the text output by the image description model to obtain a corrected image description text;
[0100] Compare the output text of the image description model with the target image description text, and correct the image description content output by the model. The specific task prompt of this process is: "The content of this image is related to {CAPTION}. Your previous description of this image is {MLLM_CAPTION}. Please further analyze and correct your description." Among them, CAPTION refers to the original description text of the image (that is, the target image description text), and MLLM_CAPTION refers to the description content of the image by the image description model. Then record the output of the image-text description model, which is recorded as the corrected description content of the image by the image-text description model. That is, the corrected image description text.
[0101] Step S223: when the image description and text correction of all the picture-text question-answer pairs are completed, a trained picture-text description model is obtained;
[0102] Step S224: Based on the trained image-text description model, generate emergency vertical domain image-text pairs corresponding to each image description correction text, and integrate all emergency vertical domain image-text pairs into image-text instruction fine-tuning data.
[0103] Step S224 mainly divides pairs according to different tasks and designs different task prompts so that the trained graphic description model can generate more refined emergency vertical domain question and answer pairs.
[0104] For example, assuming that the task division result of the target emergency image is "combustible material analysis and fire hazard analysis", the corresponding task prompt can be: "Please generate a question-answer pair related to combustible material analysis and fire hazard analysis based on this picture and the content related to {MLLM_REFINE_CAPTION}. The length of the answer should be greater than 200 words. The format is similar to the following example: {Question: \"What fire hazards are there in the picture?\", Answer: \"The battery and the stove are placed too close. \"
[0105] For another example, assuming that the task division result of the target emergency image is "fire safety education", its task prompt can be: "Please generate a question-answer pair related to fire safety education based on this picture, the content is related to {MLLM_REFINE_CAPTION}, and the length of the answer should be greater than 200 words. The format is similar to the following example: {Question: \"When encountering a fire at home, what basic escape steps should be followed to ensure your own safety?\", Answer: \"When encountering a fire at home, it is crucial to stay calm and quickly take the correct escape steps. Once a fire is discovered, call the local fire alarm number (such as 119 in China) immediately and clearly inform the specific address of the fire, the size of the fire, and whether there are people trapped. "
[0106] Among them, MLLM_REFINE_CAPTION refers to the image description correction text. Finally, all emergency vertical domain image-text pairs generated around image content and task category, that is, image-text instruction fine-tuning data, can be obtained from the output of the trained image-text description model InternVL2-76B.
[0107] Step 103, performing secondary instruction fine-tuning on the pre-built initial multimodal emergency model according to the graphic instruction fine-tuning data and the general task data set to obtain a multimodal emergency vertical domain model;
[0108] In some embodiments, the process of performing secondary instruction fine-tuning on the pre-built initial multimodal emergency model according to the graphic instruction fine-tuning data and the general task data set to obtain the multimodal emergency vertical domain model can be achieved by executing the following sub-steps S31 to S34:
[0109] Step S31: constructing an initial multimodal emergency model based on the multimodal large language model framework, and introducing a LoRA model fine-tuning module into the initial multimodal emergency model;
[0110] An initial multimodal emergency model is constructed based on the multimodal large language model InternVL2-8B, and the LoRA (Low-Rank Adaptation) model fine-tuning module is introduced into the initial multimodal emergency model.
[0111] That is to say, the model that needs fine-tuning in the embodiment of the present invention is the initial multimodal emergency model InternVL2-8B. Specifically, the initial multimodal emergency model can be fine-tuned by combining the graphic instruction fine-tuning data with the general task data set.
[0112] It should be pointed out that the initial multimodal emergency model can also be constructed using other sizes of multimodal large language models InternVL2. For example, InternVL2-4B is selected. The purpose of selecting InternVL2-8B in the present invention is that compared with models of other sizes, the 8B model has a higher cost-effectiveness in reasoning, can be reasoned on a single consumer-grade graphics card, and the reasoning effect is better than that of the largest model (InternVL2-76B), and the loss is also smaller. It is understandable that the present invention is not limited to this.
[0113] In the specific training, the present invention introduces the LoRA module to the initial multimodal emergency model. The LoRA module is a technology for fine-tuning large pre-trained models. Unlike traditional fine-tuning methods, the LoRA module achieves efficient adaptive adjustment by only updating a small number of parameters in the model. By introducing the LoRA module, the computing resources and time costs can be significantly reduced while maintaining the competitiveness of the model performance.
[0114] Step S32: setting the fine-tuning training parameters of the LoRA model fine-tuning module;
[0115] According to the recommended parameters of the official training document of InternVL2, set the fine-tuning training parameters of the LoRA model fine-tuning module.
[0116] Step S33: Mixing the graphic and text instruction fine-tuning data with the general task data set to obtain an emergency graphic and text mixed data set;
[0117] In order to maintain the general image-text question-answering capability of InternVL2, during the model training process, it is necessary to mix the image-text instruction fine-tuning data (which can be understood as the vertical domain task data set) with the general task data set and then use it in fine-tuning training.
[0118] Specifically, first retain all the fine-tuning data of InternVL2 (i.e., the fine-tuning data of graphic instructions), modify the general task data set into the training data format of InternVL2, mix the two types of data, and finally organize the obtained new data into the Jsonl file of the original data for subsequent fine-tuning training.
[0119] Step S34: The initial multimodal emergency model is trained using the emergency image-text mixed data set, and during the training process, the initial multimodal emergency model is fine-tuned by secondary instructions based on fine-tuning training parameters to obtain a multimodal emergency vertical domain model.
[0120] Therefore, in the knowledge enhancement stage, the embodiment of the present invention combines the multimodal large language model with the thought chain knowledge enhancement technology to achieve invalid information filtering and thought chain image annotation expansion of emergency-related graphic data. The data obtained in the knowledge enhancement stage will be used for knowledge injection for fine-tuning training of the multimodal large language model. As a result, the fine-tuned multimodal large language model can be combined with the retrieval enhancement generation technology to build a more reliable multimodal emergency knowledge popularization platform.
[0121] Step 104: Using the multimodal emergency vertical domain model, a search-enhanced graph analysis is performed on the emergency-related query information input by the user, and emergency-related response information is generated.
[0122] In practical applications, the present invention introduces a small-sized multimodal large language model InternVL2-1B for image annotation to summarize the content of the image input by the user. Therefore, in the subsequent processing process, the image annotation and the question input by the user can be used as retrieval queries (Retrieval Query) to retrieve emergency-related data from the vector database (Vector Database). At the same time, the retrieved key can be stored and used as traceability output (that is, which information is related to the current response of the output model). Among them, the stored key can be used as a quick search keyword when traceability is required later. The emergency-related data can then be used as context enhancement for the user-input question to achieve retrieval enhancement, assist the model in conducting questions and answers in the emergency field, and generate emergency-related response information, thereby improving the relevance of the answers of the multimodal large language model.
[0123] In combination with the above discussion, in some embodiments, the process of performing retrieval-enhanced graph analysis on the emergency-related query information input by the user and generating emergency-related reply information through the multimodal emergency vertical domain model can be achieved by executing the following sub-steps S41 to S45:
[0124] Step S41: Acquire emergency-related query information input by a user, where the emergency-related query information includes images to be analyzed and text information to be answered;
[0125] Step S42: inputting the image to be analyzed into the pre-trained multimodal image annotation model for image annotation to obtain an image content overview corresponding to the image to be analyzed;
[0126] Step S43: using the image content overview and the text information to be answered as the key information for retrieval, and retrieving emergency-related data from the vector database;
[0127] Step S44: inputting emergency-related data, images to be analyzed, and text information to be answered into the multimodal emergency vertical domain model;
[0128] Step S45: Perform image recognition analysis around the text information to be answered on the image to be analyzed through the multimodal emergency vertical domain model, and use emergency-related data to perform context enhancement on the text information to be answered during the image recognition analysis process to generate emergency-related reply information.
[0129] In the content generation stage, through comprehensive analysis of user input images and text information, combined with retrieval enhancement generation technology, relevant professional knowledge can be quickly obtained to generate emergency-related content with clear logic, rich pictures and texts, and easy to understand, and a reliable basis for generation can be provided. Among them, the generated results include but are not limited to knowledge questions and answers in emergency scenarios, targeted processing steps, etc. All outputs are presented in an intuitive and concise manner to help users quickly grasp key information and take effective actions. Therefore, by combining image understanding and language generation technology, a deep understanding of emergency scenarios and accurate knowledge generation can be achieved. It can provide efficient and intelligent solutions for complex and changing emergency needs, and is widely applicable to various emergency scenarios such as firefighting, building accidents, and medical treatment.
[0130] From the perspective of the overall technical solution, the present invention designs two key personalized multimodal tasks. One is the identification of hidden dangers in emergency scenarios (such as fire hidden dangers), and the other is the popularization of multimodal emergency knowledge. The former is based on image understanding and natural language reasoning, and can provide users with accurate hidden danger detection services. The latter, through optimized information presentation, enables users to obtain immediate, relevant and easy-to-understand help, thereby improving the overall efficiency of emergency response.
[0131] Taking fire-related emergency scenarios as an example, the technical solution provided by the embodiment of the present invention is adopted, through the advanced large language model combined with pixel-level image understanding technology, to support users to input image queries and intelligently identify potential fire risk points in the image. For example, improper stacking of flammable materials, aging of electrical lines, etc. At the same time, the model can also generate professional prevention suggestions based on the recognition results to help non-professionals to self-check and self-correct, so that fire prevention extends from professional fields to daily life and becomes an important part of every family. In addition, by popularizing fire knowledge in the form of pictures and texts, emergency information can be transformed into more intuitive, interesting and professionally detailed learning content. By integrating text descriptions with multimedia elements, not only can the public easily learn knowledge related to fire prevention, but also improve memory effect and practicality, thereby comprehensively improving the public's emergency response capabilities and self-protection awareness.
[0132] In an embodiment of the present invention, a method for generating emergency information based on multimodality is proposed in view of the limitations of traditional emergency information generation methods in multimodal applications. By combining a large language model, multimodal technology, retrieval enhancement, and thought chain knowledge distillation, a multimodal emergency science popularization large language model based on thought chain data enhancement with the ability to recognize images and distinguish things in emergency scenarios is constructed. In actual applications, based on the multimodal emergency science popularization large language model, accurate, real-time and effective emergency-related response information can be quickly generated for emergency-related query information input by users. By adopting the technical solution provided by the present invention, the adaptability problem of traditional methods in complex emergency scenarios can be effectively solved, the generation efficiency of emergency information can be significantly improved, and the accuracy, practicality and diversity of emergency information generation can be enhanced.
[0133] For better explanation, refer to Figure 2 , taking fire-related graphic data as an example, a schematic diagram of the overall process of a multi-modal emergency information generation method provided by an embodiment of the present invention is shown. It should be pointed out that this embodiment only briefly describes the general process of multi-modal emergency information generation. The specific implementation process of each step can be understood by referring to the relevant content in the aforementioned embodiment. It will not be repeated here. It can be understood that the present invention does not limit this.
[0134] S1: Collect emergency-related graphic data for fine-tuning
[0135] First, we obtain an open-source image and text dataset from the Internet, and perform frequency screening on the open-source image and text dataset based on text keywords ["fire", "flammable", "explosive", "fire scene", "fire", "fire fighting", "fire extinguishing", "fire prevention", "burning"] to obtain highly relevant image and text data for emergency situations.
[0136] Then, based on the pre-trained multimodal task division model InternVL2-76B, multimodal data screening and expansion (image classification and task division) were performed on the obtained emergency highly relevant image and text data to obtain image-text question and answer pair data.
[0137] S2: Construct fine-tuning data of graphic instructions based on thought chain knowledge distillation and train multi-modal vertical domain large model
[0138] The multimodal large language model InternVL2-76B is selected as the framework to build an initial picture and text description model;
[0139] Using the picture-text question-answer pair data, the initial picture-text description model InternVL2-76B is trained for picture-text extension generation based on thought chain knowledge distillation (image description and correction description content), and emergency-related question-answer pairs (emergency vertical domain picture-text pairs) are generated based on the trained picture-text description model to obtain the vertical domain task data set (i.e., picture-text instruction fine-tuning data);
[0140] The pre-acquired general task dataset and the constructed vertical domain task dataset are mixed, and the mixed dataset is used to train the initial multimodal emergency model InternVL2-8B which introduces the LoRA model fine-tuning module. During the training process, the initial multimodal emergency model InternVL2-8B is fine-tuned for the second time based on the fine-tuning training parameters to obtain the multimodal emergency vertical domain model InternVL2-8B.
[0141] S3: Combine retrieval-enhanced generation technology with multimodal large language models to build an emergency information question-answering system
[0142] Obtaining emergency-related query information input by a user, the emergency-related query information including images to be analyzed and text information to be answered;
[0143] Input the image to be analyzed into the pre-trained multimodal image annotation model InternVL2-1B for image annotation to obtain an overview of the image content corresponding to the image to be analyzed;
[0144] The image content overview and the text information to be answered are used as the key information for retrieval, and the emergency-related data are retrieved from the vector database;
[0145] Input emergency-related data, images to be analyzed, and text information to be answered into the multimodal emergency vertical domain model InternVL2-8B;
[0146] The multimodal emergency vertical domain model InternVL2-8B is used to perform image recognition analysis on the image to be analyzed around the text information to be answered. During the image recognition analysis process, emergency-related data are used to enhance the context of the text information to be answered and generate emergency-related response information.
[0147] Reference Figure 3 , shows a structural block diagram of a multi-modal emergency information generation device provided by an embodiment of the present invention, which may specifically include:
[0148] The multimodal data screening and expansion unit 301 is used to obtain emergency-related graphic data and a general task data set, and perform multimodal data screening and expansion on the emergency-related graphic data to obtain graphic-text question-answer pair data;
[0149] A graphic-text instruction fine-tuning data construction unit 302, configured to construct graphic-text instruction fine-tuning data based on thought chain knowledge distillation according to the graphic-text question-answer pair data;
[0150] A secondary instruction fine-tuning unit 303 is used to perform secondary instruction fine-tuning on the pre-built initial multimodal emergency model according to the graphic instruction fine-tuning data and the general task data set to obtain a multimodal emergency vertical domain model;
[0151] The image recognition and analysis unit 304 is used to perform retrieval-enhanced image recognition analysis on the emergency-related query information input by the user through the multimodal emergency vertical domain model, and generate emergency-related response information.
[0152] In an optional embodiment, the emergency-related graphic data includes a plurality of valid graphic pairs, each of which includes a valid image description text and a valid emergency image; the multimodal data screening and expansion unit 301 includes:
[0153] The instruction text construction unit is used to integrate a variety of refined multi-modal emergency tasks and construct instruction texts based on task division instruction requirements, task division category enumeration, and specific examples;
[0154] A target emergency image generation unit is used to perform task division on all the valid emergency images based on the instruction text through a pre-trained multimodal task division model, eliminate emergency-irrelevant images based on the task division results, and perform image annotation on all the retained valid emergency images to obtain corresponding target emergency images;
[0155] A picture-text question-answer pair determination unit, configured to use the image description text corresponding to the target emergency image as the target image description text, and use the target emergency image and the target image description text as a picture-text question-answer pair;
[0156] The picture-text question-answer pair data integration unit is used to integrate all the picture-text question-answer pairs into picture-text question-answer pair data.
[0157] In an optional embodiment, the graphic instruction fine-tuning data construction unit 302 includes:
[0158] An initial image-text description model building unit, used to build an initial image-text description model based on a multimodal large language model framework;
[0159] The image-text instruction fine-tuning data generation unit is used to use the image-text question-answer pair data to perform image-text extension generation training based on thought chain knowledge distillation on the initial image-text description model, and generate image-text instruction fine-tuning data based on the trained image-text description model.
[0160] In an optional embodiment, the graphic instruction fine-tuning data generating unit includes:
[0161] An image description unit is used to describe the target emergency image for each of the image-text question-answer pairs based on the task division result corresponding to the target emergency image by using the initial image-text description model to generate an image description model output text;
[0162] A text correction unit, used for comparing the content of the text output by the image description model with the target image description text, and based on the content comparison difference, performing text correction on the text output by the image description model to obtain an image description correction text;
[0163] A trained image-text description model output unit, used to obtain a trained image-text description model when the image description and text correction of all the image-text question-answer pairs are completed;
[0164] The emergency vertical domain image-text pair generation unit is used to generate an emergency vertical domain image-text pair corresponding to each image description correction text based on the trained image-text description model, and integrate all the emergency vertical domain image-text pairs into image-text instruction fine-tuning data.
[0165] In an optional embodiment, the secondary instruction fine-tuning unit 303 includes:
[0166] An initial multimodal emergency model construction unit is used to construct an initial multimodal emergency model based on a multimodal large language model framework, wherein a LoRA model fine-tuning module is introduced into the initial multimodal emergency model;
[0167] A fine-tuning training parameter setting unit, used to set the fine-tuning training parameters of the LoRA model fine-tuning module;
[0168] A data mixing unit, used for mixing the graphic-text instruction fine-tuning data with the general task data set to obtain an emergency graphic-text mixed data set;
[0169] The model training and fine-tuning unit is used to train the initial multimodal emergency model using the emergency graphic and text mixed data set, and during the training process, perform secondary instruction fine-tuning on the initial multimodal emergency model based on the fine-tuning training parameters to obtain a multimodal emergency vertical domain model.
[0170] In an optional embodiment, the image recognition and analysis unit 304 includes:
[0171] An emergency-related query information acquisition unit, used to acquire emergency-related query information input by a user, wherein the emergency-related query information includes an image to be analyzed and text information to be answered;
[0172] An image content summary generating unit, used for inputting the image to be analyzed into a pre-trained multimodal image annotation model for image annotation, and obtaining an image content summary corresponding to the image to be analyzed;
[0173] An emergency-related data retrieval unit, configured to retrieve emergency-related data from a vector database by using the image content overview and the text information to be answered as retrieval key information;
[0174] A data input unit, used for inputting the emergency-related data, the image to be analyzed and the text information to be answered into the multimodal emergency vertical domain model;
[0175] The image recognition and analysis subunit is used to perform image recognition analysis on the image to be analyzed around the text information to be answered through the multimodal emergency vertical domain model, and during the image recognition and analysis process, use the emergency-related data to contextually enhance the text information to be answered to generate emergency-related reply information.
[0176] In an optional embodiment, the emergency information generating device further includes:
[0177] A keyword word library construction unit, used to construct a keyword word library related to emergency scenarios;
[0178] A network open source image and text data set acquisition unit, used to acquire a network open source image and text data set, wherein the network open source image and text data set includes a plurality of open source images, and each of the open source images corresponds to an image description text;
[0179] A keyword matching and screening unit, configured to perform keyword matching and screening on each of the image description texts based on the keyword vocabulary;
[0180] A valid image description text determining unit, configured to determine that the image description text is a valid image description text when the number of occurrences of a keyword in the image description text is greater than a preset number threshold, and that the open source image corresponding to the valid image description text is a valid emergency image;
[0181] A valid image-text pair determination unit, used for taking the valid image description text and the valid emergency image as a valid image-text pair;
[0182] The emergency-related image-text data integration unit is used to integrate all the valid image-text pairs into emergency-related image-text data.
[0183] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the aforementioned method embodiment.
[0184] An embodiment of the present invention further provides an electronic device, the device comprising a processor and a memory:
[0185] The memory is used to store the program code and transmit the program code to the processor;
[0186] The processor is used to execute the multi-modal emergency information generation method of any embodiment of the present invention according to the instructions in the program code.
[0187] An embodiment of the present invention further provides a computer-readable storage medium, which is used to store program code, and the program code is used to execute the multi-modal emergency information generation method of any embodiment of the present invention.
[0188] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0189] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0190] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0191] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0192] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0193] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-modal emergency information generation method, characterized in that: include: Acquire emergency-related graphic and text data and general task data sets, and perform multimodal data screening and expansion on the emergency-related graphic and text data to obtain graphic and text question-answer pair data; According to the picture-text question-answer pair data, construct picture-text instruction fine-tuning data based on thought chain knowledge distillation; Performing secondary instruction fine-tuning on the pre-built initial multimodal emergency model according to the graphic instruction fine-tuning data and the general task data set to obtain a multimodal emergency vertical domain model; Through the multimodal emergency vertical domain model, the emergency-related query information input by the user is subjected to retrieval-enhanced graph analysis, and emergency-related response information is generated.
2. The multimodal emergency information generation method according to claim 1, characterized in that: The emergency-related image-text data includes a plurality of valid image-text pairs, each of which includes a valid image description text and a valid emergency image; the multimodal data screening and expansion of the emergency-related image-text data to obtain image-text question-answer pair data includes: Integrate a variety of refined multi-modal emergency tasks, and construct instruction texts based on task division instruction requirements, task division category enumeration, and specific examples; Based on the instruction text, all the valid emergency images are divided into tasks through a pre-trained multimodal task division model, emergency irrelevant images are eliminated based on the task division results, and all the retained valid emergency images are respectively annotated to obtain corresponding target emergency images; The image description text corresponding to the target emergency image is used as the target image description text, and the target emergency image and the target image description text are used as an image-text question-answer pair; All of the picture-text question-answer pairs are integrated into picture-text question-answer pair data.
3. The multimodal emergency information generation method according to claim 2, characterized in that: The step of constructing the image-text instruction fine-tuning data based on the thought chain knowledge distillation according to the image-text question-answer pair data includes: Build an initial picture and text description model based on the multimodal large language model framework; The image-text question-answer pair data is used to perform image-text extension generation training based on thought chain knowledge distillation on the initial image-text description model, and image-text instruction fine-tuning data is generated based on the trained image-text description model.
4. The multimodal emergency information generation method according to claim 3, characterized in that: The method adopts the picture-text question-answer pair data to perform picture-text extension generation training based on thought chain knowledge distillation on the initial picture-text description model, and generates picture-text instruction fine-tuning data based on the trained picture-text description model, including: For each of the image-text question-answer pairs, based on the task division result corresponding to the target emergency image, the target emergency image is described by the initial image-text description model to generate an image description model output text; Comparing the content of the image description model output text with the target image description text, and based on the content comparison difference, performing text correction on the image description model output text to obtain an image description correction text; When the image description and text correction of all the image-text question-answer pairs are completed, a trained image-text description model is obtained; Based on the trained image-text description model, emergency vertical domain image-text pairs corresponding to each image description correction text are generated respectively, and all the emergency vertical domain image-text pairs are integrated into image-text instruction fine-tuning data.
5. The multi-modal emergency information generation method according to claim 3, characterized in that: The pre-built initial multimodal emergency model is subjected to secondary instruction fine-tuning according to the graphic instruction fine-tuning data and the general task data set to obtain a multimodal emergency vertical domain model, including: An initial multimodal emergency model is constructed based on a multimodal large language model framework, wherein a LoRA model fine-tuning module is introduced into the initial multimodal emergency model; Set the fine-tuning training parameters of the LoRA model fine-tuning module; Mixing the graphic-text instruction fine-tuning data with the general task data set to obtain an emergency graphic-text mixed data set; The emergency graphic-text mixed data set is used to train the initial multimodal emergency model, and during the training process, the initial multimodal emergency model is fine-tuned by secondary instructions based on the fine-tuning training parameters to obtain a multimodal emergency vertical domain model.
6. The multimodal emergency information generation method according to claim 1, characterized in that: The multimodal emergency vertical domain model is used to perform retrieval-enhanced image analysis on the emergency-related query information input by the user, and generate emergency-related response information, including: Acquire emergency-related query information input by a user, wherein the emergency-related query information includes an image to be analyzed and text information to be answered; Inputting the image to be analyzed into a pre-trained multimodal image annotation model for image annotation to obtain an image content overview corresponding to the image to be analyzed; Retrieving emergency-related data from a vector database using the image content overview and the text information to be answered as retrieval key information; Inputting the emergency-related data, the image to be analyzed, and the text information to be answered into the multimodal emergency vertical domain model; The multimodal emergency vertical domain model is used to perform image recognition analysis on the image to be analyzed around the text information to be answered, and during the image recognition analysis process, the emergency-related data is used to perform context enhancement on the text information to be answered to generate emergency-related reply information.
7. The multimodal emergency information generation method according to any one of claims 1 to 6, characterized in that: The process of collecting emergency-related graphic and text data includes: Build a keyword lexicon related to emergency scenarios; Obtaining an open-source image and text dataset on the Internet, wherein the open-source image and text dataset on the Internet includes a plurality of open-source images, and each of the open-source images corresponds to an image description text; Perform keyword matching screening on each of the image description texts based on the keyword word library; When the number of occurrences of the keyword in the image description text is greater than a preset number threshold, the image description text is determined to be a valid image description text, and the open source image corresponding to the valid image description text is a valid emergency image; The valid image description text and the valid emergency image are used as a valid image-text pair; All of the valid image-text pairs are integrated into emergency-related image-text data.
8. A multi-modal emergency information generation device, characterized in that: include: A multimodal data screening and expansion unit, used to obtain emergency-related graphic data and general task data sets, and perform multimodal data screening and expansion on the emergency-related graphic data to obtain graphic-text question-answer pair data; A graphic-text instruction fine-tuning data construction unit, used to construct graphic-text instruction fine-tuning data based on thought chain knowledge distillation according to the graphic-text question-answer pair data; A secondary instruction fine-tuning unit, used for performing secondary instruction fine-tuning on the pre-built initial multimodal emergency model according to the graphic instruction fine-tuning data and the general task data set to obtain a multimodal emergency vertical domain model; The image recognition and analysis unit is used to perform retrieval-enhanced image recognition analysis on the emergency-related query information input by the user through the multimodal emergency vertical domain model, and generate emergency-related response information.
9. An electronic device, characterized in that: The device comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the multimodal-based emergency information generation method described in any one of claims 1-7 according to the instructions in the program code.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program code, and the program code is used to execute the multi-modal based emergency information generation method described in any one of claims 1-7.
Citation Information
Patent Citations
Chinese and cross-language query expansion method based on retrieval enhancement and knowledge distillation
CN118673095A
Label calculation method based on hybrid expert LLM
CN118760736A
Enhanced search result generation using multi-document summarization
US20240281487A1
Task processing method and task processing system
WO2025007892A1