A method and system for extracting and summarizing multimodal key information based on a large language model
By using a multimodal key information extraction and summarization system based on a large language model, combining preprocessing and traditional algorithms, key information is identified and extracted. The system monitors performance in real time and provides alerts, solving the efficiency problem of large language models in text recognition and information extraction, and improving the efficiency and accuracy of information extraction.
Patent Information
- Application Number
- CN202411646774.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-11-18
AI Technical Summary
In existing technologies, when large language models are combined with text recognition and information extraction techniques, it is difficult to efficiently extract document information, especially when dealing with unstructured text and irregularly formatted documents, resulting in performance degradation.
A multimodal key information extraction and summarization system based on a large language model is adopted, including preprocessing, key information extraction, data acquisition, parameter setting, core processing and display and reminder units. By combining the large language model with traditional algorithms, it identifies named entities and event information, obtains the importance index and performance evaluation index of key information, and monitors and alerts in real time.
It achieves efficient extraction of key information, improves information extraction efficiency, and enhances system performance through real-time monitoring and alerts, enabling timely detection and handling of performance issues, thereby improving the system's processing capacity and the accuracy of information extraction.
Smart Images

Figure CN119597898B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data processing technology, and specifically relates to a method and system for extracting and summarizing multimodal key information based on a large language model. Background Technology
[0002] Key information extraction is a crucial task in Natural Language Processing (NLP), aiming to extract useful information from unstructured text and transform it into structured data for further analysis and application. The development of information extraction techniques aims to address the information overload problem in large-scale text data processing, providing powerful analytical tools for enterprises, research institutions, and government departments. Traditional information extraction methods typically rely on rule-based and key information matching, which is limited by the text structure and the presence of specific key information. This means that the performance of traditional information extraction methods degrades when the document format is irregular or the user's question is not explicitly key information. With the development of large language models, they have achieved significant success in NLP tasks. However, challenges remain in their application to information extraction, particularly in effectively combining large language models with text recognition and information extraction techniques to achieve efficient document information extraction. Summary of the Invention
[0003] To address the shortcomings of existing technologies, this invention provides a method and system for extracting and summarizing multimodal key information based on a large language model, thereby solving the technical problem of combining large language models with text recognition and information extraction technologies to achieve efficient document information extraction.
[0004] To solve the above-mentioned technical problems, the present invention adopts the following technical solution.
[0005] This invention first discloses a multimodal key information extraction and summarization system based on a large language model, including: a preprocessing unit, a key information extraction unit, a data acquisition unit, a data transmission unit, a parameter setting unit, a core processing unit, a summary generation unit, and a display and reminder unit;
[0006] The preprocessing unit is used to perform format unification, noise reduction, and cleaning on the data from which key information needs to be extracted, and to perform operations such as word segmentation, part-of-speech tagging, removal of stop words, removal of special characters, removal of punctuation marks, and removal of HTML tags on the data.
[0007] The key information extraction unit is used to identify and extract named entities and event information through a large language model, identify the semantic relationships between entities, and extract the attribute information of entities and events.
[0008] The data acquisition unit is used to collect the number of keywords, the number of characters, the duration, and the time data required by the multimodal key information extraction and summarization system based on a large language model.
[0009] The data transmission unit is used to transmit the data collected by the data acquisition unit to the core processing unit.
[0010] The parameters set by the parameter setting unit are transmitted to the core processing unit and the display reminder unit. At the same time, the report generated by the core processing unit is transmitted to the summary generation unit and the display reminder unit.
[0011] The parameter setting unit is used to set various parameters required for analysis of the multimodal key information extraction and summarization system based on a large language model;
[0012] The core processing unit is used to receive the data collected by the data acquisition unit, analyze and process the received data, and generate a report on the multimodal key information extraction and summary system based on the analysis and processing results.
[0013] The summary generation unit is used to summarize the extracted multimodal key information based on the report on the multimodal key information extraction and summary system based on the large language model transmitted by the core processing unit.
[0014] The display and reminder unit is used to acquire the report on the multimodal key information extraction and summary system based on the large language model generated by the core processing unit, display the acquired report, and issue a warning when the report results reach a preset reminder threshold.
[0015] The present invention further includes the following preferred embodiments:
[0016] The data acquisition unit further includes a keyword quantity acquisition unit, a text quantity acquisition unit, a duration acquisition unit, and a time acquisition unit. The keyword quantity acquisition unit is used to acquire various data related to the number of keywords required by the multimodal key information extraction and summarization system based on a large language model. The text quantity acquisition unit is used to acquire various data related to the number of characters required by the multimodal key information extraction and summarization system based on a large language model. The duration acquisition module is used to acquire duration data required by the multimodal key information extraction and summarization system based on a large language model. The time acquisition module is used to acquire various time data required by the multimodal key information extraction and summarization system based on a large language model.
[0017] The parameter setting unit includes a modality weight setting module and a reminder threshold setting module; the modality weight setting module is used to set the weight coefficient of each modality of the key information extracted by the multimodal key information extraction and summarization system based on the large language model; the reminder threshold setting module is used to set the reminder threshold required by the multimodal key information extraction and summarization system based on the large language model.
[0018] The core processing unit includes a data receiving module, a data analysis module, and a report generation module. The data receiving module receives all data required by the multimodal key information extraction and summarization system based on a large language model, collected by the data acquisition unit, and the weight coefficients of each modality of the key information extracted by the multimodal key information extraction and summarization system based on a large language model, set by the parameter setting unit. The data analysis module analyzes and processes the data required by the multimodal key information extraction and summarization system based on a large language model, received by the data receiving module. The report generation module generates a report on the multimodal key information extraction and summarization system based on a large language model, based on the analysis and processing results of the data analysis module.
[0019] Obtain the number of each key piece of information in each modality, the total number of key pieces of information in each modality, the number of words in the first occurrence of each key piece of information in the data, and the number of words in the last occurrence. Based on the above data analysis and processing, obtain the importance index ω of each key piece of information.
[0020]
[0021] Where, N a The total amount of information for each modality, n a The number of key pieces of information in each modality, I a The weight coefficient for each modality is the key information, where x is the number of modalities in the data, and n is the number of modalities. f n is the number of words when each key piece of information first appears in the data. l This represents the last occurrence of each key piece of information in the data. The importance of each key piece of information is then determined based on its importance index.
[0022] The system obtains the number of correctly extracted key information, the total number of key information, the total number of all key information that should be extracted, the total number of information for each modality, the request start time, the processing end time, and the processing duration. It then analyzes and processes this data to obtain a performance evaluation index for the multimodal key information extraction and summarization system based on a large language model.
[0023]
[0024] Where RG represents the number of correctly extracted key information, ZG represents the total number of key information, YG represents the total number of key information that should be extracted, and N represents the number of key information that should be extracted. a The total amount of information for each modality, x is the number of modalities in the data, and t is the total amount of information for each modality. S To request the start time, t E To handle the end time, d C The processing time is considered. Therefore, the performance of the multimodal key information extraction and summarization system based on the large language model is derived from the performance evaluation index.
[0025] The display and reminder unit includes a data acquisition module, a report display module, and a warning and reminder module. The data acquisition module is used to acquire the reminder threshold set by the parameter preset unit and the report generated by the core processing unit regarding the multimodal key information extraction and summary system based on a large language model. The report display module is used to display the report on the multimodal key information extraction and summary system based on a large language model acquired by the data acquisition module. The warning and reminder module is used to issue a warning and reminder when the report results on the multimodal key information extraction and summary system based on a large language model acquired by the data acquisition module reach the reminder threshold set by the parameter preset unit.
[0026] This invention also discloses a method for extracting and summarizing multimodal key information based on a large language model, utilizing the aforementioned system for extracting and summarizing multimodal key information based on a large language model, comprising:
[0027] Step 1: Preprocess the data from which key information needs to be extracted and summarized;
[0028] Step 2: Based on the large language model, identify and extract named entities and key event information from the preprocessed data;
[0029] Step 3: Collect the data required for the multimodal key information extraction and summarization system based on the large language model. The data includes data on the number of keywords, data on the number of words, duration data, and time data.
[0030] Step 4: Set the parameters required for the multimodal key information extraction and summarization system based on the large language model. The parameters include the weight coefficient of each modality of the key information and the reminder threshold.
[0031] Step 5: Analyze and process the parameters and data, and generate a report on the multimodal key information extraction and summary system based on a large language model;
[0032] Step 6: Summarize the extracted key information and the generated report, and display the report on the multimodal key information extraction and summary system based on the large language model. At the same time, issue a warning when the report results reach the preset reminder threshold.
[0033] Accordingly, this application also discloses a terminal, including a processor and a storage medium;
[0034] The storage medium is used to store instructions;
[0035] The processor is configured to operate according to the instructions to perform the steps of the aforementioned method for extracting and summarizing multimodal key information based on a large language model.
[0036] Accordingly, this application also discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned method for extracting and summarizing multimodal key information based on a large language model.
[0037] The beneficial effects of this invention are as follows: Compared with the prior art, this invention provides a method and system for extracting and summarizing multimodal key information based on a large language model. It achieves efficient key information extraction by combining a large language model with traditional algorithms, thereby improving the efficiency of key information extraction. Simultaneously, it monitors the processing performance of the system in real time, issuing warnings when system performance reaches a threshold. This allows for timely detection and handling of performance issues in the system to improve its performance. During analysis, the method obtains the quantity of each key piece of information in each modality, the total quantity of key information in each modality, the number of characters in the first occurrence of each key piece of information, and the number of characters in the last occurrence. Based on this data analysis, it obtains the importance index of each key piece of information. A higher importance index indicates greater importance of the key information, requiring more attention during summarization; conversely, a lower importance index indicates lower importance, requiring less attention during summarization. Simultaneously, the performance evaluation index of the multimodal key information extraction and summarization system based on a large language model can be obtained by analyzing and processing the following data: the number of correctly extracted key information, the total number of key information, the total number of all key information that should be extracted, the total number of information for each modality, the request start time, the processing end time, and the processing duration. A higher performance evaluation index indicates better performance of the multimodal key information extraction and summarization system based on a large language model, and vice versa. Attached Figure Description
[0038] Figure 1This is a schematic diagram of the structure of the multimodal key information extraction and summarization system based on a large language model in this invention. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0040] The embodiments described in this application are merely some, not all, embodiments of the present invention. Based on the spirit of the present invention, other embodiments obtained by those skilled in the art without inventive effort are all within the protection scope of the present invention.
[0041] To address the shortcomings of existing technologies, this invention proposes a multimodal key information extraction and summarization method and system based on a large language model. By combining a large language model with traditional algorithms, it achieves efficient key information extraction, thereby improving the efficiency of key information extraction. At the same time, it monitors the processing performance of the multimodal key information extraction and summarization system, including video, audio, images, and text, in real time. When the system performance reaches the warning threshold, it issues a warning and promptly addresses the issue to improve system performance.
[0042] See Figure 1 As shown, the multimodal key information extraction and summarization system based on a large language model disclosed in this invention includes a preprocessing unit, a key information extraction unit, a data acquisition unit, a data transmission unit, a parameter setting unit, a core processing unit, a summary generation unit, and a display and reminder unit.
[0043] The preprocessing unit is used to perform format unification, noise reduction, and cleaning on the data from which key information needs to be extracted, and to perform operations such as word segmentation, part-of-speech tagging, removal of stop words, removal of special characters, removal of punctuation marks, and removal of HTML tags on the data.
[0044] The key information extraction unit is used to identify and extract named entities and event information through a large language model, identify the semantic relationships between entities, and extract the attribute information of entities and events.
[0045] The data acquisition unit is used to collect the number of keywords, the number of words, the duration, and the time data required by the multimodal key information extraction and summarization system based on a large language model.
[0046] In a further embodiment, the data acquisition unit includes: a keyword quantity acquisition unit, a text quantity acquisition unit, a duration acquisition unit, and a time acquisition unit. The keyword quantity acquisition unit is used to acquire data regarding the quantity of keywords required by the multimodal key information extraction and summarization system based on a large language model. The text quantity acquisition unit is used to acquire data regarding the quantity of text required by the multimodal key information extraction and summarization system based on a large language model. The duration acquisition unit is used to acquire duration data required by the multimodal key information extraction and summarization system based on a large language model. The time acquisition unit is used to acquire time data required by the multimodal key information extraction and summarization system based on a large language model.
[0047] The data transmission unit is used to transmit the data collected by the data acquisition unit to the core processing unit, transmit the parameters set by the parameter setting unit to the core processing unit and the display reminder unit, and transmit the report generated by the core processing unit to the summary generation unit and the display reminder unit.
[0048] The parameter setting unit is used to set various parameters required for analysis of the multimodal key information extraction and summarization system based on a large language model.
[0049] In a further embodiment, the parameter setting unit includes: a modality weight setting module and a reminder threshold setting module. The modality weight setting module is used to set the weight coefficient of each modality of the key information extracted by the multimodal key information extraction and summarization system based on a large language model. The reminder threshold setting module is used to set the reminder threshold required by the multimodal key information extraction and summarization system based on a large language model.
[0050] The core processing unit is used to receive the data collected by the data acquisition unit, analyze and process the received data, and generate a report on the multimodal key information extraction and summary system based on the analysis and processing results.
[0051] In a further embodiment, the core processing unit includes: a data receiving module, a data analysis module, and a report generation module. The data receiving module receives the data collected by the data acquisition unit for the multimodal key information extraction and summarization system based on a large language model, as well as the weight coefficients of each modality of the key information extracted by the multimodal key information extraction and summarization system based on a large language model, set by the parameter setting unit. The data analysis module analyzes and processes the data received by the data receiving module for the multimodal key information extraction and summarization system based on a large language model. The report generation module generates a report on the multimodal key information extraction and summarization system based on a large language model, based on the analysis and processing results of the data analysis module.
[0052] In a further embodiment, when the data analysis module performs analysis and processing, it obtains the quantity of each key piece of information in each modality, the total quantity of key information in each modality, the number of words in the first occurrence of each key piece of information in the data, and the number of words in the last occurrence, and obtains the importance index ω of each key piece of information based on the above data analysis and processing.
[0053]
[0054] Where, N a The total amount of information for each modality, n a The number of key pieces of information in each modality, I a The weight coefficient for each modality is the key information, where x is the number of modalities in the data, and n is the number of modalities. f n is the number of words when each key piece of information first appears in the data. l This represents the last occurrence of each key piece of information in the data. The importance of each key piece of information is determined by its importance index; a higher index indicates greater importance and warrants more attention in the summary, while a lower index indicates lower importance and warrants less attention.
[0055] In a further embodiment, the data analysis module, during analysis and processing, obtains the number of correctly extracted key information, the total number of key information, the total number of all key information that should be extracted, the total number of information for each modality, the request start time, the processing end time, and the processing duration. It then analyzes and processes this data to obtain a performance evaluation index for the multimodal key information extraction and summarization system based on a large language model.
[0056]
[0057] Where RG represents the number of correctly extracted key information, ZG represents the total number of key information, YG represents the total number of key information that should be extracted, and N represents the number of key information that should be extracted. a The total amount of information for each modality, x is the number of modalities in the data, and t is the total amount of information for each modality. S To request the start time, t E To handle the end time, d C The processing time is considered. Therefore, the performance of the multimodal key information extraction and summarization system based on the large language model is determined by its performance evaluation index. A higher performance evaluation index indicates better performance, and vice versa.
[0058] The summary generation unit is used to summarize the extracted multimodal key information based on the report on the multimodal key information extraction and summary system based on the large language model transmitted by the core processing unit.
[0059] The display and reminder unit is used to acquire and display the report generated by the core processing unit regarding the multimodal key information extraction and summarization system based on a large language model. At the same time, it provides a warning when the report results reach a preset reminder threshold so that operators can promptly identify problems in the multimodal key information extraction and summarization system based on a large language model and take action to improve the performance of the system.
[0060] In a further embodiment, the display and reminder unit includes a data acquisition module, a report display module, and a warning and reminder module. The data acquisition module is used to acquire the reminder threshold set by the parameter preset unit and the report generated by the core processing unit regarding the multimodal key information extraction and summary system based on a large language model. The report display module is used to display the report on the multimodal key information extraction and summary system based on a large language model acquired by the data acquisition module. The warning and reminder module is used to issue a warning and reminder when the report results on the multimodal key information extraction and summary system based on a large language model acquired by the data acquisition module reach the reminder threshold set by the parameter preset unit.
[0061] This invention also discloses a method for extracting and summarizing multimodal key information based on a large language model, based on the aforementioned system for extracting and summarizing multimodal key information based on a large language model, comprising:
[0062] Step 1: Preprocess the data from which key information needs to be extracted and summarized.
[0063] Specifically, the preprocessing unit performs format unification, noise reduction, and cleaning on the data from which key information needs to be extracted, and performs operations such as word segmentation, part-of-speech tagging, removal of stop words, removal of special characters, removal of punctuation marks, and removal of HTML tags on the data.
[0064] Step 2: Based on the large language model, identify and extract named entities and key event information from the preprocessed data.
[0065] The key information extraction unit identifies and extracts named entities and event information based on a large language model, identifies semantic relationships between entities, and extracts attribute information of entities and events.
[0066] Step 3: Collect the data required for the multimodal key information extraction and summarization system based on the large language model. The data includes data on the number of keywords, data on the number of words, duration data, and time data.
[0067] The data acquisition unit uses the keyword quantity acquisition unit, text quantity acquisition unit, duration acquisition unit, and time acquisition unit to collect data on the number of keywords, the number of texts, the duration, and the time required by the multimodal key information extraction and summarization system based on the large language model.
[0068] Step 4: Set the parameters required for the multimodal key information extraction and summarization system based on the large language model. The parameters include the weight coefficient of each modality of the key information and the reminder threshold.
[0069] The modality weight setting module of the parameter setting unit sets the weight coefficient of each modality of the key information extracted by the multimodal key information extraction and summarization system based on the large language model. The reminder threshold setting module sets the reminder threshold required by the multimodal key information extraction and summarization system based on the large language model.
[0070] Step 5: Analyze and process the parameters and data, and generate a report on the multimodal key information extraction and summary system based on a large language model.
[0071] When the data analysis module performs analysis and processing, it can obtain the quantity of each key piece of information in each modality, the total quantity of key information in each modality, the number of words in the first occurrence of each key piece of information in the data, and the number of words in the last occurrence of each key piece of information. Based on the above data analysis and processing, it can obtain the importance index of each key piece of information, and thus determine the importance of each key piece of information. The higher the importance index, the higher the importance of the key piece of information, and the more attention it should be given in the summary. Conversely, the lower the importance of the key piece of information, the less attention it should be given in the summary.
[0072] Meanwhile, during data analysis, the module can analyze and process data such as the number of correctly extracted key information, the total number of key information, the total number of all key information that should be extracted, the total number of information for each modality, the request start time, the processing end time, and the processing duration to obtain a performance evaluation index for the multimodal key information extraction and summarization system based on a large language model. Based on this performance evaluation index, the system's performance can be assessed; a higher index indicates better performance, and vice versa.
[0073] Step 6: Summarize the extracted key information and the generated report, and display the report on the multimodal key information extraction and summary system based on the large language model. At the same time, issue a warning when the report results reach the preset reminder threshold.
[0074] Specifically, the summary generation unit summarizes the extracted multimodal key information based on the report on the multimodal key information extraction and summary system based on the large language model transmitted by the core processing unit, and displays the report on the multimodal key information extraction and summary system based on the large language model through the report display module of the display and reminder unit. At the same time, when the report results reach the preset reminder threshold, the warning and reminder module will issue a warning.
[0075] The beneficial effects of this invention are as follows: Compared with the prior art, this invention provides a method and system for extracting and summarizing multimodal key information based on a large language model. It achieves efficient key information extraction by combining a large language model with traditional algorithms, thereby improving the efficiency of key information extraction. Simultaneously, it monitors the processing performance of the system in real time, issuing warnings when system performance reaches a threshold. This allows for timely detection and handling of performance issues in the system to improve its performance. During analysis, the method obtains the quantity of each key piece of information in each modality, the total quantity of key information in each modality, the number of characters in the first occurrence of each key piece of information, and the number of characters in the last occurrence. Based on this data analysis, it obtains the importance index of each key piece of information. A higher importance index indicates greater importance of the key information, requiring more attention during summarization; conversely, a lower importance index indicates lower importance, requiring less attention during summarization. Simultaneously, the performance evaluation index of the multimodal key information extraction and summarization system based on a large language model can be obtained by analyzing and processing the following data: the number of correctly extracted key information, the total number of key information, the total number of all key information that should be extracted, the total number of information for each modality, the request start time, the processing end time, and the processing duration. A higher performance evaluation index indicates better performance of the multimodal key information extraction and summarization system based on a large language model, and vice versa.
[0076] Based on the spirit of this invention, those skilled in the art will readily conceive of a computer program product derived from the aforementioned method for extracting and summarizing multimodal key information based on a large language model. The computer program product may include a computer-readable storage medium on which computer-readable program instructions are loaded to enable a processor to implement various aspects of this disclosure. Specifically, this application also includes a terminal comprising a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the aforementioned method for extracting and summarizing multimodal key information based on a large language model.
[0077] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0078] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0079] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. A multimodal key information extraction and summarization system based on a large language model, characterized in that, It includes a preprocessing unit, a key information extraction unit, a data acquisition unit, a data transmission unit, a parameter setting unit, a core processing unit, a summary generation unit, and a display and reminder unit; The preprocessing unit is used to perform format unification, noise reduction, and cleaning on the data from which key information needs to be extracted, and to perform operations such as word segmentation, part-of-speech tagging, removal of stop words, removal of special characters, removal of punctuation marks, and removal of HTML tags on the data. The key information extraction unit is used to identify and extract named entities and event information through a large language model, identify the semantic relationships between entities, and extract the attribute information of entities and events. The data acquisition unit is used to collect the number of keywords, the number of characters, the duration, and the time data required by the multimodal key information extraction and summarization system based on a large language model. The data transmission unit is used to transmit the data collected by the data acquisition unit to the core processing unit. The parameters set by the parameter setting unit are transmitted to the core processing unit and the display reminder unit. At the same time, the report generated by the core processing unit is transmitted to the summary generation unit and the display reminder unit. The parameter setting unit is used to set various parameters required for analysis of the multimodal key information extraction and summarization system based on a large language model; The core processing unit is used to receive the data collected by the data acquisition unit, analyze and process the received data, and generate a report on the multimodal key information extraction and summary system based on the analysis and processing results. The summary generation unit is used to summarize the extracted multimodal key information based on the report on the multimodal key information extraction and summary system based on the large language model transmitted by the core processing unit. The display and reminder unit is used to acquire the report on the multimodal key information extraction and summary system based on the large language model generated by the core processing unit, display the acquired report, and issue a warning when the report results reach a preset reminder threshold.
2. The multimodal key information extraction and summarization system based on a large language model according to claim 1, characterized in that, The data acquisition unit further includes a keyword quantity acquisition unit, a text quantity acquisition unit, a duration acquisition unit, and a time acquisition unit; the keyword quantity acquisition unit is used to collect various data related to the number of keywords required by the multimodal key information extraction and summarization system based on a large language model. The text quantity acquisition unit is used to collect various data related to the text quantity required by the multimodal key information extraction and summarization system based on a large language model. The duration acquisition module is used to acquire duration data required by the multimodal key information extraction and summarization system based on a large language model. The time acquisition module is used to collect various time-related data required by the multimodal key information extraction and summarization system based on a large language model.
3. The multimodal key information extraction and summarization system based on a large language model according to claim 2, characterized in that, The parameter setting unit includes a modality weight setting module and an alert threshold setting module; the modality weight setting module is used to set the weight coefficient of each modality of the key information extracted by the multimodal key information extraction and summarization system based on the large language model. The reminder threshold setting module is used to set the reminder threshold required by the multimodal key information extraction and summarization system based on a large language model.
4. The multimodal key information extraction and summarization system based on a large language model according to claim 3, characterized in that, The core processing unit includes a data receiving module, a data analysis module, and a report generation module. The data receiving module receives all data required by the multimodal key information extraction and summarization system based on a large language model, collected by the data acquisition unit, as well as the weight coefficients of each modality of the key information extracted by the multimodal key information extraction and summarization system based on a large language model, set by the parameter setting unit. The data analysis module analyzes and processes the data required by the multimodal key information extraction and summarization system based on a large language model received by the data receiving module. The report generation module is used to generate a report on the multimodal key information extraction and summary system based on the analysis and processing results of the data analysis module.
5. The multimodal key information extraction and summarization system based on a large language model according to claim 4, characterized in that, The data analysis module is further used to obtain the quantity of each key piece of information in each modality, the total quantity of key information in each modality, the number of words in the first occurrence of each key piece of information in the data, and the number of words in the last occurrence, and to obtain the importance index ω of each key piece of information based on the above data analysis processing: Where, N a The total amount of information for each modality, n a The number of key pieces of information in each modality, I a The weight coefficient for each modality is the key information, where x is the number of modalities in the data, and n is the number of modalities. f n is the number of words when each key piece of information first appears in the data. l The number of words in the last occurrence of each key piece of information in the data; the importance of each key piece of information is determined based on its importance index.
6. The multimodal key information extraction and summarization system based on a large language model according to claim 5, characterized in that, The data analysis module is further used to obtain the number of correctly extracted key information, the total number of key information, the total number of all key information that should be extracted, the total number of information for each modality, the request start time, the processing end time, and the processing duration, and to analyze and process the above data to obtain the performance evaluation index φ of the multimodal key information extraction and summarization system based on a large language model. Where RG represents the number of correctly extracted key information, ZG represents the total number of key information, YG represents the total number of key information that should be extracted, and N represents the number of key information that should be extracted. a The total amount of information for each modality, x is the number of modalities in the data, and t is the total amount of information for each modality. S To request the start time, t E To handle the end time, d C To determine the processing time, the performance of the multimodal key information extraction and summarization system based on the large language model is derived from the performance evaluation index.
7. The multimodal key information extraction and summarization system based on a large language model according to claim 6, characterized in that, The display and reminder unit includes a data acquisition module, a report display module, and a warning and reminder module. The data acquisition module is used to acquire the reminder threshold set by the parameter preset unit and the report on the multimodal key information extraction and summary system based on the large language model generated by the core processing unit. The report display module is used to display the report on the multimodal key information extraction and summary system based on the large language model acquired by the data acquisition module. The warning and reminder module is used to issue a warning and reminder when the report results of the multimodal key information extraction and summary system based on the large language model obtained by the data acquisition module reach the reminder threshold set by the parameter preset unit.
8. A method for extracting and summarizing multimodal key information based on a large language model, based on the multimodal key information extraction and summarization system based on any one of claims 1-7, characterized in that, include: Step 1: Preprocess the data from which key information needs to be extracted and summarized; Step 2: Based on the large language model, identify and extract named entities and key event information from the preprocessed data; Step 3: Collect the data required for the multimodal key information extraction and summarization system based on the large language model. The data includes data on the number of keywords, data on the number of words, duration data, and time data. Step 4: Set the parameters required for the multimodal key information extraction and summarization system based on the large language model. The parameters include the weight coefficient of each modality of the key information and the reminder threshold. Step 5: Analyze and process the parameters and data, and generate a report on the multimodal key information extraction and summary system based on a large language model; Step 6: Summarize the extracted key information and the generated report, and display the report on the multimodal key information extraction and summary system based on the large language model. At the same time, issue a warning when the report results reach the preset reminder threshold.
9. A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to perform the steps of the multimodal key information extraction and summarization method based on a large language model according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the multimodal key information extraction and summarization method based on a large language model as described in any one of claims 1-7.
Citation Information
Patent Citations
Multi-modal recognition method and device based on large language model, electronic equipment and storage medium
CN118410457A
Long text report intelligent creation method and system based on large language model
CN118862856A