Information extraction method, system and device based on target detection model and multi-modal large model
Through the combination of the object detection model and multimodal large model, the problem of OCR's inability to understand the characteristics of certificates and the identification of non-standardized materials in the financial field is solved, efficient and accurate information extraction of complex materials is achieved, and the overall efficiency of credit business is improved.
Patent Information
- Application Number
- CN202510207762.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-18
AI Technical Summary
The existing OCR technology cannot effectively understand the image characteristics and layout structure of documents in the financial field, cannot extract important information, and lacks recognition capabilities when facing non-standardized materials, cannot correct errors and understand semantics, resulting in inefficient credit business.
A multi-layer hybrid model of the object detection model and a multi-modal large model is adopted to accurately locate key information areas through data acquisition, labeling and training, and extract key information contents in combination with text analysis, image recognition and semantic understanding.
It improves the ability to identify non-standardized materials, improves the accuracy and efficiency of information extraction, reduces the burden of model calculation, and ensures the efficient operation of credit business.
Smart Images

Figure CN120340036A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information recognition and extraction, and more specifically, to an information extraction method, system and device based on a target detection model and a multimodal large model. Background Art
[0002] In the current digital age, the scale of bank credit is increasing continuously. Along with the development of financial technology, the demand for digitalization and standardization of user information is getting higher and higher. For example, standardized user information is used to input into a risk control model for calculation and scoring. However, this also brings a series of challenges, one of which is the information extraction from various non-standard materials. The information extraction of most materials requires people to check one by one with the eyes, which affects the overall operation efficiency of credit business. To solve such problems, the optical character recognition (OCR) technology is widely adopted to identify some highly standardized materials such as identity cards, value-added tax invoices, etc.
[0003] In the current field of fintech, the optical character recognition (OCR) technology has been used in many banks. This method uses a dedicated small model for OCR to analyze image-based materials, thereby reducing the cost and error rate of manual reading and inputting materials, and improving the productivity and efficiency of banks. However, although such technology has made certain progress in improving efficiency, there are still the following problems in the existing OCR recognition solutions:
[0004] 1. Insufficient information interpretation ability: OCR cannot understand the image features of documents, cannot extract important information according to the layout structure, and does not know what type of document the target is.
[0005] 2. Unable to understand semantics: There is such a situation that when the text has errors or is unclear, it is difficult for OCR recognition to effectively correct and supplement it, and thus it is impossible to thoroughly understand the semantic content contained in the document as a whole.
[0006] 3. Weak recognition ability: Since the OCR model is mainly trained based on text data, when encountering differences in the style, background, lighting conditions, etc. of document images, its generalization ability is relatively weak, and it is easily interfered by the changes in the appearance of documents, ultimately resulting in unstable text extraction effects.
[0007] The reason is that: The OCR (optical character recognition) technology mainly focuses on converting the text information in pictures into editable text. However, the forms of credit materials are extremely diverse. In addition to text, they also contain rich non-text visual elements such as seals, signatures, table structures, and special format markings. These elements carry key business information. The simple text extraction of OCR cannot fully exploit and utilize them, easily causing information omission and understanding deviation, and restricting the accuracy of subsequent processes such as credit risk assessment. Summary of the Invention
[0008] The object of the present invention is to provide an information extraction method, system and device based on a target detection model and a multimodal large model, which can better identify and understand the content and key information of a document and extract them.
[0009] The above technical object of the present invention is achieved by the following technical solutions: An information extraction method based on a target detection model and a multimodal large model, comprising the following steps:
[0010] Construct a multi-layer hybrid model including a target detection model and a multimodal large model;
[0011] Connect to the data source, collect training pictures, and obtain the labels of the training pictures;
[0012] According to the training pictures and their labels, train the target detection model until the trained target detection model, when inputting a picture, adds labels to the picture and outputs a picture with labels;
[0013] According to the training pictures and their labels, train the multimodal large model until the trained multimodal large model, when inputting a picture with labels, outputs the content of key information;
[0014] Input the picture to be extracted into the multi-layer hybrid model, and output the content of the key information of the picture to be extracted.
[0015] As a preferred technical solution of the present invention, the labels include position labels and category labels; the position labels are used to mark the range and position of key information on the picture; the category labels are used to classify and mark the selected key information.
[0016] As a preferred technical solution of the present invention, the process of precise processing is as follows: according to the position labels and category labels of the picture, cut and extract the picture with the position of key information; remove the interference content in the cut image segment to obtain several image segments containing key information.
[0017] An information extraction system based on a target detection model and a multimodal large model, comprising:
[0018] A data collection and annotation module, used to connect to the data source, collect training pictures, and obtain the labels of the training pictures;
[0019] The multi-layer hybrid model includes a target detection layer and a multimodal processing layer;
[0020] The target detection layer is used to train the target detection model according to the training pictures and their labels until the trained target detection model, when inputting pictures, adds labels to the pictures and outputs pictures with labels.
[0021] The multimodal processing layer is used to train the multimodal large model according to the training pictures and their labels until the trained multimodal large model outputs key information content when inputting pictures with labels.
[0022] As a preferred technical solution of the present invention, the multi-layer hybrid model further includes a post-processing layer;
[0023] The post-processing layer is used to precisely process the input pictures according to the labels of the input pictures and output several image segments containing key information; the input items of the multimodal processing layer further include image segments containing key information.
[0024] As a preferred technical solution of the present invention, the multimodal large model is configured with text parsing, image recognition, and semantic understanding components.
[0025] As a preferred technical solution of the present invention, the multi-layer hybrid model further includes a result output layer;
[0026] The result output layer is used to convert the key information content output by the multimodal processing layer into a standard format and output it.
[0027] An information extraction device based on a target detection model and a multimodal large model includes: a processor and a memory, the memory stores a computer program executable by the processor, and the processor implements the above method when executing the computer program.
[0028] In summary, the present invention has the following beneficial effects: By combining the multimodal large model with the traditional visual deep neural network, it can better identify and understand the content and key information of documents. The multimodal large model can simultaneously process various types of data such as text and images, integrate different modal information, and can capture the associations between text, format, and image features in complex materials to a certain extent, providing a relatively comprehensive material analysis to help staff better grasp the connotation of complex materials. The target detection model can accurately locate the key areas in complex materials in advance, screen out these high-value information, and then input them into the multimodal large model for targeted processing. This not only reduces the computing burden of the multimodal large model and improves the analysis speed, but also ensures that the model focuses on the most decision-influential information, making the final complex material analysis more efficient, accurate, and deeply meeting the actual needs of the complex and changeable actual business. Description of the Drawings
[0029] Figure 1It is the method flow chart of the present invention;
[0030] Figure 2 It is the schematic diagram of the multi-modal large model of the present invention;
[0031] Figure 3 It is the architecture diagram of the object detection model of the present invention;
[0032] Figure 4 It is the training result diagram of the object detection model of the present invention;
[0033] Figure 5 It is the flow chart of the embodiment of the present invention;
[0034] Figure 6 It is the schematic diagram of the output of the object detection model of the present invention;
[0035] Figure 7 It is the schematic diagram of the output of the post-processing layer. Detailed implementation manners
[0036] The present invention will be further described in detail below with reference to the accompanying drawings.
[0037] As Figure 1 and 2 shown, the present invention provides an information extraction method based on an object detection model and a multi-modal large model, including the following steps:
[0038] S1. Construct a multi-layer hybrid model including an object detection model and a multi-modal large model; specifically, the multi-layer hybrid model further includes a post-processing layer and a result output layer;
[0039] That is, build the architecture of the entire multi-layer hybrid model.
[0040] S2. Connect to the data source, collect training pictures, and obtain the labels of the training pictures;
[0041] The labels include position labels and category labels; the position labels are used to mark the range and position of key information on the picture; the category labels are used to classify and mark the key information selected by the frame.
[0042] Specifically, the connected data source includes the databases of various business systems within the bank. Through timed or manual triggering, the incoming picture materials submitted by users are collected; and then through the manual marking method of the staff, various materials and objects in the pictures are framed. For example: the core content area of the real estate certificate, including information such as the property owner, location, building area, real estate certificate number, etc., and also includes work permits, watermark texts, and so on.
[0043] S3. Input the training pictures and their labels into the multi-layer hybrid model for training. First, input them into the model of the object detection layer.
[0044] Based on the training images and their labels, the object detection model is trained until, when an image is input into the trained object detection model, the model adds labels to the image and outputs the image with labels.
[0045] Specifically, based on a batch of images after data annotation, that is, training images with labels, the Yolo object detection model framework is used for model training. The trained model performs object detection on the images, annotates common key objects in the bank user materials, such as the key information area of the ID card, the key indicator section of the financial statement, the important clause paragraph of the contract, the core content area of various certification documents, the seal and signature positions, etc., decomposes the complex materials into multiple key information segments with clear business directions, and outputs image segments with position labels and category labels.
[0046] S4. The data output by the object detection layer will also undergo precise processing by the post-processing layer. The process of precise processing is as follows: According to the position labels and category labels of the images, the images with key information positions are cropped and extracted; the interference content in the cropped image segments is removed to obtain several image segments containing key information.
[0047] Specifically: Based on the position labels and category labels of various objects output by the object detection layer, the key objects in the images are cropped and extracted; or the interference items in the images, such as watermarks, work permits, etc., are removed. Finally, the information source part in the complex materials is enlarged, the noise part is reduced, and it is transformed into image segments with more clear business information.
[0048] S5. The content output by the post-processing layer is finally input into the multi-modal processing layer. Based on the training images and their labels, the multi-modal large model is trained until, when an image with labels is input into the trained multi-modal large model, the key information content is output;
[0049] Specifically: The multi-modal large model is introduced to receive the images output by the object detection layer and the post-processing layer, combined with text prompts, and through integrating multi-modal ability components such as text parsing, image recognition, and semantic understanding, deeply explores the internal relationships between different modal information, accurately analyzes the specific content in the key image information segments, such as "householder or relationship with the householder" in the household register, the property owner, building area, etc. in the real estate certificate, and outputs in a specified format such as json.
[0050] S6. The output of the multi-modal large model is input into the result output layer: The key information content output by the multi-modal processing layer is converted into a standard format and output.
[0051] Specifically: According to the requirements of the bank's front-end business system, the data extracted and integrated by the multi-modal processing layer is further transformed and encapsulated, structured into a standard format, and fed back to the credit approval system, customer risk assessment system, customer portrait construction system, etc. in the form of API interfaces or data file pushes, etc., to achieve automated business decision-making driven by data.
[0052] S7. Input the picture to be extracted into the multi-layer hybrid model, and the key information content of the picture to be extracted is output. This is the application step after the multi-layer hybrid model is trained. When there is a need for information extraction, input the picture into the trained multi-layer hybrid model, and the corresponding key information content can be obtained.
[0053] Corresponding to the above method, the present invention also provides an information extraction system based on a target detection model and a multi-modal large model, including:
[0054] A data collection and annotation module, used to connect to the data source, collect training pictures, and obtain the labels of the training pictures;
[0055] A multi-layer hybrid model, including a target detection layer, a post-processing layer, a multi-modal processing layer, and a result output layer;
[0056] The target detection layer is used to train the target detection model according to the training pictures and their labels until the trained target detection model, when inputting a picture, adds a label to the picture and outputs a picture with a label;
[0057] The post-processing layer is used to perform precise processing on the input picture according to the label of the input picture and output several image segments containing key information;
[0058] The multi-modal processing layer is used to train the multi-modal large model according to the training pictures and their labels, and the image segments containing key information until the trained multi-modal large model, when inputting a picture with a label, outputs the key information content.
[0059] Specifically, the multi-modal large model is configured with text parsing, image recognition, and semantic understanding components.
[0060] Furthermore, the base model of the multi-modal large model is the general multi-modal large model Qwen2VL-72B. As an open-source model, it has the ability to parse documents and recognize characters that exceed the GPT-4o commercial model in the current (2024.Q4) and reaches SOTA in related fields. Input the image information segments processed by the aforementioned post-processing layer into the multi-modal large model, and by constructing prompt information, extract the corresponding fields and information in the image segments to finally obtain an accurate parsing result.
[0061] The result output layer is used to convert the key information content output by the multi-modal processing layer into a standard format and output it.
[0062] Corresponding to the above methods and systems, the present invention also provides an information extraction device based on an object detection model and a multi-modal large model, including: a processor and a memory. The memory stores computer programs executable by the processor, and when the processor executes the computer programs, the methods of S1-S7 described above are implemented.
[0063] As an embodiment of the present invention, as Figure 3 shown, certain optimizations have also been made to the structure of the object detection model of the present invention, and the specific content is as follows:
[0064] For the model structure of the object detection model, a version suitable for processing small object detection in complex scenes in the YOLO series is selected, such as YOLO11, which includes multiple modules, specifically: 1. Convolutional layer; 2. Spatial Pyramid Pooling Fast (SPPF) and Cross Stage Partial with Spatial Attention (C2PSA); 3. Neck network; 4. Attention mechanism; 5. Head; 6. C3k2 module; 7. CBS module; 8. Final convolutional layer and detection layer.
[0065] 1. Convolutional layer:
[0066] The initial convolutional layer is used to downsample the image. These layers form the basis of the feature extraction process, increasing the number of channels while gradually reducing the spatial dimension. The C3k2 module is introduced. The C3k2 module is a more computationally efficient implementation of the Cross Stage Partial (CSP) bottleneck structure. It uses two smaller convolutions instead of one large convolution. "k2" in C3k2 represents the smaller convolution kernel size, which helps to achieve faster processing speed while maintaining performance.
[0067] 2. Spatial Pyramid Pooling Fast (SPPF) and Cross Stage Partial with Spatial Attention (C2PSA):
[0068] The YOLO11 backbone retains the Spatial Pyramid Pooling Fast (SPPF) module from previous versions and introduces a new Cross Stage Partial with Spatial Attention (C2PSA) module after it. The C2PSA module is a significant addition, which enhances the spatial attention in the feature map. This spatial attention mechanism enables the model to more effectively focus on important regions within the image. By performing spatial pooling on the features, the C2PSA module enables the model to focus on specific regions of interest, potentially improving the detection accuracy of objects of different sizes and positions.
[0069] 3. Neck network:
[0070] The neck network combines features of different scales and transmits them to the head for prediction. This process typically involves upsampling and concatenating feature maps from different levels, enabling the model to effectively capture multi-scale information.
[0071] 4. Attention mechanism:
[0072] The attention to spatial attention is enhanced through the C2PSA module. This attention mechanism enables the model to focus on key regions within the image, potentially achieving more accurate detection, especially for smaller or partially occluded objects.
[0073] 5. Head
[0074] The head is responsible for generating the final prediction results for object detection and classification. It processes the feature maps passed from the neck network and finally outputs the bounding boxes and class labels of the objects within the image.
[0075] 6. C3k2 module
[0076] In the head part, multiple C3k2 modules are used to efficiently process and refine the feature maps. The C3k2 modules are placed in multiple paths within the head to process multi-scale features of different depths. The C3k2 module exhibits different flexibilities according to the value of the C3k parameter:
[0077] When C3k = False, the behavior of the C3k2 module is similar to that of the C2f module, adopting a standard bottleneck structure. When C3k = True, the bottleneck structure is replaced by the C3 module, which allows for deeper and more complex feature extraction.
[0078] Key features of the C3k2 module:
[0079] Faster processing speed: Using two smaller convolutions reduces the computational overhead compared to a single large convolution, thus enabling faster feature extraction. Parameter efficiency: C3k2 is a more compact version of the CSP bottleneck structure, making the architecture more efficient in terms of the number of trainable parameters.
[0080] It should be noted that the C3k module provides greater flexibility by allowing custom convolution kernel sizes. The adaptability of C3k is more significant for extracting more detailed features from images, helping to improve detection accuracy.
[0081] 7. CBS module:
[0082] The head contains several CBS (Convolution - Batch Norm - SiLU) layers after the C3k2 module. These layers further refine the feature maps in the following ways:
[0083] Extract relevant features for accurate object detection.
[0084] Stabilize and normalize the data flow through Batch Norm.
[0085] Utilize the Sigmoid linear unit (SiLU) activation function to achieve non-linearity, thereby enhancing the model performance.
[0086] The CBS module is a basic component in both the feature extraction and detection processes, ensuring that the refined feature maps are passed to the subsequent layers for bounding box and classification predictions.
[0087] 8. Final convolutional layer and detection layer:
[0088] Each detection branch ends with a set of Conv2D layers that reduce the features to the number of outputs required for bounding box coordinates and class predictions. The final detection layer integrates these predictions, including:
[0089] Bounding box coordinates for locating objects in the image; objectness scores indicating the presence or absence of objects; class scores for determining the classes of the detected objects.
[0090] Based on the above, according to the target feature distribution of bank user materials, some layers of the model network structure can be modified, such as adding shallow feature extraction branches to enhance the perception ability of small-sized texts, titles, etc.; adjusting the combination of convolutional kernel scales and strides to adapt to different material layouts and target size differences.
[0091] Specifically, data augmentation techniques can be adopted. Use image-specific annotation tools such as Labelme to manually annotate and select 500 pieces of in-house real data within a limited time, generating a total of more than 2,000 target samples. For the problem of relatively limited and class-imbalanced annotation data materials (more of some key file types and fewer of some special files), adopt diversified data augmentation strategies, including but not limited to image operations such as random cropping, rotation, scaling, brightness and contrast adjustment, affine transformation, and text data augmentation based on text semantic replacement, insertion, and deletion to expand the diversity of training data and improve the generalization ability of the model.
[0092] Specifically, training optimization techniques can also be adopted. Adopt an adaptive learning rate adjustment strategy, combined with the cosine annealing algorithm, to quickly converge in the early stage of model training and finely adjust the weights in the later stage; introduce the focal loss function to mitigate the negative impact of sample class imbalance on model training, ensure the accuracy and stability of the model for detecting various key targets. After multiple rounds of iterative training, make the model reach the leading target detection accuracy index on the bank material test set. Finally, the measured model performance accuracy is 99.5% and the recall rate is 98.7%. Increasing the amount of annotated data can further optimize the generalization ability of the model. For example Figure 4As shown, it is the result of the object detection model in a certain training and various scores after 400 rounds of training.
[0093] As an embodiment of the present invention, the solution of the present invention is docked with the existing business system of the bank, and standardized, secure and reliable interface components are developed based on the solution to realize two-way data interaction with the existing bank infrastructure such as the credit system and the customer management system, ensuring that the new material analysis and extraction system can obtain material data in real time and feedback processing results without affecting the stability of the original business process of the bank, and realizing a business closed-loop.
[0094] As Figure 5 shown, the steps are as follows:
[0095] Step 100: The original picture to be parsed.
[0096] Step 101: The object detection model obtained after training for the real picture data of customers in the bank as described in the above technical solution, and its processing effect is as follows Figure 6 shown.
[0097] Step 102: Process the marked item categories and positions output in 101, crop and extract the key area categories, or remove and mask the interference items such as watermarks and work permits to prevent interference with the work of the visual multi-modal large model (such as extracting wrong names, etc.), and obtain 103.
[0098] Step 103: The image fragments generated after being processed by 102, or the original picture after being processed, as Figure 7 shown.
[0099] 104: The visual multi-modal large model layer receives the input of 103 and extracts relevant information in the key fragments according to the prompt word prompt.
[0100] 105: Receive the output of the large model layer and format it into a standard format such as json as interface data.
[0101] Compared with traditional OCR, the solution of the present invention has the following advantages:
[0102] 1. Stronger ability to handle non-standard materials:
[0103] OCR is mainly applicable to standard category materials (such as ID cards, VAT invoices, etc.) and large text-intensive scenarios. In financial business scenarios such as the credit field, there are a large number of non-standard materials. These materials have a small amount of text but a complex document structure, the text is widely distributed in the document, and the acquisition environment is variable. For example, for relevant materials such as customers' bank statement information, asset information, marriage and guarantor information, OCR performs poorly in such tasks.
[0104] Multimodal large models can integrate information from multiple modalities, such as visual information and text semantic information. It can understand the structure of documents and the logical relationships of the content, and has more advantages in parsing and processing non-standard materials. Through the overall understanding of documents, it can better extract text, even when factors such as text layout, font, and background are complex and variable.
[0105] 2. Superior in-depth integration and application capabilities:
[0106] Multimodal large models can be deeply integrated with the front-end operation interface of credit domain applications. It can not only extract text, but also perform preliminary semantic understanding and association on the text information after extraction. For example, when checking the quality of materials uploaded by customer managers during business development, it matches the extracted text information with business rules to determine whether the materials are complete and the information is accurate. OCR is often just a simple text extraction tool, lacking the ability to deeply integrate with business systems and perform semantic understanding. A large amount of manual work is required later to process and judge the information association after text extraction.
[0107] 3. Better scalability and versatility:
[0108] When processing text extraction, multimodal large models can also consider information from other modalities, such as voice and images. This provides more possibilities for the expansion of financial business scenarios. For example, in the customer-facing process, it can perform comprehensive processing based on image information and voice information to guide customers to handle business autonomously.
[0109] 4. Greater potential to adapt to future development trends:
[0110] The processing of multimodal information is an important ability in future development trends such as AI Agent and embodied intelligence. Multimodal large models are at the forefront of this development wave. It can continuously improve its related functions such as text extraction through data feedback and model updates. OCR technology is relatively mature, and there is relatively little room for improvement and expansion when dealing with new and complex business scenarios and technology integration trends.
[0111] In summary, the technical advantages of the solution of the present invention are as follows:
[0112] This solution combines multimodal large models with traditional visual deep neural networks to better identify and understand the content and key information of documents. Multimodal large models can simultaneously process various types of data such as text and images, integrate different modal information, and can capture the associations between text, format, and image features in credit materials to a certain extent, providing a relatively comprehensive material analysis and helping credit review personnel better grasp the connotations of complex materials. This is its advantage compared to OCR.
[0113] However, when dealing with a vast amount of complex credit materials, directly analyzing all data at once will face the dilemmas of high cost and low efficiency for multi-modal large models. At this time, the pre-intervention of the object detection model becomes particularly necessary. The object detection model can accurately locate the key areas in credit materials in advance, such as the core clause paragraphs, important financial data tables, key parts of collateral asset pictures, etc., screen out these high-value information, and then input them into the multi-modal large model for targeted processing. In this way, it not only reduces the computational burden of the multi-modal large model and improves the parsing speed, but also ensures that the model focuses on the information with the most decision-making influence, making the final credit material parsing more efficient, accurate, and deeply meeting the complex and ever-changing actual needs of the credit business.
[0114] More specifically, the innovative advantages of the present invention are as follows:
[0115] 1. Precise positioning and extraction:
[0116] Traditional OCR (Optical Character Recognition) often performs text recognition on the entire page of the incoming document materials, lacking the ability to precisely locate fragments of the target document image. The self-trained object detection model can first accurately identify the fragments of the target document image that need attention, such as specific contract clause areas, key identity information areas, etc., avoiding the recognition and processing of irrelevant content, reducing subsequent noise interference, and thus improving the accuracy of key information extraction.
[0117] 2. Strong adaptability to complex layouts:
[0118] The layout forms of customer incoming document materials are diverse, and there may be complex situations such as mixed text and charts, and mixed use of different fonts and sizes. Traditional OCR is prone to problems such as recognition errors and inaccurate character segmentation when dealing with such complex layouts. The self-trained object detection model can focus on the target fragments with relatively stronger regularity, and then the multi-modal large model extracts information, which can better adapt to complex layouts and improve the overall recognition accuracy and information extraction quality.
[0119] 3. Reducing data complexity:
[0120] When using the multi-modal large model alone, it is necessary to directly process the entire incoming document materials, which contain a large amount of complex information, including both text and image and other multi-modal data mixed together, increasing the difficulty for the model to understand and extract key information. However, first extracting the fragments of the target document image through the self-trained object detection model is equivalent to a preliminary screening and simplification of the original data, reducing the data complexity faced by the multi-modal large model, enabling it to more focusedly extract effective information in the key areas.
[0121] 4. Focusing on key content:
[0122] Although multimodal large models are powerful in themselves, when faced with a vast amount of incoming materials, without prior target screening and guidance, they may disperse their energy on some non-critical information, affecting the extraction effect of core and key information. After the self-training object detection model clearly extracts the target document image fragments, the multimodal large model can specifically analyze these contents related to the business core, such as focusing on key parts like income certificates, images and text descriptions related to collateral in loan applications, enhancing the pertinence and accuracy of key information extraction.
[0123] 5. Improve the stability of model performance:
[0124] Since it reduces the amount and complexity of data processed by the multimodal large model, it avoids situations such as model overload and incorrect judgments that may be caused by processing too much irrelevant or redundant information, helping to maintain the stability of the multimodal large model's performance. In the long-term and large-scale key information extraction work of customer incoming materials, it can always maintain a good output effect and serve the bank's business process more reliably.
[0125] In summary, the method of using a self-training object detection model combined with a multimodal large model to extract key information from customer incoming materials has significant advantages in terms of accuracy, efficiency, and model performance.
[0126] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as within the protection scope of the present invention.
Claims
1. An information extraction method based on a target detection model and a multimodal large model, characterized in that: It includes the following steps: Construct a multi-layer hybrid model including an object detection model and a multi-modal large model; Connect to the data source, collect training images, and obtain the labels of the training images; Train the object detection model according to the training images and their labels until the trained object detection model adds labels to the input images and outputs images with labels when the images are input; Train the multi-modal large model according to the training images and their labels until the trained multi-modal large model outputs the key information content when the images with labels are input; Input the image to be extracted into the multi-layer hybrid model, and output the key information content of the image to be extracted.
2. The information extraction method based on a target detection model and a multi-modal large model according to claim 1, wherein: The labels include position labels and category labels; the position labels are used to mark the range and position of the key information on the image; the category labels are used to classify and mark the selected key information.
3. The information extraction method based on an object detection model and a multi-modal large model according to claim 2, wherein: The process of precision processing is as follows: according to the position label and category label of the image, crop and extract the image with the position of the key information; remove the interference content in the cropped image fragment to obtain several image fragments containing the key information.
4. An information extraction system based on an object detection model and a multimodal large model, characterized in that: It includes: A data collection and annotation module, which is used to connect to the data source, collect training images, and obtain the labels of the training images; The multi-layer hybrid model includes an object detection layer and a multi-modal processing layer; The object detection layer is used to train the object detection model according to the training images and their labels until the trained object detection model adds labels to the input images and outputs images with labels when the images are input; The multi-modal processing layer is used to train the multi-modal large model according to the training images and their labels until the trained multi-modal large model outputs the key information content when the images with labels are input.
5. An information extraction system based on an object detection model and a multi-modal large model according to claim 4, characterized in that: The multi-layer hybrid model further includes a post-processing layer; The post-processing layer is used to perform precision processing on the input image according to the label of the input image and output several image fragments containing the key information; the input item of the multi-modal processing layer further includes the image fragments containing the key information.
6. The information extraction system based on the object detection model and the multimodal large model according to claim 5, characterized in that: The multi-modal large model is configured with text parsing, image recognition, and semantic understanding components.
7. An information extraction system based on an object detection model and a multi-modal large model according to claim 6, characterized in that: The multi-layer hybrid model further includes a result output layer; The result output layer is used to convert the key information content output by the multi-modal processing layer into a standard format and output it.
8. An information extraction device based on an object detection model and a multimodal large model, characterized in that: It includes: A processor and a memory, the memory stores a computer program executable by the processor, and when the processor executes the computer program, it implements the method described in any one of claims 1-3.
Citation Information
Cited By
Contract signature identification method based on large and small model collaboration
CN121095966A