Image information extraction method and device, computer equipment and storage medium
By constructing a multimodal large model and training and fine-tuning it, the accuracy of the image information extraction model was improved, the problem of low accuracy in extracting key image information was solved, the high accuracy requirement of insurance underwriting was met, and the underwriting efficiency was improved.
Patent Information
- Application Number
- CN202510896997.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies have low accuracy rates for extracting key information from images, which cannot meet the high accuracy requirements of insurance underwriting.
A multimodal large model is constructed, including a visual encoder, a feature converter, and a decoder. The visual feature extraction capability of the visual encoder and the entity extraction capability of the decoder are improved through training and fine-tuning. Combined with image data augmentation technology, an augmented dataset is constructed for model fine-tuning, and finally a high-performance image information extraction model is obtained.
It improves the accuracy of the image information extraction model, meets the high accuracy requirements of insurance underwriting scenarios, and improves underwriting efficiency.
Smart Images

Figure CN120876877A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and is applied to online processing business scenarios in financial technology. In particular, it relates to a method, apparatus, computer equipment, and storage medium for extracting image information. Background Technology
[0002] In the insurance industry, underwriting is a crucial step, directly impacting an insurance company's ability to accurately assess risk and reasonably determine premium rates. With continuous technological advancements, major insurance companies have begun exploring the use of artificial intelligence (AI) for underwriting tasks—AI underwriting. The core of AI underwriting lies in using artificial intelligence technology to review the underwriting materials submitted by clients, assisting or replacing manual underwriting, thereby improving underwriting efficiency and reducing labor costs.
[0003] In underwriting materials, inpatient medical record images uploaded by clients constitute a significant proportion. Accurately extracting diagnostic information from these images is crucial for underwriting accuracy. Currently, the industry typically employs solutions based on OCR (Optical Character Recognition) + BERT (a pre-trained language model), OCR + LLM (Large Language Model), or VLLM (Vision LLM) for entity recognition to extract key information from medical records. While these solutions improve underwriting efficiency to some extent, their entity recognition accuracy is generally around 85%, making further breakthroughs difficult. This has become one of the bottlenecks restricting the development of AI underwriting technology.
[0004] In recent years, with the rapid development of large-scale models and multimodal large-scale models, the parameter scale of base models has become sufficiently large, and the semantic understanding capabilities of models have been significantly improved. This provides a new opportunity to solve the aforementioned problems. However, in current medical record diagnosis and recognition tasks, the accuracy of fine-tuning of multimodal large-scale models remains low, resulting in low accuracy of extracted key information from medical record images, which cannot meet the urgent need for high accuracy in insurance underwriting. Summary of the Invention
[0005] The purpose of this application is to provide an image information extraction method, apparatus, computer device, and storage medium to solve the technical problem of low accuracy in extracting key image information in the prior art.
[0006] Firstly, an image information extraction method is provided, which employs the following technical solution:
[0007] Obtain a pre-trained image dataset, perform text recognition on the pre-trained image dataset to obtain text recognition results, and stitch the text recognition results with the corresponding images to obtain an image text set;
[0008] The image-text set is input into a pre-constructed multimodal large model, wherein the multimodal large model includes a visual encoder, a feature converter, and a decoder;
[0009] The image text set is input into the multimodal large model for training, and the text prediction result is output. Based on the text prediction result and the text recognition result, the model parameters of the visual encoder and the feature converter are adjusted, and the training continues iteratively until the stopping condition is met to obtain the pre-trained large model.
[0010] Construct a disease database dictionary, and generate a pseudo-image dataset containing diagnostic result labels based on the disease database dictionary;
[0011] The pseudo-image dataset is input into the pre-trained large model to obtain the predicted diagnosis results. Based on the diagnosis result labels and the predicted diagnosis results, the model parameters of the feature converter and the decoder are adjusted, and iterative training continues until the model converges to obtain a candidate model.
[0012] Construct an augmented dataset, and use the augmented dataset to fine-tune the candidate model to obtain the final image information extraction model;
[0013] The image to be extracted is obtained, and the image to be extracted is input into the image information extraction model to obtain diagnostic result information.
[0014] Secondly, an image information extraction device is provided, which adopts the following technical solution:
[0015] The acquisition module is used to acquire a pre-trained image dataset, perform text recognition on the pre-trained image dataset to obtain text recognition results, and concatenate the text recognition results with the corresponding images to obtain an image text set.
[0016] An input module is used to input the image text set into a pre-constructed multimodal large model, wherein the multimodal large model includes a visual encoder, a feature converter, and a decoder;
[0017] The first training module is used to input the image text set into the multimodal large model for training, output text prediction results, adjust the model parameters of the visual encoder and the feature converter based on the text prediction results and the text recognition results, and continue iterative training until the stopping condition is met to obtain the pre-trained large model.
[0018] A construction module is used to build a disease database dictionary and generate a pseudo-image dataset containing diagnostic result labels based on the disease database dictionary;
[0019] The second training module is used to input the pseudo-image dataset into the pre-trained large model to obtain the predicted diagnosis results, adjust the model parameters of the feature converter and the decoder based on the diagnosis result labels and the predicted diagnosis results, and continue iterative training until the model converges to obtain a candidate model.
[0020] The fine-tuning module is used to construct an augmented dataset and use the augmented dataset to fine-tune the candidate model to obtain the final image information extraction model.
[0021] The extraction module is used to acquire the image to be extracted, input the image to be extracted into the image information extraction model, and obtain diagnostic result information.
[0022] Thirdly, a computer device is provided that adopts the technical solution described below:
[0023] The computer device includes a memory and a processor. The memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the image information extraction method described above.
[0024] Fourthly, a computer-readable storage medium is provided, which adopts the technical solution described below:
[0025] The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the image information extraction method described above.
[0026] Compared with the prior art, this application has the following main advantages:
[0027] This application provides an image information extraction method. It constructs a multimodal large-scale model for image information extraction, comprising a visual encoder, a feature converter, and a decoder. First, the decoder is frozen while the visual encoder is trained to improve its visual feature extraction capabilities. Then, the visual encoder is frozen again to improve the decoder's entity extraction capabilities. Next, an augmented dataset is constructed using image data augmentation methods to fine-tune the candidate model, ultimately resulting in a high-performance image information extraction model. This model is able to more deeply understand and learn key information in images, improving the accuracy of image information extraction and ensuring that it meets the high accuracy requirements of insurance underwriting scenarios, thereby improving underwriting efficiency. Attached Figure Description
[0028] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is an exemplary system architecture diagram to which this application can be applied;
[0030] Figure 2 This is a flowchart of one embodiment of the image information extraction method according to this application;
[0031] Figure 3 This is a schematic diagram of the structure of an embodiment of the image information extraction device according to this application;
[0032] Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0034] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0035] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0036] like Figure 1As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.
[0037] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0038] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers, etc.
[0039] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0040] It should be noted that the image information extraction method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the image information extraction device is generally set in the server / terminal device.
[0041] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0042] Continue to refer to Figure 2 The flowchart illustrates an embodiment of the image information extraction method according to this application, including the following steps:
[0043] Step S201: Obtain the pre-trained image dataset, perform text recognition on the pre-trained image dataset to obtain the text recognition result, and stitch the text recognition result with the corresponding image to obtain the image text set.
[0044] Obtaining a pre-training image dataset can be achieved by collecting image report data from medical institutions, including but not limited to various medical images such as X-rays, CT scans, and MRI scans, along with their corresponding diagnostic report texts. These image report data typically contain rich textual information, such as basic patient information, examination results, and diagnostic opinions. The size of the pre-training image dataset is usually between 100,000 and 1 million images to ensure that the model can learn sufficiently rich features.
[0045] In this embodiment, the image information extraction method operates on an electronic device (e.g., Figure 1 The server / terminal device shown can acquire the pre-trained image dataset via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra wideband) connections, and other currently known or future-developed wireless connection methods.
[0046] When performing text recognition on a pre-trained image dataset, optical character recognition (OCR) technology is used to identify the text portions of the images. Deep learning models such as CRNN or Transformer-based OCR models can be used during text recognition, achieving an accuracy of over 95%. The text recognition result includes the identified text content and its location information within the image. The location information includes text detection boxes, the confidence score of each text detection box, and text coordinates, which refer to the angular coordinates of the four corners of the text detection box containing the text.
[0047] In some embodiments, before performing text recognition on the pre-trained image dataset, the images are first preprocessed, including grayscale conversion, binarization, and denoising, to improve the efficiency and accuracy of text recognition.
[0048] The text recognition results are obtained by sorting the text according to its coordinates in the image. Then, the text recognition results are concatenated with the corresponding image to obtain the image text set.
[0049] In some alternative implementations, the steps described above for concatenating the text recognition results with the corresponding images to obtain an image-text set include:
[0050] Obtain the confidence score of the text detection box corresponding to the text recognition result, and filter out the text lines with a confidence score greater than or equal to a preset threshold;
[0051] Obtain the center coordinates of each text detection box corresponding to a text line, and perform line fitting based on the center coordinates to obtain the fitted line for the high confidence line;
[0052] The average slope of the fitted straight line is taken as the rotation slope of each image in the pre-trained image dataset, and the text order of the text recognition result is obtained based on the rotation slope.
[0053] By concatenating the text recognition results in the order of the text, an image text set is obtained.
[0054] Specifically, the center coordinates of the text detection boxes in each image of the text recognition results are obtained, and the text lines corresponding to each image are determined based on the center coordinates; all text detection boxes corresponding to each text line are obtained, the average confidence of all text detection boxes corresponding to each text line is calculated as the text line confidence, and text lines with a confidence greater than or equal to a preset threshold are selected as high-confidence lines to ensure the accuracy of text recognition.
[0055] Obtain the center coordinates of each text detection box corresponding to the high-confidence row. Based on the coordinates of multiple center points in the same row, use the least squares method to fit a straight line y = kx + b, where k is the slope and b is the intercept. The fitting process determines the optimal fitting line parameters by minimizing the sum of squared errors.
[0056]
[0057] In the formula, i represents the i-th text detection box in each high-confidence row; n represents the number of text detection boxes in each high-confidence row; x i The x-axis coordinate of the i-th text detection box is represented by the y-axis coordinate. i This represents the y-axis coordinate of the i-th text detection box.
[0058] In some alternative implementations, the RANSAC algorithm is used instead of the least squares method to improve robustness to outliers. The RANSAC algorithm, through random sampling consensus, can find the best-fit line even in the presence of a large number of noisy points.
[0059] The average of the fitted lines for all high-confidence lines is used as the rotation slope. This slope is used to correct the reading order of the text, especially when the text in the image is tilted. Based on the rotation slope, the text order of the recognition results is obtained. Following this determined text order, the recognition results of each line are sequentially concatenated to form the complete text content. Simultaneously, the correspondence between the text and the original image is preserved, forming an image-text set. Each sample in the image-text set contains the original image and the corresponding text content.
[0060] By piecing together the text recognition results according to a predetermined text order, the model is able to better understand the text structure.
[0061] Step S202: Input the image text set into the pre-constructed multimodal large model, wherein the multimodal large model includes a visual encoder, a feature converter, and a decoder.
[0062] In this embodiment, the pre-built multimodal large model is a deep learning model capable of simultaneously processing image and text data. It adopts the InternVL1b model as its main architecture and includes three main components: a visual encoder, a feature converter, and a decoder. The visual encoder converts the input image into a high-dimensional feature vector to extract visual features; the feature converter maps the visual features to the same language feature space as the decoder, achieving the fusion of visual and linguistic information; the decoder uses a large language model to process text input and generate text output.
[0063] Step S203: Input the image text set into the multimodal large model for training, output the text prediction result, adjust the model parameters of the visual encoder and feature converter based on the text prediction result and text recognition result, and continue iterative training until the stopping condition is met to obtain the pre-trained large model.
[0064] In this embodiment, the visual encoder and feature converter are trained to improve the visual encoder's ability to extract visual features in vertical fields, such as extracting more accurate and richer visual features from medical images.
[0065] In some alternative implementations, the steps described above for training a multimodal large model by inputting the image-text set into the model and outputting text prediction results include:
[0066] The image-text set is input into a multimodal large model, and the visual encoder extracts visual features from the image-text set to obtain the image visual features.
[0067] Image visual features are transformed using a feature converter to obtain image mapping features;
[0068] The image mapping features are autoregressively decoded using a decoder to generate text prediction results.
[0069] Specifically, the visual encoder includes an image segmentation layer, an image embedding layer, and a multi-layer attention fusion layer. The steps described above for extracting visual features from an image-text set using the visual encoder to obtain image visual features include:
[0070] Each image in the image text set is segmented by image segmentation layering to obtain a set of image blocks;
[0071] The image embedding layer transforms the feature vector of each image block in the image block set to obtain the image embedding vector.
[0072] The image embedding vector is input into a multi-layer attention fusion layer, and attention is calculated on the image embedding vector through the attention mechanism to obtain the image visual features.
[0073] The image segmentation layer divides the input image into a grid of preset size, resulting in a set of image blocks. For example, segmenting at 16×16 pixels yields 196 image blocks for a standard 224×224 pixel input image. The image embedding layer converts each image block into a fixed-dimensional (e.g., 768-dimensional) feature vector and adds positional encoding information, enabling the model to perceive the spatial relationships between different image blocks, i.e., the image embedding vector. The multi-layer attention fusion layer consists of multiple attention fusion layers, each containing a multi-head self-attention mechanism and a feedforward neural network. Through the self-attention mechanism, the model can capture the correlations between different regions of the image, forming a global visual representation.
[0074] In some optional implementations, the steps of segmenting each image in the image text set to obtain an image block set through image segmentation include:
[0075] A sliding window approach is used, with a preset window size and preset step size, to slide across each image to obtain multiple initial image blocks;
[0076] The image segmentation algorithm is used to segment multiple initial image blocks to obtain the final set of image blocks.
[0077] Specifically, the image is first initially segmented using a sliding window. The preset window size and preset step size can be set according to actual needs. For example, a 64×64 window size and a step size of 32 pixels are used to slide on the enhanced image to obtain multiple initial image blocks. Then, the multiple initial image blocks are further segmented using an image segmentation algorithm to obtain finer-grained image blocks.
[0078] Image segmentation algorithms can employ models such as Mask R-CNN and Unet. Taking Mask R-CNN as an example, the Mask R-CNN model includes a feature extraction layer, a region proposal layer, an alignment layer, a multi-task branching layer, and an output layer. The feature extraction layer extracts features from each initial image block to obtain semantic features. The semantic features are then input into the region proposal layer to extract regions of interest (ROIs), resulting in candidate regions. The candidate regions are then input into the alignment layer to obtain aligned ROIs. The multi-task branching layer classifies, regresses bounding boxes, and segments the aligned ROIs, yielding corresponding category, bounding box, and instance mask images. The instance mask images are then output as the final image blocks.
[0079] The image embedding layer can use a pre-trained Img2Vec model to convert m image patches into image embedding vectors respectively.
[0080] A multi-layer attention fusion layer can consist of multiple Transformer encoder layers, each containing a multi-head self-attention mechanism and a feedforward neural network. Through the self-attention mechanism, the model can capture the relationships between different regions of the image, generating image visual features that contain global contextual information.
[0081] In some alternative implementations, the multi-layer attention fusion layer can also employ the Performer architecture, a Transformer variant based on a fast attention mechanism. By using linear approximation of self-attention computation, it reduces computational complexity from quadratic to linear, making it suitable for handling large-scale image blocks. For example, the Performer architecture contains 16 layers, each consisting of a connected multi-head self-attention layer and a feature pyramid network. The attention computation in the multi-head self-attention layers incorporates Locality Sensitive Hashing (LSH) to further accelerate large-scale attention computation. Simultaneously, the feature pyramid network structure added after each attention layer is used to fuse feature information at different scales, enhancing the model's ability to understand multi-scale visual information.
[0082] Visual encoders can process detailed information in images more effectively, especially for medical report images containing a lot of text and complex structures. The extracted visual features are more accurate and richer, providing a better foundation for subsequent image information extraction.
[0083] In this embodiment, the feature converter employs a multilayer perceptron structure to map the visual feature space to the semantic feature space. The conversion process includes dimensionality adjustment, nonlinear activation, and normalization operations to ensure that the visual features are compatible with the input format of the large language model (i.e., the decoder).
[0084] In some optional implementations, the feature converter can employ a cross-attention mechanism. This mechanism allows for deeper interaction between visual and textual features, improving the effectiveness of feature conversion. Specifically, the feature converter comprises multiple cross-attention layers, each with multiple attention heads. Cross-attention calculations are performed on the image's visual features through these multiple cross-attention layers to obtain the image-mapped features.
[0085] In this embodiment, the decoder uses the LLM large language model, which can be a model based on the Transformer architecture and uses an autoregressive approach to generate text. During the decoding process, text tags are generated one by one, and each generation depends on the previously generated text content and image mapping features.
[0086] In some alternative implementations, the LLM large language model can adopt a GPT-2-based architecture, containing multiple Transformer decoder layers, each with multiple attention heads. For example, it could contain 12 Transformer decoder layers, each with 16 attention heads, and a hidden layer dimension of 1024. This architecture has a larger decoder scale, enabling the generation of more complex and accurate text content.
[0087] During training, the cross-entropy loss function is used to calculate the difference between the text prediction and text recognition results, and the model parameters are updated using the backpropagation algorithm. During training iterations, the model parameters of the visual encoder and feature converter are adjusted based on the difference between the text prediction and text recognition results until a stopping condition is met, resulting in a pre-trained large model. The stopping condition means that training stops when the loss on the validation set no longer decreases for 5 consecutive epochs, or when the preset maximum number of training epochs (e.g., 100 epochs) is reached.
[0088] By training the visual encoder and feature converter of a multimodal large model, visual features can be better mapped to the semantic feature space, thereby improving the quality of subsequent text generation.
[0089] Step S204: Construct a disease database dictionary and generate a pseudo-image dataset containing diagnostic result labels based on the disease database dictionary.
[0090] In this embodiment, a disease database dictionary is constructed, and a pseudo-image dataset is created using the disease database dictionary to train the decoder, thereby improving the decoder's ability to extract entities in the vertical domain.
[0091] Furthermore, the steps for constructing the disease database dictionary described above include:
[0092] The text recognition results are input into the trained entity extraction model to obtain high-frequency diagnostic entities;
[0093] By combining high-frequency diagnostic entities with manual annotations, we can obtain annotations for the diagnostic entities.
[0094] Obtain all publicly available diagnostic entities, and convert the publicly available diagnostic entities and labeled diagnostic entities according to the preset dictionary structure format to obtain the disease database dictionary.
[0095] In constructing the disease database dictionary, the first step is to use a trained entity extraction model to identify diagnosis-related entities from the text recognition results. The entity extraction model, pre-trained and fine-tuned on medical text corpora, can accurately identify medical diagnostic entities such as disease names, symptom descriptions, and anatomical locations in the text. The entity extraction model can employ the QWen7b large language model.
[0096] High-frequency diagnostic entities refer to disease names and related descriptions that appear frequently in the text recognition results. Initial diagnostic entities are identified from the text recognition results using an entity extraction model. The frequency of each initial diagnostic entity is counted, and entities whose frequency exceeds a preset threshold (e.g., 10 times) are selected as high-frequency diagnostic entities.
[0097] By combining high-frequency diagnostic entities with manually annotated diagnostic entities, a more comprehensive and accurate set of diagnostic entities is formed. The manual annotation is completed by medical experts, ensuring accuracy and professionalism. Professional medical annotation tools can be used during the annotation process, supporting multi-dimensional annotation based on entity category, attributes, relationships, and other dimensions.
[0098] At the same time, standardized diagnostic entities are obtained from public medical knowledge bases such as ICD-10 (International Classification of Diseases, 10th Revision), Chinese Classification of Diseases and Codes, and SNOMED CT (Systematic Medical Terminology).
[0099] All collected publicly available diagnostic entities and labeled diagnostic entities are converted according to a predefined dictionary structure format to form a disease database dictionary. The dictionary structure typically includes the following fields:
[0100] 1) Entity ID: A unique identifier, using UUID format;
[0101] 2) Entity Name: The standard name of the diagnostic entity;
[0102] 3) Entity categories: including major categories such as diseases, symptoms, and anatomical locations;
[0103] 4) Entity subcategories: more granular classifications, such as cardiovascular diseases, respiratory diseases, etc.;
[0104] 5) Synonym set: All synonyms for this entity;
[0105] 6) Superordinate concept: The parent concept of this entity;
[0106] 7) Subordinate concepts: The child concepts of this entity;
[0107] 8) Related entities: Other entities that are highly related to this entity;
[0108] 9) Severity: Divided into four levels: mild, moderate, severe, and extremely severe;
[0109] 10) Source: Indicate the data source for this entity;
[0110] 11) Confidence level: Reliability score of entity information;
[0111] 12) Update time: The timestamp of the most recent update.
[0112] To handle the semantic relationships between diagnostic entities, this embodiment also constructs an entity relationship graph, which stores and manages various relationships between entities through the graph database Neo4j, such as relationship types like "is a", "part-whole", "leads to", and "accompanying".
[0113] During the construction of the disease database dictionary, entity alignment technology was employed to merge and unify the same entity from different sources. The alignment process used a combination of BERT-based semantic similarity calculation and rule matching, achieving high alignment accuracy. The disease database dictionary is stored in JSON format and provides an efficient search interface, supporting various query methods such as exact matching, fuzzy matching, and synonym expansion.
[0114] The disease database dictionary constructed in the above manner has higher coverage, accuracy, and structure, providing more comprehensive and accurate diagnostic entity resources for the subsequent generation of pseudo-image datasets.
[0115] In some alternative implementations, the steps described above for generating a pseudo-image dataset containing diagnostic result labels based on a disease database dictionary include:
[0116] Randomly select a preset number of diagnostic entities from the disease database dictionary;
[0117] Randomly combine the random diagnostic entities and fill them into the preset pseudo-image text template to obtain the pseudo-image diagnostic text.
[0118] Generate a corresponding pseudo-image dataset containing diagnostic result labels based on the pseudo-image diagnostic text.
[0119] A preset number of diagnostic entities are obtained from the disease database dictionary using a random sampling algorithm. These entities are then randomly combined using a random shuffling algorithm and filled into a preset pseudo-image text template to construct pseudo-image diagnostic text.
[0120] The preset pseudo-image text template includes the following main parts:
[0121] 1. Report header: Includes examination type, examination time, patient basic information, etc.;
[0122] 2. Examination Findings: Describe the objective findings observed in the images;
[0123] 3. Diagnostic Results: List the diagnostic conclusions based on the findings of the examination;
[0124] 4. Recommendations: Further examination or treatment recommendations based on the diagnostic results.
[0125] Specifically, the preset pseudo-image text template can follow the following format:
[0126] {key}{split_symbol}{random combination of diagnostic entities};
[0127] The key is randomly selected from the following options: ["Discharge diagnosis", "Discharge Western medicine diagnosis", "Western medicine diagnosis", "Confirmed diagnosis", "Diagnosis", "Postoperative diagnosis", "Intraoperative diagnosis"]; split_symbol represents the symbol or character used to split the string.
[0128] Random combinations of diagnostic entities can be concatenated in the following ways: space-separated, comma-separated, semicolon-separated, no-separation, sequence number-separated (1, 2, 3), sequence number-separated (1)2)3)).
[0129] The preset pseudo-image text template contains multiple variable placeholders to populate the extracted diagnostic entities and their related attributes. The population process considers not only the entities themselves, but also their attributes (such as severity, location, and morphology) and the relationships between entities, generating a more natural and coherent text description.
[0130] To enhance the realism of the text, this embodiment also introduces Natural Language Generation (NLG) technology. A pre-trained GPT-3.5 model is used to refine and expand the filled template text, generating more natural and fluent pseudo-image diagnostic text. The GPT-3.5 model has been fine-tuned with medical text data to generate content that conforms to medical language style.
[0131] Based on the generated pseudo-image diagnostic text, corresponding medical images are generated using text-to-image generation technology, or images matching the pseudo-image diagnostic text are selected from an existing image database. Each generated image is labeled with a corresponding diagnostic result, forming a pseudo-image dataset.
[0132] In this embodiment, when generating the corresponding image based on the pseudo-image diagnostic text, a strategy combining retrieval-based and generation-based methods can be adopted:
[0133] 1) Retrieval-based approach: From existing medical image databases, a semantic matching algorithm is used to retrieve the most similar real image to the generated pseudo-image diagnostic text. The matching algorithm uses the CLIP (Contrastive Language-Image Pretraining) model to calculate the similarity between the text and the image, and selects the image with the highest similarity as the basis.
[0134] 2) Generative Approach: The Stable Diffusion model, fine-tuned with medical image data, is used to generate realistic medical images based on text descriptions. Textual conditional controls and anatomical constraints are employed during the generation process to ensure that the generated images conform to medical standards.
[0135] The images generated by the two methods are blended according to a preset ratio to form the final pseudo-image dataset. Each pseudo-image is labeled with a corresponding diagnostic result, which contains the diagnostic entity and its attribute information, and is stored in structured JSON format.
[0136] The pseudo-image dataset generated by the above method has higher quality and diversity, and is closer to the distribution of real medical data, providing more effective data support for model training.
[0137] Step S205: Input the pseudo-image dataset into the pre-trained large model to obtain the predicted diagnosis results. Adjust the model parameters of the feature converter and decoder based on the diagnosis result labels and the predicted diagnosis results, and continue iterative training until the model converges to obtain the candidate model.
[0138] The fake image dataset is input into a pre-trained large model to obtain predicted diagnostic results. By comparing the differences between the predicted diagnostic results and the true diagnostic results labels, the loss function value is calculated, and the model parameters of the feature converter and decoder are adjusted using the backpropagation algorithm.
[0139] The training process uses a small learning rate for fine-tuning to preserve the visual features learned in the pre-training phase while adapting to the specific needs of the diagnostic entity extraction task. Model convergence refers to reaching the preset maximum number of iterations, or when the current loss function value no longer decreases compared to the previous iteration's loss function value, at which point training stops, and a candidate model is obtained.
[0140] Step S206: Construct an augmented dataset, use the augmented dataset to fine-tune the candidate model, and obtain the final image information extraction model.
[0141] An augmentation dataset is constructed and used to fine-tune the candidate model to obtain the final image information extraction model. The augmentation dataset enhances the generalization ability and robustness of the model by augmenting the original pre-trained image dataset.
[0142] Image data enhancement techniques include displacement enhancement and noise enhancement. Displacement enhancement includes image rotation (e.g., ±10 degrees), scaling (e.g., 0.9-1.1 times), brightness adjustment (e.g., ±10%), and contrast adjustment (e.g., ±10%).
[0143] The candidate model is fine-tuned using an augmented dataset with a smaller learning rate and fewer training epochs to avoid overfitting. During fine-tuning, all components of the model, including the visual encoder, feature converter, and decoder, are optimized simultaneously to achieve the best overall performance.
[0144] After fine-tuning, the model is evaluated on the test set, and metrics such as accuracy, precision, recall, and F1 score are calculated. When the model performance meets the expected goals, it is saved as the final image information extraction model.
[0145] This application uses the F1 score to evaluate the overall performance of the model. Experiments on the final image information extraction model show that the F1 score has been improved by 12.5%, indicating that the entity extraction accuracy of the image information extraction model has been improved by 12.5%.
[0146] Step S207: Obtain the image to be extracted, input the image to be extracted into the image information extraction model, and obtain the diagnostic result information.
[0147] In practical underwriting scenarios, users can upload images (medical records) to be analyzed. The system will input the images to be analyzed into a trained image information extraction model and automatically extract diagnostic information from the images, including key information such as disease name, severity, and anatomical location.
[0148] It should be emphasized that, in order to further ensure the privacy and security of the images to be extracted, the images to be extracted can also be stored in a blockchain node.
[0149] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0150] This application constructs a multimodal large model, including a visual encoder, feature converter, and decoder. During training, the decoder is first frozen while the visual encoder is trained to improve its visual feature extraction capability. Then, the visual encoder is frozen again to improve the entity extraction capability of the decoder. Next, an augmented dataset is constructed using image data augmentation methods to fine-tune the candidate model. Finally, a high-performance image information extraction model is obtained, enabling the model to understand and learn key information in images more deeply, thus improving the accuracy of image information extraction and ensuring that the model can meet the high accuracy requirements of insurance underwriting scenarios.
[0151] In some optional implementations, the steps described above, which involve randomly combining random diagnostic entities and filling them into a preset pseudo-image text template to obtain pseudo-image diagnostic text, include:
[0152] The random combination method is determined by preset logical judgment, and random diagnostic entities are combined according to the random combination method to obtain a sequence of combined entities;
[0153] Obtain a preset pseudo-image text template, and embed the combined entity sequence into the slot corresponding to the preset pseudo-image text template through a template filling algorithm to obtain preliminary pseudo-image diagnostic text.
[0154] A text synthesis algorithm is used to perform semantic smoothing on the preliminary pseudo-image diagnosis text to obtain the final pseudo-image diagnosis text.
[0155] The preset logical judgment can employ a screening algorithm based on co-occurrence frequency. This involves extracting the co-occurrence probabilities of random diagnostic entities from a disease database dictionary, and selecting entity pairs with co-occurrence probabilities higher than a preset probability threshold for combination, resulting in a combined entity sequence. For example, the extracted random diagnostic entities include six entities: "hypertension, diabetes, coronary heart disease, cerebral infarction, cholecystitis, and arthritis." Assuming the co-occurrence probability of "hypertension" and "diabetes" is 0.65, and the co-occurrence probability of "hypertension" and "cholecystitis" is 0.15, and setting a co-occurrence probability threshold of 0.3, entity pairs with co-occurrence probabilities higher than the threshold are selected to obtain the candidate combination {hypertension, diabetes, coronary heart disease, cerebral infarction}, while "cholecystitis" and "arthritis" are removed.
[0156] Obtain a preset pseudo-image text template, and use a filling algorithm tool to map the combined entity sequence one by one to the template slots to obtain the initial filled template content; use a content verification tool to determine the matching degree between the initial filled template content and the diagnostic content. If they do not match, adjust the positions to determine the adjusted filling result; from the adjusted filling result, obtain the entity data of each slot of the preset pseudo-image text template, and use a text generation tool to match it according to business constraints to generate the initial pseudo-image diagnostic text.
[0157] A text synthesis algorithm is used to perform semantic smoothing on the preliminary pseudo-image diagnosis text. If there are semantic conflicts in the preliminary pseudo-image diagnosis text, the entity order in the preliminary pseudo-image diagnosis text is adjusted to obtain the final pseudo-image diagnosis text.
[0158] Specifically, a semantic conflict detection tool is used to compare the diagnostic entities in the preliminary pseudo-image diagnostic text with the context to determine whether there is a conflict. If a conflict exists, the conflicting diagnostic entities are marked to obtain a conflict marking result set. Based on the conflict marking result set, a semantic distance-based sorting algorithm is used to rearrange the marked diagnostic entities, prioritizing the placement of diagnostic entities with higher semantic similarity in adjacent positions to obtain an adjusted entity sequence. Finally, a text content smoothing tool is used to perform semantic consistency processing on the adjusted entity sequence to obtain a smoothed text fragment, which is the final pseudo-image diagnostic text.
[0159] This application improves the text quality of the preliminary pseudo-image diagnosis text by performing semantic smoothing processing, thereby enhancing the realism of the images generated based on the text description.
[0160] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0161] Foundational technologies in artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0162] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0163] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0164] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of an image information extraction device, which is similar to... Figure 2Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0165] like Figure 3 As shown, the image information extraction device 300 described in this embodiment includes: an acquisition module 301, an input module 302, a first training module 303, a construction module 304, a second training module 305, a fine-tuning module 306, and an extraction module 307. Wherein:
[0166] The acquisition module 301 is used to acquire a pre-trained image dataset, perform text recognition on the pre-trained image dataset to obtain text recognition results, and concatenate the text recognition results with the corresponding images to obtain an image text set.
[0167] The input module 302 is used to input the image text set into a pre-constructed multimodal large model, wherein the multimodal large model includes a visual encoder, a feature converter, and a decoder;
[0168] The first training module 303 is used to input the image text set into the multimodal large model for training, output text prediction results, adjust the model parameters of the visual encoder and the feature converter based on the text prediction results and the text recognition results, and continue iterative training until the stopping condition is met to obtain the pre-trained large model.
[0169] The construction module 304 is used to construct a disease database dictionary and generate a pseudo-image dataset containing diagnostic result labels based on the disease database dictionary;
[0170] The second training module 305 is used to input the pseudo-image dataset into the pre-trained large model to obtain the prediction and diagnosis results, adjust the model parameters of the feature converter and the decoder based on the diagnosis result labels and the prediction and diagnosis results, and continue iterative training until the model converges to obtain a candidate model.
[0171] The fine-tuning module 306 is used to construct an augmented dataset and use the augmented dataset to fine-tune the candidate model to obtain the final image information extraction model.
[0172] The extraction module 307 is used to acquire the image to be extracted, input the image to be extracted into the image information extraction model, and obtain diagnostic result information.
[0173] It should be emphasized that, in order to further ensure the privacy and security of the images to be extracted, the images to be extracted can also be stored in a blockchain node.
[0174] The aforementioned image information extraction device 300 constructs a multimodal large model, including a visual encoder, a feature converter, and a decoder. During training, the decoder is first frozen while the visual encoder is trained to improve its visual feature extraction capability. Subsequently, the visual encoder is frozen to improve the entity extraction capability of the decoder. Then, an augmented dataset is constructed using image data augmentation methods to fine-tune the candidate model, ultimately obtaining a high-performance image information extraction model. This enables the model to understand and learn key information in images more deeply, improving the accuracy of image information extraction and ensuring that the model can meet the high accuracy requirements of insurance underwriting scenarios.
[0175] In some alternative implementations, module 301 includes:
[0176] The filtering submodule is used to obtain the confidence level of the text detection box corresponding to the text recognition result and filter out the text lines with a confidence level greater than or equal to a preset threshold as high confidence lines.
[0177] The fitting submodule is used to obtain the center coordinates of each text detection box corresponding to the high confidence line, and to perform straight line fitting based on the center coordinates to obtain the fitting straight line of the high confidence line.
[0178] The acquisition submodule is used to take the average slope of the fitted line as the rotation slope of each image in the pre-trained image dataset, and obtain the text order of the text recognition result based on the rotation slope;
[0179] The splicing submodule is used to splice the text recognition results according to the text order to obtain an image text set.
[0180] By piecing together the text recognition results according to a predetermined text order, the model is able to better understand the text structure.
[0181] In some alternative implementations, the first training module 303 includes:
[0182] The visual feature extraction submodule is used to input the image text set into the multimodal large model, and extract visual features from the image text set through the visual encoder to obtain image visual features;
[0183] The feature conversion submodule is used to convert the visual features of the image through the feature converter to obtain image mapping features;
[0184] The decoding submodule is used to perform autoregressive decoding on the image mapping features through the decoder to generate text prediction results.
[0185] By training the visual encoder and feature converter of a multimodal large model, visual features can be better mapped to the semantic feature space, thereby improving the quality of subsequent text generation.
[0186] In some optional implementations of this embodiment, the visual encoder includes an image segmentation layer, an image embedding layer, and a multi-layer attention fusion layer, and the visual feature extraction submodule includes:
[0187] The segmentation unit is used to segment each image in the image text set through the image segmentation layer to obtain an image block set;
[0188] An embedding unit is used to transform the feature vector of each image block in the image block set through the image embedding layer to obtain an image embedding vector;
[0189] The attention fusion unit is used to input the image embedding vector into the multi-layer attention fusion layer, and perform attention calculation on the image embedding vector through the attention mechanism to obtain the image visual features.
[0190] Visual encoders can process detailed information in images more effectively, especially for medical report images containing a lot of text and complex structures. The extracted visual features are more accurate and richer, providing a better foundation for subsequent image information extraction.
[0191] In some alternative implementations, builder 304 includes:
[0192] The entity extraction submodule is used to input the text recognition results into the trained entity extraction model to obtain high-frequency diagnostic entities.
[0193] The annotation submodule is used to combine the high-frequency diagnostic entities with manual annotations to obtain annotated diagnostic entities;
[0194] A submodule is constructed to obtain all publicly available diagnostic entities, and to convert the publicly available diagnostic entities and the labeled diagnostic entities according to a preset dictionary structure format to obtain a disease database dictionary.
[0195] The disease database dictionary constructed in the above manner has higher coverage, accuracy, and structure, providing more comprehensive and accurate diagnostic entity resources for the subsequent generation of pseudo-image datasets.
[0196] In some alternative implementations, building module 304 also includes:
[0197] An extraction submodule is used to randomly extract a preset number of random diagnostic entities from the disease database dictionary;
[0198] The template filling submodule is used to randomly combine the random diagnostic entities and fill them into a preset pseudo-image text template to obtain pseudo-image diagnostic text.
[0199] The image generation submodule is used to generate a corresponding pseudo-image dataset containing diagnostic result labels based on the pseudo-image diagnostic text.
[0200] The pseudo-image dataset generated by the above method has higher quality and diversity, and is closer to the distribution of real medical data, providing more effective data support for model training.
[0201] In some optional implementations of this embodiment, the template filling submodule includes:
[0202] The combination unit is used to determine the random combination method through preset logic judgment, and combine the random diagnostic entities according to the random combination method to obtain a combined entity sequence;
[0203] The filling unit is used to obtain a preset pseudo-image text template, and to embed the combined entity sequence into the slot corresponding to the preset pseudo-image text template through a template filling algorithm to obtain preliminary pseudo-image diagnostic text.
[0204] The semantic smoothing unit is used to perform semantic smoothing processing on the preliminary pseudo-image diagnostic text using a text synthesis algorithm to obtain the final pseudo-image diagnostic text.
[0205] By performing semantic smoothing on the preliminary pseudo-image diagnostic text, the text quality of the pseudo-image diagnostic text can be improved, thereby enhancing the realism of the images generated based on the text description.
[0206] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0207] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with memory 41, processor 42, and network interface 43 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0208] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0209] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for image information extraction methods. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.
[0210] In some embodiments, the processor 42 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions for the image information extraction method.
[0211] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.
[0212] By constructing a multimodal large model, including a visual encoder, feature converter, and decoder, the training process first freezes the decoder and trains the visual encoder to improve its visual feature extraction capabilities. Then, the visual encoder is frozen to improve the entity extraction capabilities of the decoder. Next, an augmented dataset is constructed using image data augmentation methods to fine-tune the candidate model. Finally, a high-performance image information extraction model is obtained, enabling the model to understand and learn key information in images more deeply, improving the accuracy of image information extraction, and thus ensuring that the model can meet the high accuracy requirements of insurance underwriting scenarios.
[0213] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the image information extraction method described above.
[0214] By constructing a multimodal large model, including a visual encoder, feature converter, and decoder, the training process first freezes the decoder and trains the visual encoder to improve its visual feature extraction capabilities. Then, the visual encoder is frozen to improve the entity extraction capabilities of the decoder. Next, an augmented dataset is constructed using image data augmentation methods to fine-tune the candidate model. Finally, a high-performance image information extraction model is obtained, enabling the model to understand and learn key information in images more deeply, improving the accuracy of image information extraction, and thus ensuring that the model can meet the high accuracy requirements of insurance underwriting scenarios.
[0215] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0216] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
[0217] The software tools or components not belonging to our company that appear in the embodiments of this application are merely examples and do not represent actual use.
Claims
1. A method for extracting image information, characterized in that, Includes the following steps: Obtain a pre-trained image dataset, perform text recognition on the pre-trained image dataset to obtain text recognition results, and stitch the text recognition results with the corresponding images to obtain an image text set; The image-text set is input into a pre-constructed multimodal large model, wherein the multimodal large model includes a visual encoder, a feature converter, and a decoder; The image text set is input into the multimodal large model for training, and the text prediction result is output. Based on the text prediction result and the text recognition result, the model parameters of the visual encoder and the feature converter are adjusted, and the training continues iteratively until the stopping condition is met to obtain the pre-trained large model. Construct a disease database dictionary, and generate a pseudo-image dataset containing diagnostic result labels based on the disease database dictionary; The pseudo-image dataset is input into the pre-trained large model to obtain the predicted diagnosis results. Based on the diagnosis result labels and the predicted diagnosis results, the model parameters of the feature converter and the decoder are adjusted, and iterative training continues until the model converges to obtain a candidate model. Construct an augmented dataset, and use the augmented dataset to fine-tune the candidate model to obtain the final image information extraction model; The image to be extracted is obtained, and the image to be extracted is input into the image information extraction model to obtain diagnostic result information.
2. The image information extraction method according to claim 1, characterized in that, The step of concatenating the text recognition result with the corresponding image to obtain the image text set includes: Obtain the confidence level of the text detection box corresponding to the text recognition result, and filter out the text lines with a confidence level greater than or equal to a preset threshold as high confidence lines; Obtain the center coordinates of each text detection box corresponding to the high confidence line, and perform line fitting based on the center coordinates to obtain the fitted line of the high confidence line; The average slope of the fitted straight line is taken as the rotation slope of each image in the pre-trained image dataset, and the text order of the text recognition result is obtained based on the rotation slope. By concatenating the text recognition results according to the text order, an image text set is obtained.
3. The image information extraction method according to claim 1, characterized in that, The step of inputting the image-text set into the multimodal large model for training and outputting text prediction results includes: The image-text set is input into the multimodal large model, and the visual encoder extracts visual features from the image-text set to obtain image visual features. The image visual features are transformed by the feature converter to obtain image mapping features; The image mapping features are autoregressively decoded using the decoder to generate text prediction results.
4. The image information extraction method according to claim 3, characterized in that, The visual encoder includes an image segmentation layer, an image embedding layer, and a multi-layer attention fusion layer. The step of extracting visual features from the image text set using the visual encoder to obtain image visual features includes: Each image in the image text set is segmented using the image segmentation layer to obtain an image block set; The image embedding layer transforms the feature vector of each image block in the image block set to obtain the image embedding vector. The image embedding vector is input into the multi-layer attention fusion layer, and attention calculation is performed on the image embedding vector through the attention mechanism to obtain the image visual features.
5. The image information extraction method according to claim 1, characterized in that, The steps for constructing the disease database dictionary include: The text recognition results are input into the trained entity extraction model to obtain high-frequency diagnostic entities; The high-frequency diagnostic entities are combined with manual annotations to obtain annotated diagnostic entities; Obtain all publicly available diagnostic entities, and convert the publicly available diagnostic entities and the labeled diagnostic entities according to a preset dictionary structured format to obtain a disease database dictionary.
6. The image information extraction method according to claim 1, characterized in that, The step of generating a pseudo-image dataset containing diagnostic result labels based on the disease database dictionary includes: A preset number of random diagnostic entities are randomly selected from the disease database dictionary; The random diagnostic entities are randomly combined and filled into a preset pseudo-image text template to obtain pseudo-image diagnostic text. Based on the pseudo-image diagnostic text, a corresponding pseudo-image dataset containing diagnostic result labels is generated.
7. The image information extraction method according to claim 6, characterized in that, The step of randomly combining the random diagnostic entities and filling them into a preset pseudo-image text template to obtain pseudo-image diagnostic text includes: The random combination method is determined by a preset logical judgment, and the random diagnostic entities are combined according to the random combination method to obtain a combined entity sequence; A preset pseudo-image text template is obtained, and the combined entity sequence is embedded into the slot corresponding to the preset pseudo-image text template through a template filling algorithm to obtain preliminary pseudo-image diagnostic text. The preliminary pseudo-image diagnostic text is semantically smoothed using a text synthesis algorithm to obtain the final pseudo-image diagnostic text.
8. An image information extraction device, characterized in that, include: The acquisition module is used to acquire a pre-trained image dataset, perform text recognition on the pre-trained image dataset to obtain text recognition results, and concatenate the text recognition results with the corresponding images to obtain an image text set. An input module is used to input the image text set into a pre-constructed multimodal large model, wherein the multimodal large model includes a visual encoder, a feature converter, and a decoder; The first training module is used to input the image text set into the multimodal large model for training, output text prediction results, adjust the model parameters of the visual encoder and the feature converter based on the text prediction results and the text recognition results, and continue iterative training until the stopping condition is met to obtain the pre-trained large model. A construction module is used to build a disease database dictionary and generate a pseudo-image dataset containing diagnostic result labels based on the disease database dictionary; The second training module is used to input the pseudo-image dataset into the pre-trained large model to obtain the predicted diagnosis results, adjust the model parameters of the feature converter and the decoder based on the diagnosis result labels and the predicted diagnosis results, and continue iterative training until the model converges to obtain a candidate model. The fine-tuning module is used to construct an augmented dataset and use the augmented dataset to fine-tune the candidate model to obtain the final image information extraction model. The extraction module is used to acquire the image to be extracted, input the image to be extracted into the image information extraction model, and obtain diagnostic result information.
9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the image information extraction method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the image information extraction method as described in any one of claims 1 to 7.
Citation Information
Cited By
Method for processing multi-modal data and electronic device
CN122596048A