Multi-modal case database construction method and device, electronic equipment and medium
By performing layout area detection and logical hierarchical tree construction on multimodal medical case documents, and combining multimodal large language models for case segmentation and information extraction, the problems of low efficiency and low accuracy in constructing multimodal case databases are solved, and efficient and accurate multimodal case database construction and storage are achieved.
Patent Information
- Application Number
- CN202511157903.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-08-19
AI Technical Summary
When constructing a multimodal case database, the existing technology has the problems of low construction efficiency and long construction time due to the complex and diverse formats of case data. Regular expressions are unable to understand the deep semantic logic and fuzzy expressions of the text, resulting in low information extraction accuracy and waste of resources.
By performing document layout area detection on multimodal medical case documents, constructing a regional logical hierarchical tree, using a multimodal large language model to perform case segmentation and information extraction, generating case information extraction prompt words, and performing multimodal information association and storage, intelligent organization and structured storage are achieved.
It improves the efficiency of building multimodal case databases, reduces storage resource waste, improves data accuracy and quality, and adapts to medical documents in different formats and updated formats.
Smart Images

Figure CN120656630A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of computer technology, and in particular to a method, device, electronic device, and medium for constructing a multimodal case database. Background Art
[0002] With the rapid development of medical technology, the complexity and volume of medical data are increasing. The presence of multimodal data in medical cases, coupled with the complex and diverse formats of medical cases, is a significant factor influencing the construction of medical databases. A common approach to building a multimodal case database is to construct a set of medical regular expressions for content matching and extraction of medical cases. Then, using this set of medical regular expressions, content matching, extraction, and integration are performed on multimodal medical cases to generate a multimodal integrated medical case database. Finally, these integrated medical cases are stored to create a multimodal case database.
[0003] However, it has been found in practice that when the above method is used to construct a multimodal case database, the following technical problems often arise: Due to the complex and diverse formats of case data, when regular expressions are used to match and extract content from case data, targeted regular expressions need to be constructed for each format of case data, and when the case data (for example, data format, extraction requirements, medical terminology) changes, the regular expression needs to be rebuilt, resulting in low efficiency in case database construction, long construction time, and a large amount of waste of resources; at the same time, regular expressions only match data and extract content through the literal meaning of the text, and cannot understand the deep semantic logic and fuzzy expressions of the text, and the efficiency and accuracy of data extraction and matching of other modalities besides text are low, and there are problems of omission and erroneous extraction of information, resulting in low accuracy in case data content extraction, a large amount of errors and redundant information, and a waste of data storage resources and low database quality.
[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention
[0005] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0006] Some embodiments of the present disclosure propose a method, apparatus, electronic device, and medium for constructing a multimodal case database to solve one or more of the technical problems mentioned in the above background technology section.
[0007] In a first aspect, some embodiments of the present disclosure provide a method for constructing a multimodal case database, comprising: performing document layout area detection processing on the acquired multimodal medical case documents to obtain a case area information set; performing a region logic hierarchy tree construction on the above case area information set to obtain a case area logic hierarchy tree; performing case segmentation processing on the above multimodal medical case documents based on the above case area logic hierarchy tree and the above case area information set to obtain a medical segmentation case set; generating case information extraction prompt word information for the above medical segmentation case set; utilizing a multimodal large language model to extract prompt word information based on the above case information, performing structured information extraction processing on the above medical segmentation case set to obtain a case text structured information set; performing position association category identification on the above medical segmentation case set based on the above case area information set to obtain a medical image category label information set; performing multimodal information association on the above case text structured information set and the above medical image category label information set to obtain a multimodal synthetic medical case information set, and performing case storage on the above multimodal synthetic medical case information set to obtain a multimodal case database.
[0008] In a second aspect, some embodiments of the present disclosure provide a multimodal case database construction device, comprising: a document layout area detection unit, which performs document layout area detection processing on the acquired multimodal medical case document to obtain a case area information set; a region logic hierarchy tree construction unit, which performs region logic hierarchy tree construction on the above case area information set to obtain a case area logic hierarchy tree; a case segmentation unit, which performs case segmentation processing on the above multimodal medical case document according to the above case area logic hierarchy tree and the above case area information set to obtain a medical segmentation case set; a generation unit, which generates case information extraction prompt word information for the above medical segmentation case set; a case The information extraction unit uses a multimodal large language model to extract prompt word information based on the above-mentioned case information, performs structured information extraction processing on the above-mentioned medical segmentation case set, and obtains a case text structured information set; the position association category identification unit performs position association category identification on the above-mentioned medical segmentation case set based on the above-mentioned case region information set, and obtains a medical image category label information set; the multimodal information association unit performs multimodal information association on the above-mentioned case text structured information set and the above-mentioned medical image category label information set to obtain a multimodal synthetic medical case information set, and performs case storage on the above-mentioned multimodal synthetic medical case information set to obtain a multimodal case database.
[0009] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0010] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.
[0011] The above-mentioned embodiments of the present disclosure have the following beneficial effects: the multimodal case database construction method of some embodiments of the present disclosure can automatically extract and integrate multimodal medical cases, improve the efficiency of database construction and reduce the waste of storage resources. Specifically, the reasons for the low efficiency of case database construction, long construction time, and waste of a large amount of resources, as well as the low database quality are: due to the complex and diverse formats of case data, when using regular expressions to match case data content extraction, it is necessary to construct targeted regular expressions for each format of case data, and when the case data (for example, data format, extraction requirements, medical terminology) changes, the regular expression needs to be reconstructed, resulting in low efficiency of case database construction, long construction time and waste of a large amount of resources; at the same time, regular expressions only match data and extract content based on the literal meaning of the text, and cannot understand the deep semantic logic and fuzzy expressions of the text. In addition, the extraction and matching of data in other modalities besides text are inefficient and inaccurate, and there are problems of omission and erroneous extraction of information, resulting in low accuracy of case data content extraction, a large amount of errors and redundant information, and waste of data storage resources and low database quality. Based on this, the multimodal case database construction method of some embodiments of the present disclosure can first perform document layout area detection processing on the acquired multimodal medical case documents to obtain a case area information set. Here, layout analysis can be performed on multimodal medical case documents of different formats and different medical fields to accurately grasp the layout of the multimodal medical case documents without the need to construct and update complex regular expressions. Thirdly, a regional logical hierarchy tree is constructed on the above-mentioned case area information set to obtain a case area logical hierarchy tree. Here, the construction of the case area logical hierarchy tree can better understand the semantic information, layout information and hierarchical structure between different areas of the multimodal medical case document, and can adapt to documents of different formats. Then, based on the above-mentioned case area logical hierarchy tree and the above-mentioned case area information set, the above-mentioned multimodal medical case document is subjected to case segmentation processing to obtain a medical segmentation case set. Here, segmenting the multimodal medical case document into cases of the same patient can improve the correlation and segmentation between multimodal data and avoid the presence of a large amount of erroneous data in subsequent associated storage. Subsequently, case information extraction prompt word information for the above-mentioned medical segmentation case set is generated. Here, this information is used to guide subsequent information extraction by the multimodal large language model, improving its efficiency and accuracy. Subsequently, the multimodal large language model is used to extract prompt word information based on the aforementioned case information. Structured information extraction is then performed on the aforementioned medical case set, resulting in a structured information set of case text. The multimodal large language model can identify the deep semantic logic and fuzzy expressions of the different model data within the case, improving efficiency and accuracy.Then, based on the above-mentioned case region information set, the above-mentioned medical segmentation case set is subjected to position association category identification to obtain a medical image category label information set. Here, through the medical image position association, the recognition and segmentation accuracy of the medical segmentation case to which the medical image belongs can be improved, and the accuracy of medical image category identification can be further improved by the assistance of different modal data. Finally, the above-mentioned case text structured information set and the above-mentioned medical image category label information set are subjected to multimodal information association to obtain a multimodal synthetic medical case information set, and the above-mentioned multimodal synthetic medical case information set is subjected to case storage to obtain a multimodal case database. Here, the accuracy of multimodal information association and the quality of the multimodal synthetic medical case information set can be improved, the waste of storage resources and the construction efficiency of the multimodal case database can be reduced, and the construction market can be shortened. It can be concluded that the multimodal case database construction method can adapt to medical documents of different formats and updated formats through layout area segmentation and logical hierarchical tree construction. Then, through customized prompt word information and multimodal large language model, it automatically extracts and associates and integrates data of different modalities such as images and texts, realizes intelligent organization and structured storage of multimodal medical case documents, improves the quality of multimodal synthetic medical case information, reduces the waste of storage resources, improves the construction efficiency of the multimodal case database and shortens the construction time. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0013] Figure 1 is a flowchart of some embodiments of the method for constructing a multimodal case database according to the present disclosure; Figure 2 is a schematic diagram of a process for constructing a multimodal case database according to some embodiments of the multimodal case database construction method of the present disclosure; Figure 3 is a schematic structural diagram of some embodiments of the multimodal case database construction device according to the present disclosure; Figure 4 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0014] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0015] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.
[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0017] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0018] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0019] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0020] Figure 1 A process 100 of some embodiments of a multimodal case database construction method according to the present disclosure is shown. The multimodal case database construction method includes the following steps: Step 101 : Perform document layout area detection on the acquired multimodal medical case document to obtain a case area information set.
[0021] In some embodiments, the execution entity (e.g., an electronic device) of the multimodal case database construction method can perform document layout area detection on the acquired multimodal medical case document to obtain a case area information set. The acquired multimodal medical case document can be acquired via a wired or wireless connection. The multimodal medical case document can be a document that records the patient's physical health status in multiple data formats. The multiple data formats can include, but are not limited to, at least one of the following: images, text, physical sign information, and time series data. The case area information in the case area information set can be the location information and element category information of different areas in the detected multimodal medical case document. The element category information can include, but is not limited to, at least one of the following: document title, chart, medical image, formula, abstract, reference, and text paragraph.
[0022] As an example, the execution entity may input the multimodal medical case document into a PP-DocLayout (document picture layout detection) model to perform document layout area detection to obtain a medical record area information set.
[0023] In some optional implementations of some embodiments, the above-mentioned document layout area detection processing on the acquired multimodal medical case document to obtain the case area information set may include the following steps: In the first step, the multimodal medical case document is input into the activated convolutional feature extraction network included in the document layout region detection model to obtain a first case document feature map. The document layout region detection model further includes a residual activation feature extraction backbone network, a cross-stage fusion network, a multi-scale visual feature enhancement network, and multiple multi-scale cross-feature fusion networks. The first case document feature map can represent the shallow visual information of the multimodal medical case document. The document layout region detection model can be a neural network model that parses and partitions the input multimodal medical case document into different regions to determine the location, semantic, and category information of each region. The activated convolutional feature extraction network can be a neural network model that extracts visual features from the input multimodal medical case document. The activated convolutional feature extraction network can include a standard convolutional layer, a batch normalization (BN) layer, and a leaky ReLU (Leaky Rectified Linear Unit) activation function. The residual activation feature extraction backbone network can be a deep neural network model that extracts deep visual features from the input first case document feature map. The residual activation feature extraction backbone network may include: an activated convolutional feature extraction network, a CSP1_1 network (Cross Stage Partial Connections), an activated convolutional feature extraction network, a CSP1_3 network, an activated convolutional feature extraction network, a CSP1_3 network, an activated convolutional feature extraction network, and an SPP (Spatial Pyramid Pooling) network. The first 1 in the CSP1_1 network indicates that the CSP1_1 network includes one activated convolutional feature extraction network, and the second 1 indicates that the CSP1_1 network includes one residual component. The cross-stage fusion network may be a deep neural network model that fuses features from the outputs of the input cross-stage fusion network. The cross-stage fusion network may include: a residual activation feature extraction backbone network, an upsampling module, a feature fusion module, a CSP2_1 network, a residual activation feature extraction backbone network, and an upsampling module. The multi-scale visual feature enhancement network can be a neural network model that uses a convolutional attention module (CBAM) to enhance the spatial position information of the input first case document feature map and the output of the cross-stage fusion network, and uses multi-branch convolution to construct dependencies at different scales to obtain enhanced shallow visual features. The multiple multi-scale cross-feature fusion networks can be a multi-scale cross-feature fusion module that replaces the concat fusion method and adaptively fuses features from different levels through an attention mechanism.
[0024] In the second step, the first case document feature map is input into the residual activation feature extraction backbone network to obtain a second case document feature map and a third case document feature map. The second case document feature map may be a feature map output by the second CSP1_3 network included in the residual activation feature extraction backbone network. The third case document feature map may be a feature map output by the final residual activation feature extraction backbone network.
[0025] In the third step, the second case document feature map and the third case document feature map are input into the cross-stage fusion network to obtain a fourth case document feature map and a fifth case document feature map. The fourth case document feature map may be a feature map output by the second activated convolutional feature extraction network included in the cross-stage fusion network. The fifth case document feature map may be a feature map output by the final cross-stage fusion network.
[0026] The fourth step is to input the first case document feature map into the multi-scale visual feature enhancement network to obtain the sixth case document feature map.
[0027] In the fifth step, the fifth case document feature map and the sixth case document feature map are input into the multiple multi-scale cross-feature fusion networks to obtain a case region information set. Each of the multiple multi-scale cross-feature fusion networks may include a standard convolutional network, a depthwise separable convolutional network, and a cross-attention mechanism layer. The number of the multiple multi-scale cross-feature fusion networks may be the same as the number of case region information included in the case region information set. For example, the multiple multi-scale cross-feature fusion networks may be three multi-scale cross-feature fusion networks. In practice, the execution entity may first input the fifth case document feature map and the sixth case document feature map into the standard convolutional network and the depthwise separable convolutional network included in the first multi-scale cross-feature fusion network, respectively, to obtain a first case document local channel feature map and a second case document local channel feature map. The standard convolutional network may be a model that performs local representation of the input fifth case document feature map and sixth case document feature map. The depthwise separable convolutional network can be a model that maps the fifth and sixth case document feature maps to a high-dimensional space and enriches channel information. The multiple multi-scale cross-feature fusion networks can be three multi-scale cross-feature fusion networks. Next, the first and second case document local channel feature maps are flattened to obtain a first document local channel flattened feature map set and a second document local channel flattened feature map set. The flattening can be performed to create non-overlapping flat blocks that facilitate long-term dependency matching of the fifth and sixth case document feature maps with spatial induction bias. The shapes of the fifth and sixth case document feature maps can both be H*W*C. The shape of the first case document local channel feature map can be H*W*d, where d>C. The shape of the first document local channel flattened feature map can be P*N*d, where P=WH and N=HW / P. H can represent the height of the feature map, W can represent the width of the feature map, and d can represent the number of channels in the feature map. The above N can represent the number of first document local channel flat feature maps included in the first document local channel flat feature map set. Again, the above first document local channel flat feature map set and the above second document local channel flat feature map set are input into the cross-attention mechanism layer to obtain the medical record document cross-fusion feature map set. Then, the above medical record document cross-fusion feature map set is convolutionally spliced to obtain the document multi-scale fusion feature map. In practice, the above execution entity can first flatten the feature map of the above case document cross-fusion feature map to obtain a flattened case document feature map. Among them, the shape of the above flattened case document feature map can be H*W*d.Secondly, the flattened case document feature map is input into a depthwise separable convolutional network to be mapped to a C-dimensional space to obtain a document mapping feature map. The shape of the document mapping feature map may be H*W*C. Then, the fifth case document feature map, the sixth case document feature map and the document mapping feature map are feature spliced to obtain a spliced feature map. The shape of the spliced feature map may be H*W*3C. Finally, the spliced feature map is input into a standard convolutional layer for feature fusion representation to obtain a document multi-scale fusion feature map. Then, the third case document feature map, the fourth case document feature map and the document multi-scale fusion feature map are input into multiple multi-scale cross feature fusion networks after removing the first multi-scale cross feature fusion network to obtain the seventh case document feature map and the eighth case document feature map. Finally, the seventh case document feature map, the eighth case document feature map and the document multi-scale fusion feature map are subjected to convolution prediction processing to obtain a case area information set. In practice, the execution entity may first input the seventh case document feature map, the eighth case document feature map, and the document multi-scale fusion feature map into a 3*3 convolutional layer and a 1*1 convolutional layer, respectively, to obtain a first bounding box set, a second bounding box set, and a third bounding box set. The first bounding box set, the second bounding box set, and the third bounding box set are then sorted by category score and filtered by non-maximum suppression to obtain case region information sets of different sizes.
[0028] In some optional implementations of some embodiments, the multi-scale visual feature enhancement network includes: a global pooling channel attention network and a variable convolutional spatial attention network. The global pooling channel attention network can capture information from each channel through a global average pooling layer and a global maximum pooling layer to explicitly model channel interdependence to enhance the learning of convolutional features and fully consider the correlation between channels from a global perspective. The variable convolutional spatial attention network can be a deep neural network model that utilizes a variable convolutional space to calculate spatial attention.
[0029] Optionally, the step of inputting the first case document feature map into the multi-scale visual feature enhancement network to obtain a sixth case document feature map may include the following steps: In the first step, the first case document feature map is input into the global average pooling layer and the global maximum pooling layer included in the global pooling channel attention network to obtain the case channel average pooling feature map and the case channel maximum pooling feature map.
[0030] The second step is to perform channel nonlinear fusion on the above-mentioned case channel average pooling feature map and the above-mentioned case channel maximum pooling feature map to obtain the case average channel nonlinear feature map and the case maximum channel nonlinear feature map. The above-mentioned channel nonlinear fusion can first be through a 1*1 convolution layer, RuLU activation function and 1*1 convolution layer to adjust the feature map dimension and reduce computational complexity, and then use the Sigmoid function to learn the nonlinear interaction relationship between channels, that is, the non-mutually exclusive relationship between channels, to ensure that the information of multiple channels can be emphasized rather than only emphasizing the nonlinear fusion processing of a single channel.
[0031] The third step is to perform weighted summation on the case average channel nonlinear feature map, the case maximum channel nonlinear feature map and the first case document feature map to obtain the case channel attention feature map.
[0032] In the fourth step, the case channel attention feature map is input into the variable convolutional spatial attention network to obtain a case spatial attention feature map. The case spatial attention feature map can represent the emphasis of feature map information in the spatial dimension, focus on regions containing important information, and capture more contextual information. The variable convolutional spatial attention network can be a deep neural network for spatial attention constructed by training variable convolution kernels to learn sampling offsets to enhance attention to important regions of multimodal medical case documents. In practice, the execution entity can first input the medical record channel attention feature map into a variable convolution layer with a 3*3 convolution kernel to obtain an offset domain with the same resolution as the medical record channel attention feature map. The offset domain can have a number of channels equal to twice the number of sampling points in the variable convolution layer. The offset domain contains learned offsets, where the number of channels corresponds to the number of sampling points. Each offset consists of a horizontal offset and a vertical offset. That is, the offset at the current pixel position is determined by both the horizontal and vertical offsets. Secondly, the offset in the offset domain is determined and added to the pixel value index in the medical record channel attention feature map to obtain the absolute coordinates of the offset domain. Then, the absolute coordinates of the offset domain are rounded down and up and bilinearly interpolated to obtain the offset coordinates. Finally, the offset coordinates are used to sample the medical record channel attention feature map to obtain offset sampling values. These offset sampling values are weighted summed using the weights corresponding to the variable convolution to obtain the case spatial attention feature map.
[0033] In the fifth step, convolution mapping is performed on the above-mentioned case channel attention feature map to obtain a first channel mapping feature map, a second channel mapping feature map, and a third channel mapping feature map. The first channel mapping feature map, the second channel mapping feature map, and the third channel mapping feature map can be feature maps obtained by using a convolution layer with a 3*3 convolution kernel and expansion rates of 1, 3, and 4 to capture dependencies at different scales and learn more nonlinear features.
[0034] In the sixth step, multi-scale fusion is performed on the case spatial attention feature map, the first channel mapping feature map, the second channel mapping feature map, and the third channel mapping feature map to obtain a multi-scale fused medical record feature map as the sixth case document feature map. The sixth case document feature map can be obtained by concatenating the medical record spatial attention feature map by element-wise multiplication with the first channel mapping feature map, the second channel mapping feature map, and the third channel mapping feature map, respectively.
[0035] Step 102: construct a regional logical hierarchy tree for the case region information set to obtain a case region logical hierarchy tree.
[0036] In some embodiments, the execution entity may construct a regional logical hierarchy tree for the case region information set to obtain a case region logical hierarchy tree. The case region logical hierarchy tree may be a tree structure in the form of a logical tree that displays the logical layout relationship between the case region information included in the case region information set.
[0037] As an example, the execution entity may first select at least one case text information set with a text block element from the case region information set as the case text block region information set. The text blocks may include titles and paragraphs. Then, using an Optical Character Recognition (OCR) model, regular expressions, and Natural Language Processing (NLP), text style extraction is performed on the case text block region information set to obtain a text block style feature information set. The text block style feature information may include, but is not limited to, at least one of the following: text semantics, text font style, and text encoding model. Finally, using a logical structure analysis engine, a logical tree is constructed for the case text block region information set based on the text block style feature information to obtain a logical hierarchy tree for the medical record region. The logical structure analysis rule engine may include, but is not limited to, at least one of the following analysis rules: a depth calculation rule based on title numbering, a hierarchical division rule based on font size, and a priority-based merging rule. The title number-based depth calculation rule may be a rule where the depth of a text block beginning with the number format "." equals the number of "."s in the number format + 1. The above-mentioned hierarchical division rule based on font size can be a rule based on a preset font size classification table, where the level of the text block = the level of the interval in the preset font size classification table corresponding to the font of the text block. The above-mentioned preset font size classification table can be a pre-set table that determines the level of the text block by font size. For example, a mapping relationship in the preset font size classification table can be: the level of the font size greater than or equal to 14 pounds and less than 18 pounds is 1. The above-mentioned priority-based merging rule can be a rule that the depth corresponding to the numbering format is greater than or equal to 0, and the final depth of the text block is equal to the depth corresponding to the numbering format; otherwise, the final depth of the text block is equal to the level corresponding to the font size.
[0038] Step 103 : performing case segmentation processing on the multimodal medical case document according to the case region logical hierarchical tree and the case region information set to obtain a medical segmentation case set.
[0039] In some embodiments, the multiple execution entities may perform case segmentation processing on the multimodal medical case document based on the case region logical hierarchy tree and the case region information set to obtain a medical segmentation case set. The medical segmentation cases in the medical segmentation case set may be cases of the same patient.
[0040] In some optional implementations of some embodiments, performing case segmentation processing on the multimodal medical case document according to the case region logical hierarchy tree and the case region information set to obtain a medical segmentation case set may include the following steps: The first step is to filter out at least one case area information with a category label as a title from the above case area information set.
[0041] In the second step, expression matching is performed on the at least one case region information and a preset case segmentation regular expression set to obtain an expression matching result set. The preset medical record segmentation regular expressions in the preset medical record segmentation regular expression set may be regular expressions for matching preset word segments in the medical record region. For example, the preset case segmentation regular expression set may include: case\d+, patient[AZ].
[0042] The third step is to filter out multiple case region information whose expression matching results represent successful matches from the at least one case region information as the title case region information set.
[0043] The fourth step is to sort the above-mentioned title case area information set by title area according to the above-mentioned case area logical hierarchy tree to obtain a title case area information sequence.
[0044] As an example, the execution entity may determine the sorting order information of the title case area information set in the medical record area logical hierarchy tree, and then sort the title medical record area information set according to the sorting order information to obtain a title case area information sequence.
[0045] In the fifth step, the above-mentioned title case region information sequence and case segmentation prompt word information are input into the case segmentation large language model to obtain a case segmentation location information set. The case segmentation location information in the above-mentioned case segmentation location information set may be information about the location where the text of the multimodal medical case document is segmented. The above-mentioned case segmentation prompt word information may be prompt word information that guides the large prediction model to achieve case segmentation. The above-mentioned case segmentation prompt word information may be prompt word information generated by the large language model. The above-mentioned case segmentation prompt word information may be "Does the following text describe the same patient? [paragraph 1][paragraph 2]". The above-mentioned case segmentation large language model may be a large prediction model that inputs the above-mentioned title case region information sequence and case segmentation prompt word information to determine whether the title case region information sequence is a case of the same patient. For example, the above-mentioned case segmentation large language model may be an LLM (Large Language Model) model.
[0046] In the sixth step, based on the case segmentation location information set, the multimodal medical case document is subjected to case segmentation processing to obtain a medical segmentation case set.
[0047] Step 104: Generate case information extraction prompt word information for the medical segmentation case set.
[0048] In some embodiments, the above-mentioned execution entity can generate case information extraction prompt word information for the above-mentioned medical segmentation case set. Among them, the above-mentioned case information extraction prompt word information can be text information used to guide the large language model to output in a predetermined format and perform text extraction. For example, the above-mentioned case segmentation prompt word information can be: Role: Senior AI assistant in the medical field; Task: Extract structured diagnosis and treatment data from case text; Extraction steps: 1. Clinical examination information: - [Clinical manifestations]: Chief complaint (reason for medical treatment), past history, current medical history (symptom characteristics / accompanying characteristics / treatment received and effect); - [Auxiliary examinations]: Laboratory tests (blood routine / urinalysis / biochemistry / etiology), imaging (X-ray / ultrasound / CT / MRI (key findings), special examinations (electrocardiogram / endoscopy / biopsy results); 2. Diagnosis Diagnosis and classification: - Classification of systemic diseases (such as respiratory / circulatory / digestive system, etc.), - Classification of etiology (infection / trauma / metabolism / tumor, etc.), - Grading of severity (mild / moderate / severe / emergency), - Differential diagnosis and basis, - Complication labeling; 3. Treatment plan: - Drug treatment (drug name, dosage, frequency, course of treatment), - Non-drug treatment (surgery / physical therapy / lifestyle intervention), - Referral indications; 4. Treatment effect: - Follow-up plan (monitoring indicators / time nodes), - Efficacy evaluation (degree of symptom relief / functional recovery / complications).
[0049] Step 105 : Using a multimodal large language model, prompt word information is extracted based on the case information, and structured information extraction processing is performed on the medical segmented case set to obtain a case text structured information set.
[0050] In some embodiments, the execution entity may perform structured information extraction on the medically segmented case set to obtain a case text structured information set. The case text structured information in the case text structured information set may include patient physical condition information extracted from basic patient information and various physical indicator information, presented in JSON format. The various physical indicator information may include, but is not limited to, at least one of the following: medical imaging feature description information, disease type determination information during disease diagnosis, medication information for treatment plans, and treatment plan information. The multimodal large language model may be a large language model that extracts prompt word information from the input case information and performs text extraction from data in different modalities from the medically segmented case set. For example, the multimodal large language model may be, but is not limited to, at least one of the following: a DeepSeek model, ChatGPT (Chat Generative Pre-trained Transformer), or a NExT-GPT model.
[0051] Step 106 : performing position-related category identification on the medical segmentation case set based on the case region information set to obtain a medical image category label information set.
[0052] In some embodiments, the execution entity may perform location-associated category identification on the medical segmentation case set based on the case region information set to obtain a medical image category label information set. The medical image category label information in the medical image category label information set may represent information about the category of the medical image of the case to which each medical image in the medical image segmentation case belongs. The categories may include, but are not limited to, at least one of the following: X-ray images, ultrasound images, and CT (Computed Tomography) / MRI (Magnetic Resonance Imaging) key findings.
[0053] In some optional implementations of some embodiments, performing position-associated category identification on the medical segmented case set based on the case region information set to obtain a medical image category label information set may include the following steps: In the first step, at least one case region information with category labels of image and text blocks is filtered out from the above medical record region information set as an image case region information set and a text case region information set.
[0054] In the second step, coordinate normalization is performed on the image coordinate set corresponding to the image case region information set and the text position set corresponding to the text case region information set, respectively, to obtain a normalized image coordinate set and a normalized text coordinate set. The coordinate normalization may be a normalization process that converts absolute coordinates into normalized coordinates relative to the multimodal medical case document to eliminate size differences between different documents.
[0055] The third step is to determine a set of spatial proximity values based on the standardized image coordinate set and the standardized text coordinate set. The spatial proximity values in the set of spatial proximity values may represent the degree of spatial proximity between any image case region information and any text case region information in the multimodal medical case document.
[0056] As an example, the execution subject of the above number can first determine the absolute value of the difference between the vertical coordinate of each image center point in the image center point vertical coordinate set corresponding to the above-mentioned standardized image coordinate set and the vertical center vertical coordinate of each text in the text vertical center vertical coordinate set corresponding to the above-mentioned standardized text coordinate set as the vertical proximity, thereby obtaining a vertical proximity group set. Then, the difference between the horizontal coordinate of each image center point in the image center point horizontal coordinate set corresponding to the standardized image coordinate set and the horizontal coordinate of each text in the text horizontal coordinate set corresponding to the above-mentioned standardized text coordinate set is used as the horizontal proximity, thereby obtaining a horizontal proximity group set. Finally, the weighted sum of each vertical proximity in the above-mentioned vertical proximity group set and the corresponding horizontal proximity in the above-mentioned horizontal proximity group set is performed to obtain a spatial proximity value group set.
[0057] The fourth step is to construct an image-text position relationship graph based on the above-mentioned spatial proximity value set. The image-text position relationship graph can be a weighted graph with the above-mentioned image case region information set and the above-mentioned text case region information set as nodes, and spatial proximity values and semantic similarity as the weight values of the connecting edges. The above-mentioned semantic similarity can be the similarity obtained by calculating the cosine similarity between the image case region information and the above-mentioned text case region information using CLIP (Contrastive Language-Image Pre-Training).
[0058] As an example, the above-mentioned execution subject can first perform text semantic extraction on the above-mentioned text case area information set to obtain a medical record text semantic feature vector. Secondly, perform image global visual extraction on the above-mentioned image case area information set to obtain an image feature vector set. Then, through CLIP, determine and calculate the cosine similarity of the image case area information set and the above-mentioned text case area information set as the image text semantic similarity to obtain an image text semantic similarity group set. Afterwards, determine the weighted sum of each image text semantic similarity in the above-mentioned image text semantic similarity group set and the spatial proximity value corresponding to the above-mentioned spatial proximity value group set to obtain a connection edge weight group set. Finally, construct a weighted graph with the above-mentioned image case area information set and the above-mentioned text case area information set as nodes and the connection edge weight group set as the weight value of the connection edge to obtain an image text position relationship graph.
[0059] Step 5: Perform association probability prediction on the image-text position relationship graph to obtain an image-text association probability set. The image-text association probability in the image-text association probability set can represent the degree of association between any image case region information and any text case region information. The association probability prediction can be performed using a convolutional neural network.
[0060] In step 6, based on the image-text association probability group set, the medical segmentation case to which each image case region information in the image case region information set belongs is determined as the target medical segmentation case, thereby obtaining a target medical segmentation case set. The target medical segmentation case may be a case consisting of the image case region information and the text case region information with the strongest degree of association.
[0061] As an example, the execution subject can be to select the medical segmentation case to which the text case region information corresponding to the maximum image-text association probability of each image case region information corresponds from the image-text association probability group set, and use the medical segmentation case to which the image case region information belongs as the target medical segmentation case to obtain the target medical segmentation case set. The seventh step is to perform image category recognition on each medical image included in each target medical segmentation case in the above-mentioned target medical segmentation case set to obtain a medical image category label information set. The above-mentioned image category recognition can be a classification recognition of the image category of the above-mentioned medical image using a medical image classification network. The above-mentioned medical image classification network can be a model for classifying the medical image corresponding to the input image case region information. The above-mentioned medical image classification network can be but is not limited to at least one of the following: a three-dimensional convolutional neural network, a Vision Transformer. For at least one image case region information whose confidence output by the medical image classification network is less than a preset confidence, the text case region information with the strongest correlation is input into the multimodal large model for auxiliary reclassification to obtain a medical image category label information set.
[0062] Optionally, the determining of the spatial proximity value set according to the standardized image coordinate set and the standardized text coordinate set may include the following steps: In the first step, for each image case region information in the above image case region information set, the following spatial proximity value generation steps are performed: Sub-step 1, determining the regional vertical distance group of the above-mentioned image case region information and the above-mentioned standardized text coordinate set. The regional vertical distance in the above-mentioned regional vertical distance group can characterize the relative position relationship between the image case region information and the text case region information in the reading flow, and the smaller the value, the closer the regional position. In practice, the absolute value of the vertical coordinate of the standardized image coordinate corresponding to the above-mentioned image case region information and the vertical coordinate of each standardized text coordinate in the above-mentioned standardized text coordinate set are determined as the regional vertical distance to obtain the regional vertical distance group.
[0063] Sub-step 2, determining a group of regional horizontal overlap values based on the above-mentioned image case area information, the above-mentioned standardized text coordinate set, and the above-mentioned text case area information set. The regional horizontal overlap values in the above-mentioned group of regional horizontal overlap values can characterize the overlap ratio of the image case area information and the text case area information in the horizontal direction, and the larger the value, the stronger the correlation in the horizontal direction. In practice, the group of regional horizontal overlap values of the above-mentioned image case area information and the above-mentioned text case area information set is determined based on the horizontal coordinates and image area width corresponding to the above-mentioned image case area information, the above-mentioned standardized text coordinate set, and the text area width values corresponding to the above-mentioned text case area information set.
[0064] As an example, the execution entity may use a horizontal overlap function to determine a set of horizontal overlap values for the image case region information and the text case region information set based on the horizontal coordinates and image region width corresponding to the image case region information, the standardized text coordinate set, and the text region width corresponding to the text case region information set. The horizontal overlap function may be expressed as: .
[0065] in, Represents the horizontal overlap function. Indicates the The standardized horizontal axis corresponding to the case area information of each image. Indicates the The standardized image area width corresponding to the image case area information. Indicates the The standardized text area width value corresponding to the text case area information. Indicates the The standardized text area width value corresponding to the text case area information.
[0066] Sub-step 3: Determine the image center coordinates of the case region information in the image. The first preset value may be a pre-set value. For example, the first preset value may be 2. In practice, the image center coordinates are determined as the sum of the horizontal coordinates of the standardized text coordinates corresponding to the case region information in the image, the corresponding image width, and the ratio of the first preset value.
[0067] Sub-step 4: Determine a text center point coordinate group for the standardized text coordinate set. The text center point coordinates in the text center point coordinate group can represent the position coordinates of the center point of the standardized text coordinates. In practice, the sum of the horizontal coordinate of each standardized text coordinate in the standardized text coordinate set, the corresponding text width, and the ratio of the first preset value is determined as the text center point coordinate group.
[0068] Sub-step 5: Determine a set of horizontal distances between the image center coordinates and the text center coordinates. In practice, the absolute value of the difference between the image center coordinates and the coordinates of each text center in the text center coordinates set is determined as the image-text center horizontal distance, thereby obtaining the set of image-text center horizontal distances.
[0069] Sub-step 6: Determine the difference between the second preset value and the horizontal distance between each image and text center point in the above-mentioned image and text center point horizontal distance group as a center alignment degree group. The center alignment degree in the above-mentioned center alignment degree group can represent the degree of horizontal alignment between the center points of the image case area information and the text case area information. The above-mentioned second preset value can be a pre-set value. For example, the above-mentioned second preset value can be 1.
[0070] Sub-step 7: generating a spatial proximity value group based on the above-mentioned region vertical distance group, the above-mentioned region horizontal overlap value group and the above-mentioned center alignment value group.
[0071] As an example, the execution entity may use a spatial proximity measurement function to generate a spatial proximity value group of the image case region information and the text case region information set based on the region vertical distance group, the region horizontal overlap value group, and the center alignment group. The spatial proximity measurement function may be: .
[0072] in, represents the spatial proximity metric function. Indicates the vertical distance of the area. Indicates the horizontal overlap value of the region, that is, the function value of the horizontal overlap function corresponding to this scene. Indicates center alignment.
[0073] In the process of adopting technical solutions to solve the above-mentioned technical problem 1, the following technical problem 2 is often accompanied: there is an inaccurate matching problem between the image case region information and the text case region information in the medical image position category recognition, which leads to inaccurate segmentation when segmenting the image case region information into individual medical segmentation cases, inaccurate category recognition of medical images, and further leads to inaccurate multimodal information association, storage of a large number of erroneous multimodal synthetic medical cases, and waste of storage resources. In response to the above-mentioned technical problem 2, the conventional solution is generally to use the Hungarian algorithm to determine the medical segmentation case to which each image case region information in the above-mentioned image case region information set belongs based on the above-mentioned image-text association probability group set, as the target medical segmentation case, to obtain the target medical segmentation case set. However, the conventional solutions mentioned above still have the following problems: Since the traditional Hungarian algorithm can only handle one-to-one matching problems, when there are one-to-many or many-to-many situations, random discarding or rule screening is usually used, which has a certain degree of objectivity and cannot be updated in a timely manner, resulting in low matching accuracy between image case area information and text case area information, inaccurate segmentation when segmenting image case area information into individual medical segmented cases, inaccurate category recognition of medical images, and inaccurate association of multimodal information, storing a large number of erroneous multimodal synthetic medical cases, wasting storage resources, reducing the quality of the multimodal case database, and extending the time it takes to build a multimodal case database. Taking into account the shortcomings of the conventional solutions mentioned above, and combining the advantages / technical status of our company's image-text multi-target matching technology, we decided to adopt the following solution: In some optional implementations of some embodiments, determining the medical segmentation case to which each image case region information in the image case region information set belongs based on the image-text association probability group set as the target medical segmentation case to obtain the target medical segmentation case set may include the following steps: The first step is to determine the document page cost value for each image-text association probability in the image-text association probability set, thereby obtaining a document page cost value set. The document page cost values in the document page cost value set can represent the absolute value of the page difference between the image case region information and the text case region information corresponding to the image-text association probability in the multimodal medical case document. A smaller value indicates closer spatial locations and a stronger association.
[0074] The second step is to generate an initial image-text association cost matrix based on the aforementioned image-text association probability set and the aforementioned document page cost value set. Each image-text association cost value in the initial image-text association cost matrix represents the association cost between the image case region information and the text case region information, and is used to measure the matching cost of whether the image case region information and the text case region information belong to the same case. A smaller image-text value indicates a greater likelihood of association between the image case region information and the text case region information.
[0075] As an example, the execution entity may first determine the first weight value and the second weight value for the image-text association probability set and the document page cost value set. The first weight value may be the importance of each image-text association probability in the image-text association probability set. The second weight value may be the importance of each document page cost value in the document page cost value set. The determination may be made by a heuristic algorithm. For example, the heuristic algorithm may include but is not limited to at least one of the following: genetic algorithm, simulated annealing algorithm, whale algorithm, particle swarm algorithm. Then, the image-text association probability set, the document page cost value set, the first weight value and the second weight value are weightedly summed to obtain the initial matrix of image-text association cost.
[0076] The third step is to perform a square matrix transformation on the above-mentioned image-text association cost initial matrix to obtain an image-text association cost matrix as an image-text bipartite graph. The above-mentioned image-text association cost matrix can be a cost matrix formed by expanding the above-mentioned image-text association cost initial matrix to form the same number of image case region information and text case region information. The above-mentioned image-text bipartite graph can be an undirected graph including two mutually non-intersecting image case region information sets and text case region information sets, and the two vertices of each edge belong to the above-mentioned image case region information set and text case region information set respectively. The above-mentioned image-text association cost matrix can be a representation of the above-mentioned image-text bipartite graph. In practice, the above-mentioned execution subject can, in response to determining that the number of images of the image case region information is greater than the number of texts of the text case region information, add the number of images minus the number of texts of the text case region information, and set the image-text association cost initial matrix corresponding to the added text case region information to 100 to obtain the image-text association cost matrix. In response to determining that the number of images of the image case area information is less than the number of texts in the text case area information, the number of texts minus the number of images of the text case area information is added, and the initial image-text association cost matrix corresponding to the added image case area information is set to 100 to obtain the image-text association cost matrix.
[0077] In the fourth step, at least one unmatched image case region information and at least one unmatched text case region information are determined based on the image-text bipartite graph as an unmatched node set. The unmatched nodes in the unmatched node set may be image nodes whose image case region information does not match the text case region information, or text nodes whose text case region information does not match the image case region information.
[0078] As an example, the execution entity may first filter out the highest image-text association cost value in the form of each row and each column from the image-text association cost matrix corresponding to the image-text bipartite graph to obtain an initial image-text matching set. The initial image-text matching set may be a set including matching nodes and matching edges. In practice, the execution entity may utilize the KM (Kuhn Munkres, maximum weight matching) algorithm to filter out the highest image-text association cost value in the form of each row and each column from the image-text association cost matrix to obtain an initial image-text matching set. Then, the initial image-text matching set is traversed to obtain at least one unmatched image case region information and at least one unmatched text case region information as an unmatched node set.
[0079] Step 5: Perform alternating path search and depth-first search on the unmatched node set to obtain a node augmenting path set. A node augmenting path in the node augmenting path set can be a path where both the starting and ending nodes are unmatched nodes, and where matching edges and non-matching edges appear alternately. A matching edge can be an edge where both the starting and ending nodes are matching nodes. A non-matching edge can be an edge where either or both the starting and ending nodes are non-matching nodes.
[0080] As an example, the execution entity may start with an unmatched node in the unmatched node set, search using an alternating search pattern of unmatched edges and matched edges, and perform the search using an improved depth-first search method that prioritizes higher image-text association costs until every unmatched node in the unmatched node set has been searched, thereby obtaining a node augmenting path set. The improved depth-first search method may be a depth-first search method optimized using the Hopcroft-Karp algorithm.
[0081] In the sixth step, the initial image-text matching set is iteratively optimized based on the node augmenting path set to obtain a target image-text matching graph. The target image-text matching graph may be an image-text matching graph obtained by adding the node augmenting path set to the initial image-text matching set.
[0082] In the seventh step, in response to determining that the target image-text matching graph contains image case region information that matches multiple text case region information, priority sorting is performed based on the image-text association probability group set, the spatial proximity value group set, and the image-text semantic relevance to obtain an image-text sorting matching graph. The image-text sorting matching graph can be a one-to-one matching graph between image case region information and text case region information. The image-text semantic relevance can be inputting the image case region information and the text case region information into the CLIP model to determine the semantic similarity between the image case region information and the text case region information.
[0083] As an example, the above-mentioned execution subject can, in response to determining that there are image case area information matching multiple text case area information in the above-mentioned target image text matching graph, first, screen out the text case area information corresponding to the image-text association probability with the largest numerical value from the multiple text case area information as the most matching text case area information. Then, in response to determining that there are also multiple text case area information corresponding to the image-text association probability with the largest numerical value, select the text case area information with the largest spatial proximity numerical value as the most matching text case area information. Finally, in response to determining that there are also multiple maximum spatial proximity numerical values, use the CLIP model to select the text case area information with the greatest semantic similarity to the image case area information as the most matching text case area information, until all conflicts between the image case area information matching multiple text case area information are resolved, and an image-text sorting matching graph is obtained.
[0084] In the eighth step, based on the above-mentioned image-text sorting matching graph, the medical segmentation case to which each image case region information in the above-mentioned image case region information set belongs is determined as the target medical segmentation case, and the target medical segmentation case set is obtained.
[0085] As an example, the above-mentioned execution entity can query the image text sorting matching graph to determine the medical segmentation case to which each image case area information in the above-mentioned image case area information set belongs, as the target medical segmentation case, and obtain the target medical segmentation case set.
[0086] The above technical solution, combined with step "step 108" and its related content as an inventive point of an embodiment of the present disclosure, solves the second technical problem mentioned in the background technology: "Since the traditional Hungarian algorithm can only handle one-to-one matching problems, when there are one-to-many or many-to-many situations, random discarding or rule screening is usually adopted, which has a certain objectivity and cannot be updated in time, resulting in a low matching accuracy of image case area information and text case area information, inaccurate segmentation when segmenting image case area information into various medical segmented cases, inaccurate category recognition of medical images, and inaccurate association of multimodal information, storing a large number of erroneous multimodal synthetic medical cases, wasting storage resources, reducing the quality of the multimodal case database, and extending the time to build a multimodal case database." The results are low matching accuracy, inaccurate segmentation and inaccurate classification of medical images, which in turn leads to waste of storage resources. The following are often the reasons: Since the traditional Hungarian algorithm can only handle one-to-one matching problems, when there are one-to-many or many-to-many situations, random discarding or rule screening is usually adopted, which has a certain objectivity and cannot be updated in time, resulting in low matching accuracy of image case region information and text case region information, inaccurate segmentation when segmenting image case region information into individual medical segmented cases, inaccurate classification of medical images, and inaccurate association of multimodal information, resulting in a large number of erroneous multimodal synthetic medical cases. If the above factors are solved, it is possible to improve the matching accuracy, segmentation accuracy and classification accuracy of medical images, thereby improving the accuracy of multimodal information association, reducing the waste of storage resources, improving the quality of the multimodal case database, and shortening the time to build the multimodal case database. In order to achieve this effect, the present disclosure firstly performs dynamic weighted fusion of document page cost values and image-text association probability based on spatial proximity values and image-text semantic similarity, which can improve the accuracy of the association degree of the initial matrix of image-text association cost. Then, an image-text bipartite graph, an initial image-text matching set, and a set of unmatched nodes are constructed. Alternating path search, an improved depth-first search, and iterative optimization are performed on the unmatched node set. This allows for batch processing of augmenting paths, reducing computational complexity. Subsequently, image-text association probability groups, spatial proximity value groups, and image-text semantic relevance are prioritized to generate an image-text sorted matching graph. This resolves many-to-many and one-to-many matching conflicts, optimizes the matching of image case region information and text case region information, and improves matching results. Finally, using the target medical segmentation case set, individual case-based multimodal information association and case storage are performed on the medical image sets corresponding to the case text structured information set and the medical image category label information set, resulting in a multimodal case database. This improves the quality of the multimodal case database, shortens the construction time, and reduces the waste of storage resources.
[0087] Step 107 , performing multimodal information association on the case text structured information set and the medical image category label information set to obtain a multimodal synthetic medical case information set, and performing case storage on the multimodal synthetic medical case information set to obtain a multimodal case database.
[0088] In some embodiments, the execution subject may perform multimodal information association on the case text structured information set and the medical image category label information set to obtain a multimodal synthetic medical case information set, and perform case storage on the multimodal synthetic medical case information set to obtain a multimodal case database. The multimodal synthetic medical case information in the multimodal synthetic medical case information set may be information that the case text structured information and the medical images corresponding to the medical image category label information belong to the same case of the same patient. The multimodal case database may be a database for associating and storing the case text structured information set and the medical image set corresponding to the medical image category label information set. The multimodal information association may be a multimodal information association based on individual cases. Figure 2 As shown, the entire process of processing multimodal medical case documents from step 101 to step 108 is shown. Figure 2 The left figure in the figure can show a flow chart of document layout area detection and regional logical hierarchical tree construction for multimodal medical case documents, the middle figure can show a flow chart of information extraction using prompt words and multimodal large language models, and the right figure can show a flow chart of medical image category recognition and multimodal association integration of text.
[0089] As an example, the execution entity may first determine, using the medical image category label information set, the medical segmentation case to which each medical image in the medical image set corresponding to the medical image category label information set belongs. Then, using the medical segmentation cases, the case text structured information set and the medical images are associated and integrated to obtain a multimodal synthetic medical case information set. Finally, the multimodal synthetic medical case information set is stored as a case to obtain a multimodal case database.
[0090] Further references Figure 3 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a multimodal case database construction device. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the multimodal case database construction device can be specifically applied to various electronic devices.
[0091] like Figure 3As shown, a multimodal case database construction device 300 includes: a document layout area detection unit 301, a region logic hierarchy tree construction unit 302, a case segmentation unit 303, a generation unit 304, a case information extraction unit 305, a position association category identification unit 306, and a multimodal information association unit 307. The document layout area detection unit 301 performs document layout area detection processing on the acquired multimodal medical case document to obtain a case region information set. The region logic hierarchy tree construction unit 302 performs a region logic hierarchy tree construction on the above case region information set to obtain a case region logic hierarchy tree. The case segmentation unit 303 performs case segmentation processing on the above multimodal medical case document based on the above case region logic hierarchy tree and the above case region information set to obtain a medical segmentation case set. The generation unit 304 generates case information extraction prompt word information for the above medical segmentation case set. The case information extraction unit 305 uses a multimodal large language model to extract prompt word information based on the case information and performs structured information extraction processing on the medical segmented case set to obtain a case text structured information set. The position association category identification unit 306 performs position association category identification on the medical segmented case set based on the case region information set to obtain a medical image category label information set. The multimodal information association unit 307 performs multimodal information association on the case text structured information set and the medical image category label information set to obtain a multimodal synthesized medical case information set, and then stores the multimodal synthesized medical case information set as a case to obtain a multimodal case database.
[0092] It is understandable that the units recorded in the multimodal case database construction device 300 and the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the multimodal case database construction device 300 and the units included therein, and will not be repeated here.
[0093] Reference below Figure 4 , which shows a structural schematic diagram of an electronic device (eg, an electronic device) 400 suitable for implementing some embodiments of the present disclosure. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0094] like Figure 4As shown, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 402 or programs loaded from a storage device 408 into a random access memory (RAM) 403. Various programs and data required for the operation of electronic device 400 are also stored in RAM 403. Processing device 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to bus 404.
[0095] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data. Figure 4 The electronic device 400 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 4 Each block shown in the figure may represent one device, or may represent multiple devices as needed.
[0096] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.
[0097] It should be noted that in some embodiments of the present disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. Furthermore, in some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.
[0098] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0099] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: performs document layout area detection processing on the acquired multimodal medical case document to obtain a case area information set; constructs a region logic hierarchy tree on the case area information set to obtain a case area logic hierarchy tree; performs case segmentation processing on the multimodal medical case document based on the case area logic hierarchy tree and the case area information set to obtain a medical segmentation case set; generates case information extraction prompt word information for the medical segmentation case set; uses a multimodal large language model to extract prompt word information based on the case information, performs structured information extraction processing on the medical segmentation case set to obtain a case text structured information set; performs position association category identification on the medical segmentation case set based on the case area information set to obtain a medical image category label information set; performs multimodal information association on the case text structured information set and the medical image category label information set to obtain a multimodal synthetic medical case information set; and stores cases on the multimodal synthetic medical case information set to obtain a multimodal case database.
[0100] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0102] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor, for example, they may be described as: a processor including a document layout area detection unit, a region logic hierarchy tree construction unit, a case segmentation unit, a generation unit, a case information extraction unit, a position association category identification unit, and a multimodal information association unit. Among them, the names of these units do not constitute a limitation on the units themselves under certain circumstances. For example, the document layout area detection unit may also be described as "a unit that performs document layout area detection processing on the acquired multimodal medical case documents to obtain a case area information set."
[0103] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0104] The above descriptions are merely some preferred embodiments of the present disclosure and illustrate the underlying technical principles. Those skilled in the art should understand that the scope of the invention encompassed by the embodiments of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the aforementioned inventive concept. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A method for constructing a multimodal case database, comprising: Perform document layout area detection on the acquired multimodal medical case documents to obtain a case area information set; Constructing a regional logical hierarchical tree for the case region information set to obtain a case region logical hierarchical tree; Performing case segmentation processing on the multimodal medical case document according to the case region logical hierarchical tree and the case region information set to obtain a medical segmentation case set; Generating case information extraction prompt word information for the medical segmentation case set; Utilizing a multimodal large language model, extracting prompt word information based on the case information, performing structured information extraction processing on the medical segmented case set, and obtaining a case text structured information set; According to the case region information set, position-associated category identification is performed on the medical segmentation case set to obtain a medical image category label information set; Multimodal information association is performed on the case text structured information set and the medical image category label information set to obtain a multimodal synthetic medical case information set, and case storage is performed on the multimodal synthetic medical case information set to obtain a multimodal case database.
2. The method according to claim 1, wherein The document layout area detection process is performed on the acquired multimodal medical case document to obtain a case area information set, including: Inputting the multimodal medical case document into an activated convolutional feature extraction network included in a document layout region detection model to obtain a first case document feature map, wherein the document layout region detection model further includes: a residual activation feature extraction backbone network, a cross-stage fusion network, a multi-scale visual feature enhancement network, and multiple multi-scale cross feature fusion networks; Inputting the first case document feature map into the residual activation feature extraction backbone network to obtain a second case document feature map and a third case document feature map; Inputting the second case document feature map and the third case document feature map into the cross-stage fusion network to obtain a fourth case document feature map and a fifth case document feature map; Inputting the first case document feature map into the multi-scale visual feature enhancement network to obtain a sixth case document feature map; The fifth case document feature map and the sixth case document feature map are input into the multiple multi-scale cross feature fusion networks to obtain a case region information set.
3. The method according to claim 2, wherein: The multi-scale visual feature enhancement network includes: a global pooling channel attention network and a variable convolutional spatial attention network; and The step of inputting the first case document feature map into the multi-scale visual feature enhancement network to obtain a sixth case document feature map comprises: Inputting the first case document feature map into the global average pooling layer and the global maximum pooling layer included in the global pooling channel attention network to obtain the case channel average pooling feature map and the case channel maximum pooling feature map; Performing channel nonlinear fusion on the case channel average pooling feature map and the case channel maximum pooling feature map to obtain a case average channel nonlinear feature map and a case maximum channel nonlinear feature map; Performing a weighted summation on the case average channel nonlinear feature map, the case maximum channel nonlinear feature map, and the first case document feature map to obtain a case channel attention feature map; Inputting the case channel attention feature map into the variable convolutional spatial attention network to obtain a case spatial attention feature map; Performing convolution mapping on the case channel attention feature map to obtain a first channel mapping feature map, a second channel mapping feature map, and a third channel mapping feature map; Multi-scale fusion is performed on the case spatial attention feature map, the first channel mapping feature map, the second channel mapping feature map and the third channel mapping feature map to obtain a multi-scale fused medical record feature map as the sixth case document feature map.
4. The method according to claim 1, wherein The step of performing case segmentation processing on the multimodal medical case document according to the case region logical hierarchical tree and the case region information set to obtain a medical segmentation case set includes: Filtering at least one case area information whose category label is a title from the case area information set; Perform expression matching on the at least one case region information and a preset case segmentation regular expression set to obtain an expression matching result set; Filtering out multiple pieces of case region information whose expression matching results indicate successful matching from the at least one case region information as a title case region information set; Sorting the title case area information set according to the case area logical hierarchy tree to obtain a title case area information sequence; Inputting the title case region information sequence and the case segmentation prompt word information into the case segmentation large language model to obtain a case segmentation position information set; According to the case segmentation position information set, case segmentation processing is performed on the multimodal medical case document to obtain a medical segmentation case set.
5. The method according to claim 1, wherein The step of performing position-associated category identification on the medical segmentation case set based on the case region information set to obtain a medical image category label information set includes: Filtering at least one case area information with a category label of image and text block from the case area information set as an image case area information set and a text case area information set; Performing coordinate normalization processing on the image coordinate set corresponding to the image case region information set and the text position set corresponding to the text case region information set, respectively, to obtain a normalized image coordinate set and a normalized text coordinate set; Determining a set of spatial proximity value values according to the standardized image coordinate set and the standardized text coordinate set; constructing an image-text position relationship graph based on the set of spatial proximity value values; Performing association probability prediction on the image-text position relationship graph to obtain an image-text association probability group set; According to the image-text association probability group set, determining the medical segmentation case to which each image case region information in the image case region information set belongs as a target medical segmentation case, thereby obtaining a target medical segmentation case set; Image category recognition is performed on each medical image included in each target medical segmentation case in the target medical segmentation case set to obtain a medical image category label information set.
6. The method according to claim 5, wherein: Determining a set of spatial proximity value values according to the normalized image coordinate set and the normalized text coordinate set includes: For each image case region information in the image case region information set, the following spatial proximity value generation steps are performed: Determine a regional vertical distance group between the image case region information and the standardized text coordinate set; Determining a region horizontal overlap value group according to the image case region information, the standardized text coordinate set, and the text case region information set; Determine the image center point coordinates of the image case area information; Determining a text center point coordinate group of the standardized text coordinate set; Determine a horizontal distance group between the image center point coordinates and the text center point coordinate group; Determine a difference between a second preset value and the horizontal distance between each image text center point in the image text center point horizontal distance group as a center alignment degree group; A spatial proximity value group is generated according to the region vertical distance group, the region horizontal overlap value group, and the center alignment group.
7. A multimodal case database construction device, comprising: A document layout region detection unit performs document layout region detection processing on the acquired multimodal medical case document to obtain a case region information set; A regional logic hierarchy tree construction unit constructs a regional logic hierarchy tree for the case region information set to obtain a case region logic hierarchy tree; a case segmentation unit, performing case segmentation processing on the multimodal medical case document according to the case region logical hierarchical tree and the case region information set to obtain a medical segmentation case set; A generating unit, generating case information extraction prompt word information for the medical segmented case set; A case information extraction unit, using a multimodal large language model to extract prompt word information according to the case information, performs structured information extraction processing on the medical segmented case set, and obtains a case text structured information set; a position-related category identification unit, which performs position-related category identification on the medical segmentation case set based on the case region information set to obtain a medical image category label information set; The multimodal information association unit performs multimodal information association on the case text structured information set and the medical image category label information set to obtain a multimodal synthetic medical case information set, and performs case storage on the multimodal synthetic medical case information set to obtain a multimodal case database.
8. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
A superpixel method for medical image segmentation
CN109035252A
Literature key information extraction method and device, computer equipment and storage medium
CN113673294A
Automatic cataloguing method, system and equipment for cases and storage medium
CN115880704A
Medical image segmentation and labeling method and system based on multi-modal information fusion
CN119251490A
Medical history information tracking method and device, computer equipment and readable storage medium
CN119314195A
Cited By
Multi-modal medical image segmentation method and device, computer equipment and storage medium
CN121837614A