Multimodal case database construction method and apparatus, electronic device, and medium
By performing document layout region detection and logical hierarchy tree construction on multimodal medical case documents, and combining multimodal large language models for case segmentation and information extraction, the problems of low efficiency and low accuracy in multimodal case database construction are solved, achieving efficient intelligent organization and structured storage.
Patent Information
- Application Number
- CN202511157903.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Existing technologies for constructing multimodal case databases suffer from several drawbacks. The complex and diverse formats of case data lead to low construction efficiency and long construction time. Furthermore, regular expressions cannot understand the deep semantic logic and ambiguous expressions of the text, resulting in low data extraction accuracy and wasted resources.
By performing document layout region detection on multimodal medical case documents, constructing a region logical hierarchy tree, using a multimodal large language model for case segmentation and information extraction, generating case information extraction prompt words, and performing multimodal information association and storage, intelligent organization and structured storage are achieved.
It improves the efficiency of building multimodal case databases, reduces storage resource waste, improves the accuracy of data extraction and database quality, and adapts to different formats and updated medical documents.
Smart Images

Figure CN120656630B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of computer technology, and in particular, to a multi-modal case database construction method and device, electronic equipment and medium. BACKGROUND
[0002] With the rapid development of medical technology, the complexity and data volume of medical data are increasing, and the existence of various modal data in medical cases and the complex and diverse formats of medical cases are important factors affecting the construction of medical databases. For the construction of a multi-modal case database, the commonly used method is to construct a medical regular expression set for content matching and extraction of medical cases, then use the medical regular expression set to perform content matching and extraction integration on multi-modal medical cases to obtain multi-modal integrated medical cases, and finally store the multi-modal integrated medical cases to obtain a multi-modal case database.
[0003] However, in practice, it is found that when the above method is used to construct a multi-modal case database, the following technical problems often exist: 1. Due to the complex and diverse formats of case data, regular expressions are used to match and extract the content of case data, and specific regular expressions need to be constructed for each format of case data. When the case data (such as data format, extraction requirements, and medical terminology) changes, the regular expressions need to be re-constructed, resulting in low efficiency, long construction time, and waste of a large amount of resources for case database construction; at the same time, regular expressions only match and extract data based on the literal meaning of the text, and cannot understand the deep semantic logic and fuzzy expression of the text, and the efficiency and accuracy of data extraction and matching of other modal data in addition to text are low, there are problems of missing information extraction and incorrect extraction, resulting in low accuracy of case data content extraction, a large amount of errors and redundant information, and waste of data storage resources and low database quality.
[0004] The above information disclosed in this BACKGROUND section is only for the purpose of enhancing the understanding of the background of the present disclosure and, therefore, can include information that does not form the prior art that is already known to those of ordinary skill in the art. SUMMARY
[0005] The summary section is provided to introduce concepts briefly in a simplified form, which will be described in detail in the specific embodiments section. The summary section is not intended to identify key or essential features of the claimed technology nor is it intended to be used to limit the scope of the claimed technology.
[0006] Some embodiments of the present disclosure propose a multi-modal case database construction method, device, electronic equipment and medium to solve one or more of the technical problems mentioned in the background section.
[0007] In a first aspect, some embodiments of the present disclosure provide a multi-modal case database construction method, comprising: performing document layout region detection processing on an obtained multi-modal medical case document to obtain a case region information set; constructing a region logical hierarchical tree based on the case region information set to obtain a case region logical hierarchical tree; performing case segmentation processing on the multi-modal medical case document based on the case region logical hierarchical tree and the case region information set to obtain a medical segmented case set; generating case information extraction prompt word information for the medical segmented case set; performing structured information extraction processing on the medical segmented case set based on the case information extraction prompt word information by using a multi-modal large language model to obtain a case text structured information set; performing position association category recognition on the medical segmented case set based on the case region information set to obtain a medical image category label information set; performing multi-modal information association on the case text structured information set and the medical image category label information set to obtain a multi-modal synthesized medical case information set; and performing case storage on the multi-modal synthesized medical case information set to obtain a multi-modal case database.
[0008] In a second aspect, some embodiments of the present disclosure provide a multi-modal case database construction apparatus, comprising: a document layout region detection unit configured to perform document layout region detection processing on an obtained multi-modal medical case document to obtain a case region information set; a region logical hierarchical tree construction unit configured to construct a region logical hierarchical tree based on the case region information set to obtain a case region logical hierarchical tree; a case segmentation unit configured to perform case segmentation processing on the multi-modal medical case document based on the case region logical hierarchical tree and the case region information set to obtain a medical segmented case set; a generation unit configured to generate case information extraction prompt word information for the medical segmented case set; a case information extraction unit configured to perform structured information extraction processing on the medical segmented case set based on the case information extraction prompt word information by using a multi-modal large language model to obtain a case text structured information set; a position association category recognition unit configured to perform position association category recognition on the medical segmented case set based on the case region information set to obtain a medical image category label information set; a multi-modal information association unit configured to perform multi-modal information association on the case text structured information set and the medical image category label information set to obtain a multi-modal synthesized medical case information set, and perform case storage on the multi-modal synthesized medical case information set to obtain a multi-modal case database.
[0009] In a third aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.
[0010] In a fourth aspect, some embodiments of the present disclosure provide a computer readable medium having stored thereon a computer program, wherein the computer program, when executed by a processor, implements the method according to any implementation manner of the first aspect.
[0011] The above various embodiments of the present disclosure have the following beneficial effects: the multi-modal case database construction method of some embodiments of the present disclosure can automatically extract and integrate multi-modal medical cases, improve the efficiency of constructing the database and reduce the waste of storage resources. Specifically, the reasons for the low efficiency of related case database construction, long construction time, waste of a large amount of resources, and low quality of the database are as follows: due to the complex and diverse formats of case data, when regular expressions are used for case data matching and content extraction, a specific regular expression needs to be constructed for each format of case data, and when the case data (such as data format, extraction needs, medical terminology) changes, the regular expression needs to be re-constructed, resulting in low efficiency of case database construction, long construction time, and waste of a large amount of resources; at the same time, regular expressions only match data and extract content through text literal meaning, cannot understand the deep semantic logic and fuzzy expression of the text, and have low efficiency and accuracy in extracting and matching data other than text, resulting in missing and incorrect extraction of information, low accuracy of case data content extraction, a large amount of errors and redundant information, and waste of data storage resources and low quality of the database. Based on this, the multi-modal case database construction method of some embodiments of the present disclosure can first perform document layout area detection processing on the obtained multi-modal medical case documents to obtain a case area information set. Here, the layout of multi-modal medical case documents of different formats and different medical fields can be analyzed to accurately grasp the layout of multi-modal medical case documents, without the need to construct and update complex regular expressions. Second, a region logical hierarchy tree is constructed based on the above case area information set to obtain a case region logical hierarchy tree. Here, the construction of the case region logical hierarchy tree can better understand the semantic information, layout information and hierarchical structure between different regions of the multi-modal medical case document, and can adapt to documents of different formats. Then, according to the above case region logical hierarchy tree and the above case area information set, the above multi-modal medical case document is subjected to case segmentation processing to obtain a medical segmented case set. Here, the multi-modal medical case document is segmented into cases of the same patient, which can improve the correlation and segmentation of multi-modal data and avoid a large amount of error data in subsequent associated storage. Subsequently, case information extraction prompt word information is generated for the above medical segmented case set. Here, the information extraction prompt word information is used to guide the subsequent multi-modal large language model to perform information extraction, improve the efficiency and accuracy of the multi-modal large language model. Then, using a multi-modal large language model, structured information extraction processing is performed on the above medical segmented case set according to the above case information extraction prompt word information to obtain a case text structured information set. Here, the multi-modal large language model can recognize the deep semantic logic and fuzzy expression of different model data in the case, improve the efficiency and improve the accuracy.Then, according to the case region information set, the medical segmentation case set is subjected to position correlation category recognition to obtain a medical image category label information set. Here, through medical image position correlation, the recognition segmentation accuracy of the medical segmentation case to which the medical image belongs can be improved, and through the assistance of different modal data, the accuracy of medical image category recognition can be further improved. Finally, the case text structured information set and the medical image category label information set are subjected to multi-modal information correlation to obtain a multi-modal synthetic medical case information set, and the multi-modal synthetic medical case information set is subjected to case storage to obtain a multi-modal case database. Here, the accuracy of multi-modal information correlation and the quality of the multi-modal synthetic medical case information set can be improved, the waste of storage resources can be reduced, and the construction efficiency of the multi-modal case database can be improved, and the construction time can be shortened. Thus, the multi-modal case database construction method can be adapted to medical documents of different formats and formats updated through layout region segmentation and logical level tree construction. Then, through customized prompt word information and a multi-modal large language model, image and text data of different modalities are automatically extracted and correlated and integrated to realize intelligent arrangement and structured storage of multi-modal medical case documents, improve the quality of multi-modal synthetic medical case information, reduce the waste of storage resources, improve the construction efficiency of the multi-modal case database, and shorten the construction time. BRIEF DESCRIPTION OF DRAWINGS
[0012] The above-described and other features and advantages of various embodiments of the present disclosure will be more apparent from the following detailed description taken in conjunction with the accompanying drawings. Throughout the drawings, like or similar reference numerals are used to refer to like or similar elements, and the accompanying text provides detailed description of the drawings. It should be noted that the drawings are in schematic form and the elements and elements are not necessarily drawn to scale.
[0013] Figure 1 is a flowchart of some embodiments of the multi-modal case database construction method according to the present disclosure;
[0014] Figure 2 is a flowchart of multi-modal case database construction in some embodiments of the multi-modal case database construction method according to the present disclosure;
[0015] Figure 3 is a structural schematic diagram of some embodiments of the multi-modal case database construction device according to the present disclosure;
[0016] Figure 4 is a structural schematic diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0017] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be more thoroughly and completely understood. It should be understood that the drawings of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.
[0018] In addition, it should be further noted that only parts related to the present application are shown in the drawings for ease of description. The embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0019] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0020] It should be noted that the adjectives "one", "multiple" mentioned in the present disclosure are illustrative and not limiting, and those skilled in the art should understand that unless otherwise explicitly stated in the context, it should be understood as "one or more".
[0021] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of the messages or information.
[0022] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0023] Figure 1 The flow 100 of some embodiments of the multi-modal case database construction method according to the present disclosure is shown. The multi-modal case database construction method includes the following steps:
[0024] In step 101, the obtained multi-modal medical case document is subjected to document layout region detection processing to obtain a case region information set.
[0025] In some embodiments, the execution subject (e.g., an electronic device) of the above multi-modal case database construction method can perform document layout region detection processing on the obtained multi-modal medical case document to obtain a case region information set. The multi-modal medical case document obtained above can be obtained through a wired connection or a wireless connection. The multi-modal medical case document can be a document in multiple data formats recording the health status of a patient. The multiple data formats can include, but are not limited to, at least one of the following: images, texts, physical sign information, and time series data. The case region information in the case region information set can be the position information and element category information of different regions in the detected multi-modal medical case document. The element category information can include, but is not limited to, at least one of the following: document title, chart, medical image, formula, abstract, reference, and text paragraph.
[0026] As an example, the execution subject can input the multi-modal medical case document into a PP-DocLayout (document picture layout detection) model to perform document layout region detection and obtain the case region information set.
[0027] In some optional implementations of some embodiments, the document layout region detection processing on the obtained multi-modal medical case document to obtain the case region information set can include the following steps:
[0028] In a first step, the multi-modal medical case document is input into an activation convolution feature extraction network included in a document layout region detection model to obtain a first case document feature map, wherein the document layout region detection model further includes: a residual activation feature extraction backbone network, a cross-stage fusion network, a multi-scale visual feature enhancement network, and a plurality of multi-scale cross-feature fusion networks. The first case document feature map can represent the shallow visual information of the multi-modal medical case document. The document layout region detection model can be a neural network model that analyzes and divides the different regions of the input multi-modal medical case document to determine the location information, semantic information, and category information of the different regions. The activation convolution feature extraction network can be a neural network model that extracts visual features from the input multi-modal medical case document. The activation convolution feature extraction network can include: a standard convolution layer, a BN (Batch Normalization) layer, and a Leaky ReLU (Leaky Rectified Linear Unit) activation function. The residual activation feature extraction backbone network can be a deep neural network model that extracts deep visual features from the input first case document feature map. The residual activation feature extraction backbone network can include: an activation convolution feature extraction network, a CSP1_1 network (Cross Stage Partial Connections), an activation convolution feature extraction network, a CSP1_3 network, an activation convolution feature extraction network, a CSP1_3 network, an activation convolution feature extraction network, and an SPP (Spatial Pyramid Pooling) network. The first 1 in the CSP1_1 network indicates that the number of activation convolution feature extraction networks included in the CSP1_1 network is 1, and the second 1 indicates that the number of residual components included in the CSP1_1 network is 1. The cross-stage fusion network can be a deep neural network model that fuses the features of the output of the cross-stage fusion network. The cross-stage fusion network can include: a residual activation feature extraction backbone network, an upsampling module, a feature fusion module, a CSP2_1 network, a residual activation feature extraction backbone network, and an upsampling module. The multi-scale visual feature enhancement network can be a neural network model that uses a convolution attention module (CBAM) to enhance the spatial position information of the input first case document feature map and the output of the cross-stage fusion network, and uses a multi-branch convolution to construct a dependency relationship of different scales to obtain an enhanced shallow visual feature. The plurality of multi-scale cross-feature fusion networks can be multi-scale cross-feature fusion modules that replace the Concat fusion method and use attention mechanisms to adaptively fuse features of different layers.
[0029] Second step, input the first case document feature map above to the residual activation feature extraction backbone network above to obtain a second case document feature map and a third case document feature map. Wherein the second case document feature map above can be a feature map located at the output of the second CSP1_3 network included in the residual activation feature extraction backbone network. The third case document feature map above can be a feature map of the final output of the residual activation feature extraction backbone network.
[0030] Third step, input the second case document feature map above and the third case document feature map above to the cross-stage fusion network above to obtain a fourth case document feature map and a fifth case document feature map. The fourth case document feature map above can be a feature map of the output of the second activation convolution feature extraction network included in the cross-stage fusion network. The fifth case document feature map above can be a feature map of the final output of the cross-stage fusion network.
[0031] Fourth step, input the first case document feature map above to the multi-scale visual feature enhancement network above to obtain a sixth case document feature map.
[0032] Fifth, input the fifth case document feature map and the sixth case document feature map into the plurality of multi-scale cross feature fusion networks to obtain a case region information set. Each multi-scale cross feature fusion network in the plurality of multi-scale cross feature fusion networks can include a standard convolution network, a depth separable convolution network, and a cross attention mechanism layer. The number of the plurality of multi-scale cross feature fusion networks can be the same as the number of case region information included in the case region information set. For example, the plurality of multi-scale cross feature fusion networks can be three multi-scale cross feature fusion networks. In practice, the execution subject can first input the fifth case document feature map and the sixth case document feature map into the standard convolution network and the depth separable convolution network included in the first multi-scale cross feature fusion network in the plurality of multi-scale cross feature fusion networks to obtain a first case document local channel feature map and a second case document local channel feature map. The standard convolution network can be a model for local representation of the input fifth case document feature map and the sixth case document feature map. The depth separable convolution network can be a model for mapping the fifth case document feature map and the sixth case document feature map to a high-dimensional space and enriching channel information. The plurality of multi-scale cross feature fusion networks can be three multi-scale cross feature fusion networks. Secondly, the first case document local channel feature map and the second case document local channel feature map are flattened to obtain a first document local channel flat feature map set and a second document local channel flat feature map set. The flattening can be a non-overlapping flat block flattening for facilitating long dependency matching of the fifth case document feature map and the sixth case document feature map with spatial induction bias. The shapes of the fifth case document feature map and the sixth case document feature map can both be H W C. The shape of the first case document local channel feature map can be H W d, d > C. The shape of the first document local channel flat feature map can be P N d, P = WH, N = HW / P. The above H can represent the height of the feature map, W can represent the width of the feature map, and d can represent the number of channels of the feature map. The above N can represent the number of first document local channel flat feature maps included in the first document local channel flat feature map set. Again, the above first document local channel flat feature map set and the above second document local channel flat feature map set are input into the cross attention mechanism layer to obtain a medical record document cross fusion feature map set. Then, the above medical record document cross fusion feature map set is subjected to convolution splicing to obtain a document multi-scale fusion feature map. In practice, the above execution subject can first perform feature map flattening on the above case document cross fusion feature map to obtain a flattened case document feature map. The shape of the above flattened case document feature map can be H W d. Secondly, the above flattened case document feature map is input into a depth separable convolution network to map to a C-dimensional space to obtain a document mapping feature map. The shape of the above document mapping feature map can be H W C. Then, the above fifth case document feature map, the above sixth case document feature map, and the above document mapping feature map are subjected to feature splicing to obtain a spliced feature map. The shape of the above spliced feature map can be H W 3C. Finally, the spliced feature map is input into a standard convolution layer for feature fusion representation to obtain a document multi-scale fusion feature map. Then, the above third case document feature map, the above fourth case document feature map, and the above document multi-scale fusion feature map are input into a plurality of multi-scale cross feature fusion networks after the above first multi-scale cross feature fusion network to obtain a seventh case document feature map and an eighth case document feature map. Finally, the above seventh case document feature map, the above eighth case document feature map, and the above document multi-scale fusion feature map are subjected to convolution prediction processing to obtain a case region information set. In practice, the above execution subject can first input the above seventh case document feature map, the above eighth case document feature map, and the above document multi-scale fusion feature map into the convolution layer of 3 3 and the convolution layer of 1 1 in sequence, respectively, to obtain a first bounding box set, a second bounding box, and a third bounding box set. Then, the first bounding box set, the second bounding box, and the third bounding box set are subjected to class score sorting and non-maximum suppression screening, respectively, to obtain a case region information set of different sizes.
[0033] In some optional implementations of some embodiments, the multi-scale visual feature enhancement network comprises a global pooling channel attention network and a variability convolution spatial attention network. The global pooling channel attention network can capture information of each channel through a global average pooling layer and a global maximum pooling layer to explicitly model channel interdependence to enhance learning of convolution features and fully consider correlation between channels in a global range.
[0034] Optionally, inputting the first case document feature map into the multi-scale visual feature enhancement network to obtain a sixth case document feature map can comprise the following steps:
[0035] Firstly, inputting the first case document feature map into a global average pooling layer and a global maximum pooling layer included in the global pooling channel attention network to obtain a case channel average pooling feature map and a case channel maximum pooling feature map.
[0036] Secondly, performing channel nonlinear fusion on the case channel average pooling feature map and the case channel maximum pooling feature map to obtain a case average channel nonlinear feature map and a case maximum channel nonlinear feature map. The channel nonlinear fusion can be first adjusted through a 1 1 convolution layer, a RuLU activation function and a 1 1 convolution layer to adjust the feature map dimension and reduce the calculation complexity, and then the Sigmoid function is used to learn the nonlinear interaction relationship between channels, that is, the non-exclusive relationship between channels, to ensure that the information of multiple channels can be emphasized rather than only emphasizing the nonlinear fusion of a single channel.
[0037] Thirdly, performing weighted summation on the case average channel nonlinear feature map, the case maximum channel nonlinear feature map and the first case document feature map to obtain a case channel attention feature map.
[0038] Fourthly, inputting the case channel attention feature map into the variability convolution spatial attention network to obtain a case spatial attention feature map. The case spatial attention feature map can represent emphasis of feature map information in the spatial dimension, attention to important information regions, and acquisition of more context information. The variability convolution spatial attention network can be a spatial attention deep neural network constructed by learning sampling offset to enhance attention to important regions of the multi-modal medical case document through training of a variability convolution kernel. In practice, the execution subject can first input the case channel attention feature map into a 3 The variable convolution layer of step 3 obtains an offset domain with the same resolution as the case channel attention feature map. The offset domain can have a channel number that is twice the number of sampling points of the variable convolution layer. The content of the offset domain is the learned offset, and the channel number corresponds to the number of 2D learned offsets. Each offset is composed of a horizontal coordinate offset and a vertical coordinate offset, that is, the offset at the current pixel position is determined by the horizontal and vertical offsets. Second, the offset in the offset domain and the pixel value index in the case channel attention feature map are added to obtain the absolute coordinates of the offset domain. Then, the absolute coordinates of the offset domain are down-rounded and up-rounded, and bilinear interpolation is performed to obtain the offset coordinates. Finally, the case channel attention feature map is sampled using the offset coordinates to obtain the offset sampling value, and the offset sampling value is weighted and summed using the corresponding weights of the variable convolution to obtain the case spatial attention feature map.
[0039] In the fifth step, the case channel attention feature map is convolved and mapped to obtain a first channel mapping feature map, a second channel mapping feature map, and a third channel mapping feature map. The first channel mapping feature map, the second channel mapping feature map, and the third channel mapping feature map can be obtained by using convolution kernels with expansion rates of 1, 3, and 4, respectively, and a convolution layer of step 3 to capture different scale dependencies and learn more nonlinear features. The convolution layer of step 3 learns more nonlinear features by capturing different scale dependencies to obtain feature maps.
[0040] In the sixth step, the case spatial attention feature map, the first channel mapping feature map, the second channel mapping feature map, and the third channel mapping feature map are multi-scale fused to obtain a multi-scale fusion case feature map as a sixth case document feature map. The sixth case document feature map can be obtained by element-wise multiplying the case spatial attention feature map with the first channel mapping feature map, the second channel mapping feature map, and the third channel mapping feature map in sequence and then concatenating the features to obtain the feature map.
[0041] In step 102, a case region logical hierarchy tree is constructed from a case region information set to obtain a case region logical hierarchy tree.
[0042] In some embodiments, the execution subject can construct a case region logical hierarchy tree from a case region information set to obtain a case region logical hierarchy tree. The case region logical hierarchy tree can be a tree structure that displays the logical layout relationship between the case region information included in the case region information set in the form of a logical tree.
[0043] As an example, the execution subject can first filter at least one case text information set with an element category of a text block from the case region information set as a case text block region information set. The text block can include a title and a text paragraph. Then, a text style extraction is performed on the case text block region information set by using an OCR (Optical Character Recognition) model, a regular expression, and an NLP (Natural Language Processing) to obtain a text block style feature information set. The text block style feature information can include, but is not limited to, at least one of the following: text semantics, text font style, and text encoding model. Finally, a logic tree is constructed for the case text block region information set according to the text block style feature information by using a logic structure analysis engine to obtain a medical record region logic hierarchical tree. The logic structure analysis rule engine can include, but is not limited to, at least one of the following analysis rules: a title number-based depth calculation rule, a font size-based hierarchical division rule, and a priority-based merging rule. The title number-based depth calculation rule can be a rule that the depth of a text block with a "number." numbering format at the beginning of the text is equal to the number of "." in the numbering format plus 1. The font size-based hierarchical division rule can be a rule that the hierarchical level of a text block is equal to the hierarchical level of an interval corresponding to the font of the text block in a preset font size grading table according to the preset font size grading table. The preset font size grading table can be a table that is preset and used to determine the hierarchical level of a text block according to the font size. For example, one mapping relationship in the preset font size grading table can be that the hierarchical level is 1 when the font size is greater than or equal to 14 pounds and less than 18 pounds. The priority-based merging rule can be a rule that the final depth of a text block is equal to the depth corresponding to the numbering format when the depth corresponding to the numbering format is greater than or equal to 0, and otherwise, the final depth of the text block is equal to the hierarchical level corresponding to the font size.
[0044] In step 103, the case segmentation processing is performed on the multi-modal medical case document according to the case region logic hierarchical tree and the case region information set to obtain a medical segmented case set.
[0045] In some embodiments, the multi-execution subject can perform case segmentation processing on the multi-modal medical case document according to the case region logic hierarchical tree and the case region information set to obtain a medical segmented case set. The medical segmented cases in the medical segmented case set can be cases of the same patient.
[0046] In some optional implementations of some embodiments, the case segmentation processing on the multi-modal medical case document according to the case region logic hierarchical tree and the case region information set to obtain a medical segmented case set can include the following steps:
[0047] In a first step, at least one case region information with a category label of title is selected from the case region information set.
[0048] In a second step, expression matching is performed between the at least one case region information and a preset case segmentation regular expression set to obtain an expression matching result set. The preset case segmentation regular expression in the preset case segmentation regular expression set can be a regular expression for matching a preset word segmentation in a medical record region. For example, the preset case segmentation regular expression set can include: case \d+, patient [A-Z].
[0049] In a third step, a plurality of case region information with a matching success represented by the expression matching result is selected from the at least one case region information as a title case region information set.
[0050] In a fourth step, title region sorting is performed on the title case region information set according to the case region logical hierarchical tree to obtain a title case region information sequence.
[0051] As an example, the execution subject can determine the sorting order information of the title case region information set in the medical record region logical hierarchical tree. Then, the title region sorting is performed on the title medical record region information set according to the sorting order information to obtain the title case region information sequence.
[0052] In a fifth step, the title case region information sequence and case segmentation prompt word information are input into a case segmentation large language model to obtain a case segmentation position information set. The case segmentation position information in the case segmentation position information set can be information of a position for text segmentation of the multi-modal medical case document. The case segmentation prompt word information can be information of a prompt word for guiding the large prediction model to implement case segmentation. The case segmentation prompt word information can be information of a prompt word generated by a large language model. The case segmentation prompt word information can be “Does the following text describe the same patient? [Paragraph 1] [Paragraph 2]”. The case segmentation large language model can be a large prediction model that inputs the title case region information sequence and the case segmentation prompt word information to determine whether the title case region information sequence is a case of the same patient. For example, the case segmentation large language model can be an LLM (Large Language Model) model.
[0053] In a sixth step, case segmentation processing is performed on the multi-modal medical case document according to the case segmentation position information set to obtain a medical segmentation case set.
[0054] In step 104, case information extraction prompt word information for the medical segmentation case set is generated.
[0055] In some embodiments, the execution subject can generate case information extraction prompt information for the medical segmentation case set. The case information extraction prompt information can be text information used to guide the large language model to output in a predetermined format and perform text extraction. For example, the case segmentation prompt information can be: Role: Senior AI assistant in medical field; Task: Extract structured diagnosis and treatment data from case text; Extraction steps: 1. Clinical examination information: [Clinical manifestations]: chief complaint (reason for consultation), past medical history, current medical history (symptom characteristics / associated characteristics / accepted treatment and effect); [Auxiliary examination]: laboratory examination (blood routine / urine routine / biochemistry / pathogen), imaging (X-ray / ultrasound / CT / MRI (key findings), special examination (ECG / endoscopy / biopsy results); 2. Diagnosis classification: - Systemic disease classification (such as respiratory / circulatory / digestive system, etc.), - Etiological classification (infection / trauma / metabolism / tumor, etc.), - Severity classification (mild / moderate / severe / acute), - Differential diagnosis and basis, - Comorbidity annotation; 3. Treatment plan: - Drug treatment (drug name, dose, frequency, course of treatment), - Non-drug treatment (surgery / physical therapy / lifestyle intervention), - Referral indications; 4. Treatment effect: - Follow-up plan (monitoring indicators / time nodes), - Efficacy evaluation (symptom relief degree / function recovery / complications).
[0056] In step 105, a multi-modal large language model is used to perform structured information extraction processing on the medical segmentation case set according to the case information extraction prompt information, to obtain a case text structured information set.
[0057] In some embodiments, the execution subject can perform structured information extraction processing on the medical segmentation case set to obtain a case text structured information set. The case text structured information in the case text structured information set can be patient body state information in json format that extracts patient basic information and body index information. The body index information can include at least one of the following: medical image feature description information, disease type judgment information in disease diagnosis, medication information of treatment plan, and treatment plan information. The multi-modal large language model can be a large language model that performs text extraction on different modal data of input case information extraction prompt information and medical segmentation case set. For example, the multi-modal large language model can be at least one of the following: DeepSeek model, ChatGPT (Chat Generative Pre-trained Transformer), NExT-GPT model.
[0058] Step 106, according to the case region information set, the medical segmentation case set is positionally associated with the category recognition, and the medical image category label information set is obtained.
[0059] In some embodiments, the above-mentioned execution subject can perform positionally associated category recognition on the above-mentioned medical segmentation case set according to the above-mentioned case region information set, and obtain a medical image category label information set. The medical image category label information in the medical image category label information set can represent the information of the category of the medical image of the case to which each medical image in the medical segmentation case belongs. The category can include but is not limited to at least one of the following: X-ray image, ultrasound image, CT (Computed Tomography, Electronic Computer Tomography) / MRI (Magnetic Resonance Imaging, Magnetic Resonance Imaging) key findings.
[0060] In some optional implementations of some embodiments, the above-mentioned positionally associated category recognition on the above-mentioned medical segmentation case set according to the above-mentioned case region information set to obtain the medical image category label information set can include the following steps:
[0061] First, at least one case region information with a category label of image and text block is filtered from the above-mentioned case region information set as an image case region information set and a text case region information set.
[0062] Second, the image coordinate set corresponding to the above-mentioned image case region information set and the text position set corresponding to the above-mentioned text case region information set are respectively subjected to coordinate standardization processing to obtain a standardized image coordinate set and a standardized text coordinate set. The coordinate standardization processing can be to convert absolute coordinates into normalized coordinates relative to the multi-modal medical case document to eliminate the standardization processing of different document size differences.
[0063] Third, according to the above-mentioned standardized image coordinate set and the above-mentioned standardized text coordinate set, a spatial proximity value group set is determined. The spatial proximity value in the spatial proximity value group set can represent the spatial proximity of any image case region information and any text case region information on the multi-modal medical case document.
[0064] As an example, the upper number execution subject can first determine the absolute value of the difference between each image center point vertical coordinate in the set of image center point vertical coordinates corresponding to the above-mentioned normalized image coordinate set and each text vertical center vertical coordinate in the set of text vertical center vertical coordinates corresponding to the above-mentioned normalized text coordinate set as the vertical proximity, obtaining a set of vertical proximity groups. Then, the difference between each image center point horizontal coordinate in the set of image center point horizontal coordinates corresponding to the above-mentioned normalized image coordinate set and each text horizontal coordinate in the set of text horizontal coordinates corresponding to the above-mentioned normalized text coordinate set is taken as the horizontal proximity, obtaining a set of horizontal proximity groups. Finally, the weighted sum of each vertical proximity in the above-mentioned set of vertical proximity groups and the corresponding horizontal proximity in the above-mentioned set of horizontal proximity groups is obtained, obtaining a set of spatial proximity numerical values.
[0065] The fourth step is to construct an image-text location relationship graph according to the above-mentioned set of spatial proximity numerical values. Wherein, the above-mentioned image-text location relationship graph can be a weighted graph with the above-mentioned set of image case region information and the above-mentioned set of text case region information as nodes, and the spatial proximity numerical value and the semantic similarity as the weight value of the connection edge. The above-mentioned semantic similarity can be the similarity obtained by calculating the cosine similarity of the image case region information and the above-mentioned text case region information through CLIP (Contrastive Language-Image Pre-Training, contrastive language-image pre-training model).
[0066] As an example, the above-mentioned execution subject can first perform text semantic extraction on the above-mentioned set of text case region information to obtain a medical record text semantic feature vector. Secondly, image global visual extraction is performed on the above-mentioned set of image case region information to obtain a set of image feature vectors. Then, through CLIP, the cosine similarity of the set of image case region information and the above-mentioned set of text case region information is determined as the image-text semantic similarity, obtaining a set of image-text semantic similarity groups. After that, the weighted sum of each image-text semantic similarity in the above-mentioned set of image-text semantic similarity groups and the spatial proximity numerical value corresponding to the above-mentioned set of spatial proximity numerical values is determined, obtaining a set of connection edge weights. Finally, a weighted graph is constructed with the above-mentioned set of image case region information and the above-mentioned set of text case region information as nodes, and the set of connection edge weights as the weight value of the connection edge, obtaining an image-text location relationship graph.
[0067] The fifth step is to perform association probability prediction on the above-mentioned image-text location relationship graph to obtain a set of image-text association probability groups. Wherein, the image-text association probability in the above-mentioned set of image-text association probability groups can represent the degree of association between any image case region information and any text case region information. The above-mentioned association probability prediction can be a prediction through a convolutional neural network.
[0068] Step 6, according to the image-text association probability group set, determine the medical segmentation case to which each image case region information in the image case region information set belongs as the target medical segmentation case, and obtain the target medical segmentation case set. The target medical segmentation case can be a case composed of the image case region information and the text case region information with the strongest association degree.
[0069] As an example, the execution subject can be a medical segmentation case to which the text case region information corresponding to the maximum image-text association probability corresponding to each image case region information in the image-text association probability group set belongs, as the medical segmentation case to which the image case region information belongs, as the target medical segmentation case, and obtain the target medical segmentation case set.
[0070] Step 7, perform image type identification on each medical image included in each target medical segmentation case in the target medical segmentation case set, and obtain a medical image class label information set. The image type identification can be classification identification of the image class of the medical image using a medical image classification network. The medical image classification network can be a model for classifying the type of medical image corresponding to the image case region information. The medical image classification network can be, but is not limited to, at least one of the following: a three-dimensional convolutional neural network, a Vision Transformer. For at least one image case region information with a confidence value less than a pre-set confidence value output by the medical image classification network, input the text case region information with the strongest association into a multi-modal large model for auxiliary reclassification, and obtain a medical image class label information set.
[0071] Optionally, determining the spatial proximity value group set according to the standardized image coordinate set and the standardized text coordinate set can include the following steps:
[0072] Step 1, for each image case region information in the image case region information set, perform the following spatial proximity value generation steps:
[0073] Sub-step 1, determine a region vertical distance group of the image case region information and the standardized text coordinate set. The region vertical distance in the region vertical distance group can represent the relative positional relationship of the image case region information and the text case region information in the reading flow, and a smaller value indicates a closer region position. In practice, the absolute value of the longitudinal coordinate of the standardized image coordinate corresponding to the image case region information and the longitudinal coordinate of each standardized text coordinate in the standardized text coordinate set are determined as the region vertical distance, and a region vertical distance group is obtained.
[0074] Sub-step 2 involves determining a set of horizontal overlap values for the regions based on the aforementioned image case region information, the standardized text coordinate set, and the text case region information set. The horizontal overlap values in this set characterize the horizontal overlap ratio between the image case region information and the text case region information; a larger value indicates a stronger horizontal correlation. In practice, the horizontal overlap values for the image case region information and the text case region information set are determined based on the horizontal coordinate and image region width corresponding to the image case region information, the standardized text coordinate set, and the text region width corresponding to the text case region information set.
[0075] As an example, the aforementioned execution entity can utilize a horizontal overlap function to determine the set of horizontal overlap values between the image case region information and the text case region information set, based on the horizontal coordinates and image region width corresponding to the image case region information, the standardized text coordinate set, and the text region width value corresponding to the text case region information set. The horizontal overlap function can be expressed as:
[0076] .
[0077] in, This represents a horizontally overlapping function. Indicates the first The standardized x-coordinates corresponding to the image case region information. Indicates the first The standardized image region width corresponding to the image case region information. Indicates the first The standardized text region width value corresponding to the text case region information. Indicates the first The standardized text region width value corresponding to the text case region information.
[0078] Sub-step 3: Determine the coordinates of the image center point of the aforementioned image case region information. The aforementioned first preset value can be a pre-set value. For example, the aforementioned first preset value can be 2. In practice, the image center point coordinates are determined as the sum of the ratio of the standardized text coordinates corresponding to the aforementioned image case region information, the corresponding image width, and the first preset value.
[0079] Sub-step 4, determining a text center point coordinate set of the above normalized text coordinate set. Wherein, the text center point coordinate in the text center point coordinate set can represent the position coordinate of the center point of the normalized text coordinate. In practice, the sum of the horizontal coordinate of each normalized text coordinate in the above normalized text coordinate set, the corresponding text width and the first preset value is determined as the text center point coordinate set.
[0080] Sub-step 5, determining an image-text center point horizontal distance set of the above image center point coordinate and the text center point coordinate set. In practice, the absolute value of the difference between the above image center point coordinate and each text center point coordinate in the text center point coordinate set is determined as the image-text center point horizontal distance, and the image-text center point horizontal distance set is obtained.
[0081] Sub-step 6, determining a center alignment degree set by the difference between the second preset value and each image-text center point horizontal distance in the above image-text center point horizontal distance set. Wherein, the center alignment degree in the center alignment degree set can represent the horizontal alignment degree of the center points of the image case region information and the text case region information. The second preset value can be a preset value. For example, the second preset value can be 1.
[0082] Sub-step 7, generating a spatial proximity value set according to the above region vertical distance set, the region horizontal overlap value set and the center alignment degree set.
[0083] As an example, the execution subject can generate the spatial proximity value set of the image case region information and the text case region information set according to the above region vertical distance set, the region horizontal overlap value set and the center alignment degree set by using a spatial proximity measurement function. Wherein, the spatial proximity measurement function can be:
[0084] .
[0085] Wherein, represents the spatial proximity measurement function. represents the region vertical distance. represents the region horizontal overlap value, i.e. the function value of the horizontal overlap function corresponding to the current scene. represents the center alignment degree.
[0086] In the process of adopting the technical solutions to solve the above technical problem one, the following technical problem two is often accompanied: there is an inaccurate matching problem of image case area information and text case area information in medical image position category recognition, leading to inaccurate segmentation when segmenting image case area information to each medical segmentation case, inaccurate category recognition of medical images, and further inaccurate multi-modal information association, storing a large number of incorrect multi-modal synthetic medical cases, wasting storage resources. In view of the above technical problem two, the conventional solution is generally: adopting the Hungarian algorithm, determining each image case area information in the image case area information set to which the medical segmentation case belongs as the target medical segmentation case according to the image text association probability group set, to obtain the target medical segmentation case set. However, the above conventional solution still has the following problems: since the traditional Hungarian algorithm can only handle one-to-one matching problems, when there is one-to-many or many-to-many situations, random discarding or rule screening is usually used, which has certain objectivity and cannot be updated in time, leading to low matching accuracy of image case area information and text case area information, inaccurate segmentation when segmenting image case area information to each medical segmentation case, inaccurate category recognition of medical images, and further inaccurate multi-modal information association, storing a large number of incorrect multi-modal synthetic medical cases, wasting storage resources, reducing the quality of the multi-modal case database, and prolonging the time to build the multi-modal case database. And the inventors consider the shortcomings of the above conventional solution, and combine the advantages / technical status of the image text multi-target matching technology owned by the company where the inventors are located, and we decide to adopt the following solution:
[0087] In some optional implementations of some embodiments, the above determining each image case area information in the image case area information set to which the medical segmentation case belongs as the target medical segmentation case according to the image text association probability group set, to obtain the target medical segmentation case set, can include the following steps:
[0088] First, determine the document page cost value of each image text association probability in the image text association probability group set, to obtain the document page cost value group set. Wherein, the document page cost value in the document page cost value group set can represent the absolute value of the difference in page between the image case area information and the text case area information corresponding to the image text association probability in the multi-modal medical case document, the smaller the value, the more adjacent the spatial position, and the stronger the association.
[0089] In the second step, the initial image-text association cost matrix is generated according to the image-text association probability set and the document page cost value set. Each image-text association cost value in the initial image-text association cost matrix can represent the association cost between the image case region information and the text case region information, and is used to measure the matching cost of whether the image case region information and the text case region information belong to the same case. The smaller the image-text value is, the greater the possibility of the association between the image case region information and the text case region information is.
[0090] As an example, the execution subject can first determine a first weight value and a second weight value for the image-text association probability set and the document page cost value set. The first weight value can be the importance of each image-text association probability in the image-text association probability set. The second weight value can be the importance of each document page cost value in the document page cost value set. The determination can be a determination by a heuristic algorithm. For example, the heuristic algorithm can include, but is not limited to, at least one of the following: genetic algorithm, simulated annealing algorithm, whale algorithm, particle swarm algorithm. Then, the image-text association probability set, the document page cost value set, the first weight value and the second weight value are weighted and summed to obtain the initial image-text association cost matrix.
[0091] In the third step, the initial image-text association cost matrix is converted into a square matrix to obtain an image-text association cost square matrix as an image-text bipartite graph. The image-text association cost square matrix can be an extension of the initial image-text association cost matrix to form a cost square matrix with the same number of image case region information and text case region information. The image-text bipartite graph can be an undirected graph including two disjoint image case region information sets and text case region information sets, and each edge has two vertices belonging to the image case region information set and the text case region information set. The image-text association cost square matrix can be a representation of the image-text bipartite graph. In practice, the execution subject can add image case region information of the number of images minus the number of texts in response to determining that the number of images of the image case region information is greater than the number of texts of the text case region information, and set the initial image-text association cost matrix corresponding to the added text case region information to 100 to obtain the image-text association cost square matrix. In response to determining that the number of images of the image case region information is less than the number of texts of the text case region information, add text case region information of the number of texts minus the number of images, and set the initial image-text association cost matrix corresponding to the added image case region information to 100 to obtain the image-text association cost square matrix.
[0092] In the fourth step, at least one image case region information and at least one text case region information that are not matched are determined as a set of unmatched nodes according to the image-text bipartite graph. The unmatched nodes in the set of unmatched nodes can be image nodes that are not matched with text case region information, or text nodes that are not matched with image case region information.
[0093] As an example, the execution subject can first filter out the highest image-text association cost values in the image-text association cost matrix corresponding to the image-text bipartite graph according to each row and each column to obtain an initial image-text matching set. The initial image-text matching set can be a set including matched nodes and matched edges. In practice, the execution subject can filter out the highest image-text association cost values in the image-text association cost matrix according to each row and each column by using a KM (Kuhn Munkres, maximum weight matching) algorithm to obtain an initial image-text matching set. Then, the initial image-text matching set is traversed to obtain at least one image case region information and at least one text case region information that are not matched as a set of unmatched nodes.
[0094] In the fifth step, an alternating path search and a depth-first search are performed on the set of unmatched nodes to obtain a set of node augmented paths. The node augmented paths in the set of node augmented paths can be paths in which the start node and the end node are both unmatched nodes, and matched edges and unmatched edges appear alternately. The matched edges can be edges in which the start node and the end node are both matched nodes. The unmatched edges can be edges in which either the start node or the end node, or both, are unmatched nodes.
[0095] As an example, the execution subject can start from the unmatched nodes in the set of unmatched nodes, search in an alternating mode of unmatched edges and matched edges, and search according to an improved depth-first search method that preferentially selects higher image-text association costs until each unmatched node in the set of unmatched nodes is searched, to obtain a set of node augmented paths. The improved depth-first search method can be a depth-first search method that is optimized by using a Hopcroft-Karp algorithm.
[0096] In the sixth step, the initial image-text matching set is iteratively optimized according to the set of node augmented paths to obtain a target image-text matching graph. The target image-text matching graph can be an image-text matching graph obtained by adding the set of node augmented paths to the initial image-text matching set.
[0097] In the seventh step, in response to determining that there are image case region information matching multiple text case region information in the target image-text matching graph, the image-text matching graph is prioritized according to the image-text association probability group set, the spatial proximity value group set, and the image-text semantic correlation, to obtain an image-text ordered matching graph. The image-text ordered matching graph can be a one-to-one matching graph of image case region information and text case region information. The image-text semantic correlation can be inputting the image case region information and the text case region information into the CLIP model to determine the semantic similarity of the image case region information and the text case region information.
[0098] As an example, in response to determining that there are image case region information matching multiple text case region information in the target image-text matching graph, the executing subject can first filter out the text case region information corresponding to the image-text association probability with the largest value from the multiple text case region information as the most matching text case region information. Then, in response to determining that the text case region information corresponding to the image-text association probability with the largest value also has multiple, the text case region information with the largest spatial proximity value is selected as the most matching text case region information. Finally, in response to determining that the largest spatial proximity value also has multiple, the text case region information with the largest semantic similarity to the image case region information is selected as the most matching text case region information by using the CLIP model, until all conflicts of image case region information matching multiple text case region information are resolved, to obtain the image-text ordered matching graph.
[0099] In the eighth step, according to the image-text ordered matching graph, each image case region information in the image case region information set is determined to belong to a medical segmentation case as a target medical segmentation case, to obtain a target medical segmentation case set.
[0100] As an example, the executing subject can query the image-text ordered matching graph to determine that each image case region information in the image case region information set belongs to a medical segmentation case as a target medical segmentation case, to obtain a target medical segmentation case set.
[0101] The above technical solution, combined with step "step 107" and its related content as an invention point of an embodiment of the present disclosure, solves the technical problem two mentioned in the background art. Due to the fact that the traditional Hungarian algorithm can only handle one-to-one matching problems, when there is a one-to-many or many-to-many situation, random discarding or rule screening is usually used, which has certain objectivity and cannot be updated in time, resulting in low matching accuracy of image case region information and text case region information, inaccurate segmentation when segmenting image case region information into each medical segmentation case, inaccurate classification of medical images, and further inaccurate correlation of multi-modal information, storage of a large number of incorrect multi-modal synthetic medical cases, waste of storage resources, reduction of the quality of the multi-modal case database, and prolongation of the time to build the multi-modal case database. The matching accuracy is low, the segmentation is inaccurate, and the classification of medical images is inaccurate, which further leads to waste of storage resources. The matching accuracy is often as follows: due to the fact that the traditional Hungarian algorithm can only handle one-to-one matching problems, when there is a one-to-many or many-to-many situation, random discarding or rule screening is usually used, which has certain objectivity and cannot be updated in time, resulting in low matching accuracy of image case region information and text case region information, inaccurate segmentation when segmenting image case region information into each medical segmentation case, inaccurate classification of medical images, and further inaccurate correlation of multi-modal information, storage of a large number of incorrect multi-modal synthetic medical cases. If the above factors are solved, the matching accuracy, segmentation accuracy, and classification accuracy of medical images can be improved, and the multi-modal information correlation accuracy can be improved, the storage resource waste can be reduced, the quality of the multi-modal case database can be improved, and the time to build the multi-modal case database can be shortened. In order to achieve this effect, the present disclosure first dynamically weights and fuses the image text association cost initial matrix by the document page cost numerical value and the image text association probability based on the spatial proximity value and the image text semantic similarity, which can improve the accuracy of the association degree of the image text association cost initial matrix. Then, the image text bipartite graph, the initial image text matching set and the unmatched node set are constructed, and the unmatched node set is searched by the improved depth-first search and the iteration optimization, which can batch process the augmented path and reduce the computational time complexity. Next, the image text sorting matching graph is obtained by performing priority sorting according to the image text association probability group set, the spatial proximity value group set and the image text semantic correlation, which can solve the matching conflict situation of many-to-many and one-to-many, can improve the optimal matching of image case region information and text case region information, and can improve the matching result. Finally, the multi-modal information correlation and case storage based on the individual case are performed on the medical image set corresponding to the case text structured information set and the medical image class label information set by the target medical segmentation case set, and the multi-modal case database is obtained, which can improve the quality of the multi-modal case database, shorten the time to build the multi-modal case database, and reduce the waste of storage resources.
[0102] Step 107, the multi-modal information association is performed on the case text structured information set and the medical image category label information set, to obtain a multi-modal synthetic medical case information set, and the multi-modal synthetic medical case information set is stored as a case to obtain a multi-modal case database.
[0103] In some embodiments, the execution subject can perform multi-modal information association on the case text structured information set and the medical image category label information set to obtain a multi-modal synthetic medical case information set, and store the multi-modal synthetic medical case information set as a case to obtain a multi-modal case database. The multi-modal synthetic medical case information in the multi-modal synthetic medical case information set can be information that the medical image corresponding to the case text structured information and the medical image category label information belongs to the same case of the same patient. The multi-modal case database can be a database for storing the medical image set corresponding to the case text structured information set and the medical image category label information set. The multi-modal information association can be multi-modal information association based on individual cases. As shown in Figure 2 The left side of FIG. 1 shows the entire process of steps 101-107 for processing multi-modal medical case documents, Figure 2 The left side of FIG. 1 shows the entire process of steps 101-107 for processing multi-modal medical case documents,
[0104] As an example, the execution subject can first determine the medical segmentation case to which each medical image in the medical image set corresponding to the medical image category label information set belongs through the medical image category label information set. Then, the case text structured information set and the medical image are associated and integrated through the medical segmentation case to obtain a multi-modal synthetic medical case information set. Finally, the multi-modal synthetic medical case information set is stored as a case to obtain a multi-modal case database.
[0105] Further referring to Figure 3 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a multi-modal case database construction device, which corresponds to the method embodiments shown in Figure 1 The multi-modal case database construction device can be applied in various electronic devices.
[0106] As Figure 3As shown, a multi-modal case database construction apparatus 300 includes a document layout area detection unit 301, a region logical hierarchy tree construction unit 302, a case segmentation unit 303, a generation unit 304, a case information extraction unit 305, a position association category identification unit 306, and a multi-modal information association unit 307. The document layout area detection unit 301 performs document layout area detection processing on an acquired multi-modal medical case document to obtain a case region information set. The region logical hierarchy tree construction unit 302 constructs a region logical hierarchy tree from the case region information set to obtain a case region logical hierarchy tree. The case segmentation unit 303 performs case segmentation processing on the multi-modal medical case document according to the case region logical hierarchy tree and the case region information set to obtain a medical segmented case set. The generation unit 304 generates case information extraction prompt word information for the medical segmented case set. The case information extraction unit 305 performs structured information extraction processing on the medical segmented case set according to the case information extraction prompt word information using a multi-modal large language model to obtain a case text structured information set. The position association category identification unit 306 performs position association category identification on the medical segmented case set according to the case region information set to obtain a medical image category label information set. The multi-modal information association unit 307 performs multi-modal information association on the case text structured information set and the medical image category label information set to obtain a multi-modal synthetic medical case information set, and stores the multi-modal synthetic medical case information set to obtain a multi-modal case database.
[0107] It can be understood that the units described in the multi-modal case database construction apparatus 300 correspond to the respective steps in the method described above. Figure 1 The operations, features, and advantages described above for the method also apply to the multi-modal case database construction apparatus 300 and the units included therein, and will not be described again.
[0108] Reference is made below to Figure 4 which shows a structural schematic diagram of an electronic device (e.g., electronic device) 400 suitable for implementing some embodiments of the present disclosure. Figure 4 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0109] As Figure 4As shown, the electronic device 400 can include a processing device (e.g., a central processor, a graphics processor, etc.) 401 that can perform various appropriate actions and processes according to programs stored in a read only memory (ROM) 402 or loaded into a random access memory (RAM) 403 from a storage device 408. Various programs and data required for the operation of the electronic device 400 are also stored in the RAM 403. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0110] Generally, the following devices can be connected to the I / O interface 405: input devices 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 408 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 409. The communication devices 409 can allow the electronic device 400 to exchange data with other devices wirelessly or wiredly. Although Figure 4 The electronic device 400 is shown with various devices, but it is understood that all of the shown devices are not required to be implemented or present. More or fewer devices can alternatively be implemented or present. Figure 4 Each block shown in the flowcharts can represent a device or multiple devices as needed.
[0111] In particular, processes described above with reference to the flowcharts can be implemented as a computer software program according to some embodiments of the present disclosure. For example, some embodiments of the present disclosure include a computer program product including a computer program carried on a computer readable medium, the computer program containing program codes for performing the methods shown in the flowcharts. In some such embodiments, the computer program can be downloaded and installed from a network through the communication devices 409, or installed from the storage devices 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-described functions defined in the methods of some embodiments of the present disclosure are performed.
[0112] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0113] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0114] The computer readable medium can be included in the electronic device, or can exist separately from the electronic device. The computer readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to perform the following operations: performing document layout region detection processing on the acquired multi-modal medical case document to obtain a case region information set; performing region logical hierarchy tree construction on the case region information set to obtain a case region logical hierarchy tree; performing case segmentation processing on the multi-modal medical case document according to the case region logical hierarchy tree and the case region information set to obtain a medical segmented case set; generating case information extraction prompt word information for the medical segmented case set; performing structured information extraction processing on the medical segmented case set according to the case information extraction prompt word information by using a multi-modal large language model to obtain a case text structured information set; performing position association category identification on the medical segmented case set according to the case region information set to obtain a medical image category label information set; performing multi-modal information association on the case text structured information set and the medical image category label information set to obtain a multi-modal synthesized medical case information set; and performing case storage on the multi-modal synthesized medical case information set to obtain a multi-modal case database.
[0115] Computer program code for carrying out operations of some embodiments of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0116] The flow and block diagrams in the drawings represent possible architectural, functional, and operational architectures of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block can represent a module, a segment, or a portion of code that comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or combinations of hardware and software.
[0117] The units described in some embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. The described units can also be arranged in a processor, for example, it can be described that: a processor includes a document layout area detection unit, a region logical hierarchy tree construction unit, a case segmentation unit, a generation unit, a case information extraction unit, a position association category identification unit, and a multi-modal information association unit. Among them, the name of these units does not constitute a limitation to the unit itself in some cases, for example, the document layout area detection unit can also be described as "a unit that performs document layout area detection processing on the obtained multi-modal medical case document to obtain a case region information set".
[0118] The functions described above in the detailed description can be performed at least in part by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), etc.
[0119] The above description is merely some embodiments of the present disclosure and a description of the principles of the technology used. Those skilled in the art should understand that the scope of the application involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above inventive concept. For example, the above features are replaced with technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form technical solutions.
Claims
1. A method for constructing a multimodal case database, comprising: The acquired multimodal medical case documents are processed by document layout region detection to obtain a case region information set; A regional logical hierarchy tree is constructed from the case region information set to obtain the case region logical hierarchy tree; Based on the case region logical hierarchy tree and the case region information set, the multimodal medical case document is segmented to obtain a medical segmentation case set; Generate case information extraction prompts for the medical segmentation case set; Using a multimodal large language model, prompt word information is extracted based on the case information, and structured information extraction processing is performed on the medical segmentation case set to obtain a case text structured information set; Based on the case region information set, location association category identification is performed on the medical segmentation case set to obtain a medical image category label information set. The step of performing location association category identification on the medical segmentation case set based on the case region information set to obtain the medical image category label information set includes: At least one case region information with category labels of image and text block is selected from the case region information set to form the image case region information set and the text case region information set; The image coordinate set corresponding to the image case region information set and the text position set corresponding to the text case region information set are respectively subjected to coordinate standardization processing to obtain the standardized image coordinate set and the standardized text coordinate set. Based on the standardized image coordinate set and the standardized text coordinate set, determine the spatial proximity value set; Based on the set of spatial proximity values, construct an image-text location relationship graph; The image-text location relationship graph is used to predict the association probability, resulting in an image-text association probability set. Based on the image-text association probability set, the medical segmentation case to which each image case region information belongs in the image case region information set is determined as the target medical segmentation case, thus obtaining the target medical segmentation case set; For each medical image included in the target medical segmentation case set, image type identification is performed to obtain a medical image category label information set; Multimodal information association is performed on the structured information set of case texts and the medical image category label information set to obtain a multimodal synthetic medical case information set. The multimodal synthetic medical case information set is then stored to obtain a multimodal case database.
2. The method according to claim 1, wherein, The process of performing document layout region detection on the acquired multimodal medical case documents yields a case region information set, including: The multimodal medical case document is input into the activation convolutional feature extraction network included in the document layout region detection model to obtain the first case document feature map. The document layout region detection model further includes: a residual activation feature extraction backbone network, a cross-stage fusion network, a multi-scale visual feature enhancement network, and multiple multi-scale cross-feature fusion networks. The first case document feature map is input into the residual activation feature extraction backbone network to obtain the second case document feature map and the third case document feature map; The second case document feature map and the third case document feature map are input into the cross-stage fusion network to obtain the fourth case document feature map and the fifth case document feature map; The first case document feature map is input into the multi-scale visual feature enhancement network to obtain the sixth case document feature map; The feature maps of the fifth and sixth case documents are input into the multi-scale cross-feature fusion network to obtain a case region information set.
3. The method according to claim 2, wherein, The multi-scale visual feature enhancement network includes: a global pooling channel attention network and a variable convolutional spatial attention network; and The step of inputting the first case document feature map into the multi-scale visual feature enhancement network to obtain the sixth case document feature map includes: The first case document feature map is input into the global average pooling layer and the global max pooling layer of the global pooling channel attention network to obtain the case channel average pooling feature map and the case channel max pooling feature map. Channel nonlinear fusion is performed on the average pooling feature map and the maximum pooling feature map of the case channel to obtain the average channel nonlinear feature map and the maximum channel nonlinear feature map of the case. The average channel nonlinear feature map of the case, the maximum channel nonlinear feature map of the case, and the first case document feature map are weighted and summed to obtain the case channel attention feature map; The case channel attention feature map is input into the variable convolutional spatial attention network to obtain the case spatial attention feature map; The case channel attention feature map is convolved and mapped to obtain the first channel mapping feature map, the second channel mapping feature map and the third channel mapping feature map; Multi-scale fusion is performed on the case spatial attention feature map, the first channel mapping feature map, the second channel mapping feature map, and the third channel mapping feature map to obtain a multi-scale fused medical record feature map, which serves as the sixth case document feature map.
4. The method according to claim 1, wherein, The step of performing case segmentation processing on the multimodal medical case document based on the case region logical hierarchy tree and the case region information set to obtain a medical segmentation case set includes: Filter at least one case region information with the category label as the title from the case region information set; The at least one case region information is matched with a preset set of case segmentation regular expressions to obtain a set of expression matching results. Multiple case region information that represent successful matching by expression matching results are selected from the at least one case region information and used as the title case region information set; Based on the logical hierarchy tree of the case regions, the title case region information set is sorted to obtain the title case region information sequence; The title case region information sequence and case segmentation prompt word information are input into the case segmentation large language model to obtain the case segmentation location information set; Based on the case segmentation location information set, the multimodal medical case documents are segmented to obtain a medical segmentation case set.
5. The method according to claim 1, wherein, The step of determining the spatial proximity value set based on the standardized image coordinate set and the standardized text coordinate set includes: For each image case region in the image case region information set, perform the following spatial proximity value generation step: Determine the region vertical distance group of the image case area information and the standardized text coordinate set; Based on the image case region information, the standardized text coordinate set, and the text case region information set, determine the region horizontal overlap value group; Determine the coordinates of the image center point of the image case region information; Determine the coordinate set of the text center point of the standardized text coordinate set; Determine the horizontal distance group between the image center point coordinates and the text center point coordinates group; The difference between the second preset value and the horizontal distance of each image text center point in the horizontal distance group is determined as the center alignment group; Based on the vertical distance group of the region, the horizontal overlap group of the region, and the center alignment group, a spatial proximity group is generated.
6. A multimodal case database construction device, comprising: The document layout region detection unit performs document layout region detection processing on the acquired multimodal medical case documents to obtain a case region information set; The regional logical hierarchy tree construction unit constructs a regional logical hierarchy tree for the case region information set to obtain the case region logical hierarchy tree. The case segmentation unit performs case segmentation processing on the multimodal medical case document based on the case region logical hierarchy tree and the case region information set to obtain a medical segmentation case set; The generation unit generates case information extraction prompts for the medical segmentation case set; The case information extraction unit uses a multimodal large language model to extract prompt word information based on the case information and performs structured information extraction processing on the medical segmentation case set to obtain a case text structured information set; The location association category identification unit performs location association category identification on the medical segmentation case set based on the case region information set to obtain a medical image category label information set. The step of performing location association category identification on the medical segmentation case set based on the case region information set to obtain the medical image category label information set includes: selecting at least one case region information from the case region information set whose category label is image and text block, as the image case region information set and the text case region information set; and performing coordinate standardization processing on the image coordinate set corresponding to the image case region information set and the text position set corresponding to the text case region information set to obtain a standardized image coordinate set and a text position set. A standardized text coordinate set is generated; based on the standardized image coordinate set and the standardized text coordinate set, a spatial proximity value set is determined; based on the spatial proximity value set, an image-text positional relationship graph is constructed; the image-text positional relationship graph is used to predict association probabilities to obtain an image-text association probability set; based on the image-text association probability set, the medical segmentation case to which each image case region information in the image case region information set belongs is determined as the target medical segmentation case, thus obtaining a target medical segmentation case set; image type identification is performed on each medical image included in each target medical segmentation case in the target medical segmentation case set to obtain a medical image category label information set. The multimodal information association unit performs multimodal information association on the structured information set of the case text and the medical image category label information set to obtain a multimodal synthetic medical case information set, and stores the multimodal synthetic medical case information set to obtain a multimodal case database.
7. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-5.
8. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Medical history information tracking method and device, computer equipment and readable storage medium
CN119314195A
Real-time data analysis and visualization method and system in big data environment
CN119396997A