A diagnosis and treatment flowchart structure restoration method and system based on a multi-modal large model

By combining a multimodal large model with task-specific prompts, the problems of manual dependence and high cost in the structured reconstruction of diagnosis and treatment flowcharts are solved, achieving efficient and accurate flowchart structure reconstruction, adapting to flowcharts of different styles, and possessing self-optimization capabilities.

CN121075664BActive Publication Date: 2026-04-17HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
Filing Date
2025-11-07
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, the digitization of diagnosis and treatment flowcharts mainly relies on manual methods, which consumes a lot of manpower and resources and is prone to errors. Traditional algorithms and large model training require a large amount of labeled data, resulting in low efficiency, high cost and weak universality in the structured reconstruction of diagnosis and treatment flowcharts.

Method used

By employing a multimodal large model combined with task-specific prompts and loss functions, and through preprocessing, fine-tuning training, and review and editing, the automatic structural reconstruction of the diagnosis and treatment flowchart is achieved.

Benefits of technology

With limited data, the model achieved efficient and accurate structural reconstruction of the diagnosis and treatment flowchart, reducing manual intervention, improving work efficiency and accuracy, and continuously optimizing model performance through iterative learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075664B_ABST
    Figure CN121075664B_ABST
Patent Text Reader

Abstract

This invention provides a method for restoring the structure of a diagnosis and treatment flowchart based on a multimodal large model. The method includes: preprocessing the diagnosis and treatment flowchart; establishing a standardized structure format and prompt templates for the flowchart; designing a graph structuring loss function for the multimodal large model; fine-tuning the multimodal large model to restore the flowchart structure; inputting the flowchart and prompt templates into the fine-tuned model, outputting structured content conforming to the flowchart format specifications; restoring the structured content to a standardized, clear flowchart; and storing the structured content output by the multimodal large model, the edited structured content, and the corresponding flowchart in a database. This invention can automatically extract the node content and its connections from the input diagnosis and treatment flowchart image and convert it into a standardized, clear flowchart structure, achieving digital restoration of the diagnosis and treatment process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical information processing and knowledge extraction technology, and in particular to a method and system for restoring the structure of a diagnosis and treatment flowchart based on a multimodal large model. Background Technology

[0002] With the continuous improvement of computer hardware performance, artificial intelligence (AI) technology is showing increasingly broad application prospects, and its practical application in various industries has become a hot research topic. By combining AI technology with business needs, workflows can be optimized and work efficiency improved. Currently, AI technology is widely used in fields such as healthcare, finance, and education.

[0003] In the medical field, clinical diagnosis and treatment processes are often presented in the form of flowcharts to guide doctors' decision-making and standardize the treatment process. These flowcharts are widely found in clinical guidelines, medical textbooks, and hospital regulations, covering decision-making points, examination items, medication regimens, and the interrelationships between these processes in disease diagnosis and treatment. Standardizing and digitizing medical processes is crucial for medical quality control, clinical decision support, and big data analysis. However, currently, many flowcharts exist as images or unstructured documents, and the styles and standards for flowcharts drawn by different medical institutions or departments are inconsistent. This makes it difficult for computers to directly read and utilize this valuable process knowledge, hindering data archiving and retrieval, and impeding the sharing and exploration of medical knowledge.

[0004] In current technologies, the digitization of diagnostic and treatment flowcharts primarily relies on manual methods: professionals read the flowcharts and manually convert them into text or tabular form before entering them into information systems. This method is not only resource-intensive but also susceptible to differences in human understanding, carrying the risk of omissions or errors. Furthermore, there have been research attempts to use Optical Character Recognition (OCR) to extract textual information from flowcharts and then attempt to reconstruct the flowchart structure using rule-based algorithms. However, due to the diverse layouts and complex node connections of flowcharts, traditional algorithms struggle to reliably reconstruct the underlying hierarchical structure and logical conditions. While deep learning-based solutions have made progress in areas such as image recognition, training these models often requires large amounts of labeled flowchart data. Acquiring and labeling this data is costly, and the significant differences in flowchart formats from different sources result in weak model versatility, making it difficult to adapt to diverse real-world application scenarios. Overall, there is currently a lack of efficient and intelligent automated tools for the structured reconstruction of diagnostic and treatment flowcharts.

[0005] With the development of multimodal large-scale model technology, the possibility of using large-scale pre-trained models to understand images and text has been seen. For example, next-generation large language models combined with visual modules can jointly process image content and text prompts, demonstrating cross-modal reasoning and understanding capabilities. This provides a new approach for the automatic parsing of diagnostic flowcharts. However, to date, there are no mature methods or systems on the market for automatically reconstructing the structure of diagnostic flowcharts based on large models. Although large models possess general knowledge and powerful language-visual understanding capabilities, their direct application to flowchart structure extraction has limited effectiveness if not tailored to specific tasks. In particular, large models require explicit prompts and appropriate fine-tuning training to output results that conform to the expected structural format. Therefore, it is necessary to provide a new technical solution that fully utilizes multimodal large models to achieve automated and accurate reconstruction of diagnostic flowchart structures without the need for massive amounts of labeled data. Summary of the Invention

[0006] To address the issues mentioned in the background section regarding the high reliance on manual labor, time and manpower in the existing diagnostic and treatment flowchart structure reconstruction process, and the high cost and weak versatility of traditional model training, which makes it difficult to quickly and efficiently convert diagnostic and treatment flowcharts into standardized structural information, this invention provides a diagnostic and treatment flowchart structure reconstruction method and system based on a multimodal large model. This method overcomes the above shortcomings, achieves intelligent and efficient structure reconstruction of diagnostic and treatment flowcharts, reduces labor costs, and improves work efficiency and accuracy.

[0007] To achieve the above objectives, in a first aspect, the present invention provides a method for reconstructing a diagnostic and treatment flowchart structure based on a multimodal large model, comprising the following steps:

[0008] Step S1: Input the diagnosis and treatment flowchart and preprocess the input flowchart; improve the quality of the flowchart through image processing technology to make it suitable for subsequent structure extraction.

[0009] Step S2: Develop a standardized format for the diagnostic and treatment flowchart structure and a template for prompt statements; predefine a standard format for representing the flowchart structure, as well as a template for prompt statements to guide the extraction of the flowchart structure from the large model.

[0010] Step S3: The user designs a graph-structured loss function for the multimodal large model based on the characteristics of the task, fine-tunes the diagnosis and treatment flowchart structure to restore the multimodal large model; and performs fine-tuning training of the model in combination with the characteristics of the task, so that the large model has the ability to recognize and output the flowchart structure.

[0011] Step S4: The pre-processed diagnosis and treatment flowchart and prompt statement template are used as image and text information, respectively. The fine-tuned diagnosis and treatment flowchart structure is input into the multimodal large model. The multimodal large model outputs structured content that conforms to the diagnosis and treatment flowchart structure format specification. The large model generates standardized text or symbol representations describing the flowchart structure based on the image and prompt information.

[0012] Step S5: Reconstruct the diagnosis and treatment flowchart structure of the structured content output by the multimodal large model, re-render it into a standardized and clear flowchart for visualization; draw a standardized and aesthetically pleasing flowchart based on the structural information output by the model for users to check and use.

[0013] Step S6: Compare the diagnosis and treatment flowchart with the restored visualization results, review and edit the structured content output by the multimodal large model, and store the structured content output by the multimodal large model, the edited structured content, and the corresponding diagnosis and treatment flowchart into the database for further optimization and iteration; manually check and correct the structural information output by the model, and store the model output, the corrected final structure, and the corresponding original flowchart data for subsequent model improvement.

[0014] Furthermore, the preprocessing in step S1 may include operations such as cropping, quality enhancement, size normalization, distortion correction, and contrast adjustment of the flowchart image.

[0015] For example, in step S11, the effective process area in the diagnosis and treatment flowchart is automatically detected and cropped to extract the effective area in the flowchart image that actually carries the process content, and irrelevant background is removed.

[0016] Step S12: Enhance the image quality of the cropped treatment flowchart, including denoising the cropped treatment flowchart and using image super-resolution technology to improve the clarity of the text areas in the treatment flowchart to ensure the recognizability of key text information.

[0017] Step S13 involves normalizing the size of the enhanced image to correct distortions caused by shooting or scanning, and adjusting the image contrast to make it a high-quality, standardized image that meets standard input requirements. This preprocessing maximizes the image quality and standardization of the original diagnostic flowchart, laying the foundation for the correct parsing of subsequent large-scale models.

[0018] Furthermore, the diagnostic flowchart structure format specification established in step S2 can be customized according to user needs. For example, the user can define the structure format as a triple of "(parent node, child node, connection condition)" or other suitable forms to uniformly represent the node relationships and condition information in the flowchart.

[0019] The prompt statement template is used to explain the rules and format requirements that need to be followed when extracting the flowchart structure from a large model, and can be customized by the user according to the characteristics of the task. For example, the prompt statement template may include the following rule descriptions: "Please note that a node may be connected to multiple nodes"; "The text on the directed edge represents the judgment condition, and the condition may also appear in the form of a diamond node", etc.

[0020] By establishing clear structural format specifications and corresponding prompt statement templates, it is ensured that large models have clear output requirements during the inference process, thereby improving the accuracy and consistency of their structured extraction results.

[0021] Furthermore, the multimodal large model fine-tuning process in step S3 includes data labeling, loss function design, and model training.

[0022] First, in the data annotation stage, after obtaining the original diagnostic and treatment flowchart data, training data is constructed based on several original diagnostic and treatment flowchart data. This can be done through a combination of model-assisted annotation and manual annotation. The annotation content includes identifying and marking the content of each node in the flowchart (such as diagnostic and treatment steps, judgment conditions, etc.), as well as the connection relationships between nodes and the corresponding condition content.

[0023] Secondly, in the loss design phase, a suitable loss function is designed or selected based on the characteristics of the diagnostic flowchart structure extraction task to guide model learning. The graph structuring loss function can consist of two parts: graph structure loss and text content loss. One part measures the accuracy of the flowchart topology (nodes and connections), such as using graph matching or relationship accuracy metrics; the other part measures the accuracy of extracting node text content and conditional text, such as using sequence matching or cross-entropy loss. The two parts of the loss can be weighted according to requirements to balance the model's learning of structure and content.

[0024] Finally, during the model training phase, a large multimodal model supporting both image and text input is selected as the base (it can be open-source or closed-source), and fine-tuned using the aforementioned annotations and a customized loss function. Through supervised fine-tuning, the large model gradually acquires the ability to reconstruct the diagnostic flowchart structure, achieving good task performance even with limited data.

[0025] Furthermore, in step S6, when reviewing and editing the structured content output by the model, four types of operations are included: adding, deleting, modifying, and merging.

[0026] Added: For any missing nodes, connections, or conditions that were not identified in the multimodal large model, they will be manually added to improve the flowchart structure;

[0027] Deletion: Redundant nodes, incorrect connections, or invalid conditions that are incorrectly identified in multimodal large models are removed to avoid interfering with the correct process structure;

[0028] Modifications: Adjustments and corrections are made to inaccuracies in node content, connection relationships, or condition descriptions in the output of multimodal large models, such as correcting node text errors and updating connection condition expressions.

[0029] Merging: Merging nodes or semantically equivalent conditional information that are repeatedly identified in a multimodal large model to eliminate duplication and improve the simplicity of the structural representation.

[0030] After the above review and editing, the final accurate structured content of the diagnosis and treatment flowchart is obtained. Subsequently, the structured content of the model's original output, the manually edited and corrected content, and the corresponding original diagnosis and treatment flowchart image are stored together in the database. On the one hand, this data can be used to trace the differences between the model output and the actual process, facilitating future analysis of model shortcomings; on the other hand, the accumulated correction data can serve as new training samples, utilized in subsequent fine-tuning of the larger model, continuously optimizing model performance and achieving iterative improvement. This establishes a closed-loop self-optimization mechanism, making the system increasingly intelligent and reliable with practical application.

[0031] Secondly, the present invention provides a diagnostic and treatment flowchart structure reconstruction system based on a multimodal large model, including a data annotation unit, a model training unit, and a result output unit.

[0032] The data annotation unit includes a model-assisted annotation module and a manual annotation module, used to construct a training dataset for the flowchart structure reconstruction model; the model training unit includes a loss design module and a training module, used to configure the objective function for large model training and execute the model training process; the result output unit includes a model inference module and a format output module, used to perform model inference on the new input flowchart and convert the model output into a predetermined structured format.

[0033] Specifically, the model-assisted annotation module uses pre-configured simple models or rules to perform preliminary annotations on the structural content of the original diagnosis and treatment flowchart, assisting manual annotation work.

[0034] The manual annotation module allows for manual review and modification of the model's annotation results, ensuring the accuracy and completeness of the training data.

[0035] The loss design module is used to design the loss function of the above-mentioned graph-structured multimodal large model according to the task requirements of the flowchart structure reconstruction, and to clarify the optimization objective of the multimodal large model training.

[0036] The training module inputs the labeled flowchart content into the multimodal large model and uses the designed loss function to train the model, enabling the multimodal large model to learn the ability to extract the structure of the diagnosis and treatment flowchart.

[0037] The model inference module is used to input unlabeled diagnosis and treatment flowchart images and corresponding prompts into the trained large model after deployment, and obtain the structured results output by the model.

[0038] The output format module is responsible for converting the model output results into a user-defined output format (such as the first triplet format or a visual flowchart) for user viewing or system storage.

[0039] Through the collaborative work of the above-mentioned units and modules, the automatic structured reconstruction of the diagnosis and treatment flowchart can be achieved.

[0040] Thirdly, the present invention provides a terminal device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for restoring the structure of a diagnosis and treatment flowchart based on a multimodal large model as described above. Therefore, using this terminal device, users can conveniently perform automatic parsing and structural reconstruction of diagnosis and treatment flowcharts, improving work efficiency.

[0041] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method for restoring the diagnostic and treatment flowchart structure based on a multimodal large model as described above.

[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0043] This invention, by introducing a large-scale pre-trained visual-language model and performing targeted instruction fine-tuning, achieves accurate and automatic extraction of the structure of diagnostic and treatment flowcharts with limited training data, significantly reducing model training and maintenance costs. The automated flowchart reconstruction process significantly improves efficiency, reduces manual intervention, and optimizes the digital flow of medical knowledge. Simultaneously, the manual review and iterative learning mechanism designed in this invention enables the system to continuously learn from new flowchart samples and correct errors, becoming increasingly intelligent with use and ensuring high accuracy and robustness even in complex and ever-changing real-world applications. Through these technical means, this invention effectively solves the problems of low efficiency, high cost, and error susceptibility in existing diagnostic and treatment flowchart structuring processes, and has significant implications for medical information standardization and clinical decision support. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings are briefly described below. The following drawings only show some exemplary embodiments of the present invention; those skilled in the art can obtain other forms of drawings based on these drawings without creative effort.

[0045] Figure 1 This is a flowchart of the diagnostic and treatment process flow diagram structure reconstruction method based on a multimodal large model according to the present invention;

[0046] Figure 2 This is a schematic diagram of the preprocessing steps in the diagnostic and treatment flowchart of this invention;

[0047] Figure 3 This is a schematic diagram illustrating the process of restoring the execution flowchart structure of the multimodal large model in this invention;

[0048] Figure 4 This is a flowchart illustrating the training process of the multimodal large model in this invention.

[0049] Figure 5 This is a schematic diagram of the system structure for restoring the diagnostic and treatment flowchart of the present invention. Detailed Implementation

[0050] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, those skilled in the art can combine and utilize the technical features of the various embodiments to obtain new embodiments.

[0051] Existing methods for reconstructing treatment flowcharts are highly reliant on manual intervention, time-consuming, and labor-intensive. Furthermore, traditional model training requires large amounts of labeled data, is costly, and lacks versatility, making it difficult to quickly and efficiently convert treatment flowcharts into standardized structural information. To address this, this invention proposes a method and system for reconstructing treatment flowchart structures based on a multimodal large-scale model. This method utilizes a visual-language large-scale model, combined with structured prompts designed for flowchart structure extraction tasks, to perform supervised fine-tuning of the large-scale model, enabling it to more accurately understand and execute the flowchart structure reconstruction task. The method introduces a chain-like thinking process similar to human flowchart analysis, designing the experience and steps of manually reconstructing flowchart structures as step-by-step prompts to guide the large-scale model in gradually extracting node and relationship information, thereby improving the model's understanding and extraction efficiency for complex flowcharts. In addition, this invention designs a supporting system architecture, allowing the model to learn new flowchart samples during practical use, continuously optimizing model performance and improving the method's practicality.

[0052] Applying the solution of this invention can bring several beneficial effects. First, automatically restoring the flowchart structure can significantly reduce the burden of manually compiling flowcharts and improve the efficiency of digitizing medical process information. Second, leveraging the vast knowledge and cross-modal understanding capabilities of the large model, this solution can better adapt to flowcharts of different styles, reducing reliance on manual intervention. Finally, by introducing result review and database iterative optimization mechanisms, this solution can quickly accumulate and learn new cases in practical use, continuously correct model errors, solve previously unforeseen new situations, and enhance the practicality and robustness of the method in complex and ever-changing scenarios.

[0053] like Figure 1 As shown, this embodiment provides a method for restoring the structure of a diagnosis and treatment flowchart based on a multimodal large model. This method mainly includes steps S1 to S6, which are interconnected and work together to achieve the structured parsing and restoration of the input diagnosis and treatment flowchart. The following example illustrates each step.

[0054] Let's take a disease diagnosis and treatment flowchart as an example. Assume the flowchart includes a starting node, examination and diagnosis nodes, treatment plan nodes, and several judgment conditions (e.g., different branches corresponding to normal or abnormal examination results). Currently, this flowchart only exists in image form. The purpose of this embodiment is to automatically convert it into a standardized digital flowchart structure. The entire process is as follows:

[0055] Step S1: Input the diagnostic flowchart and preprocess the input flowchart image. Image processing techniques are used to improve the quality of the flowchart image and standardize it to make it suitable for subsequent structural information extraction.

[0056] Step S2: Define the structural format specifications and prompt statement templates for the diagnostic and treatment flowchart. Predefine a standard format for representing the flowchart structure (e.g., a list of triples) and write corresponding prompt statement templates to guide the correct extraction of flowchart nodes and their relationships in the multimodal large model.

[0057] Step S3: Design a graph-structured loss function for the multimodal large model based on the characteristics of the task, and fine-tune the multimodal large model. Supervised fine-tuning is performed on the pre-trained multimodal large model using a small amount of labeled data, enabling it to recognize the structure of the diagnosis and treatment flowchart and output standardized results.

[0058] Step S4: Input the preprocessed diagnostic flowchart image and prompt statement template into the fine-tuned multimodal large model. The multimodal large model outputs structured content conforming to the stated structural format specification. The multimodal large model combines the image and prompt information to parse the flowchart and generate standardized text or symbolic representations describing the flowchart structure.

[0059] Step S5: Based on the structured content output by the multimodal large model, the flowchart structure is restored and re-rendered into a standardized and clear flowchart for visualization. The system draws a standardized and aesthetically pleasing new flowchart based on the structured content, intuitively displaying the restoration results for easy user comparison and verification.

[0060] Step S6: Compare the original diagnosis and treatment flowchart with the restored visualization results, review and edit the structured content output by the multimodal large model, and store the relevant data in the database for further optimization and iteration. Professionals proofread and correct the structural information output by the multimodal large model, record the original output and the final correction results, and use them for subsequent multimodal large model improvements to continuously improve model performance.

[0061] like Figure 2 As shown, this embodiment preprocesses the diagnostic flowchart image in step S1 to improve image quality and standardization, preparing it for subsequent structure extraction. Specifically, the preprocessing process includes the following steps:

[0062] Step S11: Detect the effective area carrying the process content in the treatment flowchart image, automatically crop out irrelevant edge backgrounds, and retain the main body of the flowchart. This removes redundant interference information and highlights the key content areas of the flowchart.

[0063] Step S12: Enhance the image quality of the cropped flowchart image. First, perform noise reduction to eliminate noise generated during scanning or photography; then, use an image super-resolution algorithm to improve the clarity of text and line areas in the image to ensure that key text information is clearly identifiable and that process nodes and connecting lines are clearly distinguishable.

[0064] Step S13: Perform size normalization and distortion correction on the enhanced image. Scale the image to a predetermined standard size, correct geometric distortions caused by factors such as shooting angle, and adjust the contrast appropriately to make the overall image clearer. After the above preprocessing operations, a high-quality standardized flowchart image is finally obtained: the text content is clearly readable, the node hierarchy and connection relationships are distinct, meeting the input requirements for subsequent multimodal large model analysis.

[0065] like Figure 3As shown, this embodiment demonstrates the process of restoring the flowchart structure using a finely tuned multimodal large model. Before starting model parsing, a unified flowchart structure format specification has been established according to step S2 (e.g., using the "(parent node, child node, connection condition)" triple format to represent node relationships), and corresponding prompt statement templates have been pre-written. These prompt templates clearly define the model extraction requirements and format rules, such as reminding the model: "Diamond nodes indicating type may have multiple output branches; text attached to arrows or connections indicates branch conditions; if there is no text on a connection, it is considered unconditional and represented by 'none'." Through such prompts, the large model can be guided to focus on the key structural information in the flowchart and organize the output content according to the specification requirements.

[0066] In this embodiment, the multimodal large model, fine-tuned and trained in step S3, is loaded and used by the model inference module of the result output unit. The clear flowchart image obtained from the preprocessing in step S1 is used as visual input, and the prompts specified in step S2 are used as text input, both fed into the trained large model. After receiving the image and prompts, the model first parses and understands the flowchart content in the image, and then extracts the nodes and their connections from the flowchart according to the rules specified in the prompts. The output of the multimodal large model is several structured triples following a preset format, as shown below:

[0067] (Beginning, Clinical examination, None)

[0068] (Clinical examination to determine if pneumonia is present; none)

[0069] (To determine if it is pneumonia, antibiotic treatment is required. "Yes")

[0070] (To determine if it is pneumonia, observe and follow up, "No")

[0071] Each triple above represents a node connection relationship and condition information in the flowchart. For example, the triple "(determine if it is pneumonia, antibiotic treatment, "yes")" means that when the judgment result of the "is it pneumonia" node is "yes", the process will enter the "antibiotic treatment" node; while "(determine if it is pneumonia, observation and follow-up, "no")" means that when the judgment result is "no", it will switch to the "observation and follow-up" branch. For connections without textual descriptions of conditions, such as a simple flow connection from the starting node to "clinical examination", the condition field is marked as none, indicating that there are no specific conditions. It should be noted that the structured content output by the multimodal large model can be in the form of a text list as shown above, or in a data format such as JSON, as long as it conforms to the structure format specifications defined in step S2. In this embodiment, the results are presented as a list of plain text triples to facilitate manual reading and verification and subsequent processing by the system. Furthermore, after the multimodal large model is output, the system's format output module performs format checks and simple organization on the results. This includes sorting output items according to the process order and removing duplicate triples to ensure the structured content conforms to specifications and is free of obvious redundancy or omissions. Based on the structured triple list output by the model, the system then executes step S5 to reconstruct the extracted process structure information into a flowchart and visualize it. The format output module calls the flowchart drawing engine to re-render a standardized flowchart graphic based on the structured content.

[0072] Specifically, this embodiment predefines a standardized style for flowchart drawing: rectangular nodes represent general operation or inspection steps, diamond nodes represent decision nodes, and the text inside the nodes uses the node content output by the model; directional arrow lines are drawn based on a triplet list to connect parent and child nodes. For each arrow line, if its triplet contains conditional text that is not "none", then that conditional text is labeled next to the arrow line. Through the above drawing rules, the system can automatically generate a new flowchart whose structure and key content correspond to the original input diagnostic flowchart, but the presentation is more neat and clear, and the symbols and layout conform to standard specifications.

[0073] use Figure 3 The schematic diagram of the large model output processing illustrates the process. The newly drawn flowchart visually demonstrates the structure obtained from the model analysis, and the restored results clearly show the nodes and relationships of the original diagnosis and treatment process. The generated standardized flowchart can be displayed on the user interface for staff to verify by comparing it with the original image. If necessary, the flowchart can also be exported as a vector file or structured data (such as XML / JSON) for storage, sharing, or further editing and use in the medical information system.

[0074] like Figure 4As shown, this embodiment elaborates on the training process of the multimodal large model in detail for step S3. This process includes three main stages: data annotation, loss function design, and model training, which are used to adjust the general pre-trained multimodal large model into a specialized model capable of performing the task of restoring the structure of the diagnosis and treatment flowchart.

[0075] The first stage is data annotation. Using the system's data annotation unit, training samples are prepared from existing diagnostic and treatment flowchart data. In this embodiment, a certain number of clinical guideline flowcharts or historical case flowcharts can be selected as raw data. The model-assisted annotation module performs preliminary structural analysis on some flowcharts: for example, using simple image analysis algorithms to detect text boxes and connecting lines in the flowchart, generating a preliminary list of nodes and relationships. Then, the manual annotation module reviews, corrects, and supplements the model-assisted results. Through this process, structured annotation data is constructed for each flowchart, including a list of nodes and a set of connections between nodes. The node list lists the standard name or content of each node in the flowchart; the connections are recorded in the form of "triples (parent node, child node, condition)," with each pair of parent and child nodes and their connection condition as an annotation entry. If some nodes are decision conditions (usually represented by a diamond), then the decision node is considered the parent node, and its multiple outward arrows are recorded as different condition branches as multiple triples. Through the above manual annotation, several large model input-output pairs were obtained for training, where the input is a flowchart image and corresponding prompts, and the output is a list of manually annotated standard triples.

[0076] Next is the loss function design phase. The loss design module in the model training unit customizes a suitable loss function to optimize model performance based on the requirements of the flowchart structure extraction task. This embodiment employs a combined loss scheme that includes structural loss and text loss. Structural loss measures the consistency between the model's output relational structure and the manually labeled standard answer. It can be calculated by comparing the nodes and connections extracted by the model with the standard triplet set. For example, metrics such as Precision, Recall, and F1 are used to evaluate the model's accuracy and recall in structure extraction. Text loss measures the degree of matching between the node names and conditional text output by the model and the standard answer text. Sequence-level cross-entropy loss or edit distance-based loss functions can be used to measure the difference between the output text and the standard answer text. During training, the structural loss and text loss are weighted and synthesized into a total loss function according to predetermined weights to jointly guide model parameter updates. Through a carefully designed loss function, the model can simultaneously identify the correct flowchart structure relationships and accurately extract node text content during training, thus learning the task's objectives more comprehensively.

[0077] Finally, the model training phase takes place. The training module uses the prepared labeled data and defined loss function to fine-tune the multimodal large model. Specifically, each training sample pair (i.e., a diagnosis flowchart image and its corresponding prompts, and a manually labeled list of structured triples) is input into the model for forward inference. The model's predicted output structure is obtained, compared with the standard answer to calculate the loss, and then the model parameters are adjusted using the backpropagation algorithm. This process is iterated continuously. After several rounds of training, the model's performance gradually converges, and it learns the knowledge and ability required to reconstruct the diagnosis flowchart structure. It should be noted that this invention does not limit the large model platform used; fine-tuning training can be performed on open-source multimodal pre-trained models (such as visual language models) or on a self-developed multimodal large model. The trained dedicated multimodal large model will be used to perform the flowchart structure parsing task in step S4, becoming the core inference engine of this invention.

[0078] Please see Figure 5 This embodiment illustrates the structural composition of the diagnostic and treatment flowchart reconstruction system of the present invention. The system mainly includes a data annotation unit, a model training unit, and a result output unit. These units work collaboratively to achieve automatic structural reconstruction of the diagnostic and treatment flowchart and iterative optimization of the results.

[0079] The data annotation unit comprises a model-assisted annotation module and a manual annotation module, used to prepare labeled datasets before model training. The model-assisted annotation module can use pre-configured simple models or rules to perform preliminary structural content annotation on the original diagnosis and treatment flowchart, automatically extracting preliminary results of node text and connection relationships, reducing manual workload; the manual annotation module provides an interactive interface for professionals to manually review and improve the annotations output by the model-assisted annotation, ensuring the accuracy and comprehensiveness of the training data.

[0080] The model training unit comprises a loss design module and a training module. The loss design module performs the loss function customization work described in step S3 above, designing a suitable combination of graph-structured loss functions based on task requirements to guide the model towards correctly extracting the flowchart structure and content. The training module is responsible for actually executing the training process of the multimodal large model: reading the labeled training dataset and loss function settings, inputting the data into the large model for training iterations, continuously adjusting the model parameters, and finally obtaining a fine-tuned model specifically for the flowchart structure reconstruction task.

[0081] The output unit comprises a model inference module and a format output module. The model inference module carries a trained multimodal large model. During deployment and runtime, it receives new unlabeled diagnostic flowchart images and corresponding prompts as input, performs large model inference to output structured content (i.e., a list of triples representing process nodes and relationships). The format output module converts the structured content generated by the model inference module into a user-defined output format. On one hand, according to the requirements of step S5, this module redraws the structural information output by the model into a standardized flowchart for user viewing; on the other hand, it is also responsible for organizing the structured content into standard data formats (such as JSON, XML, or database records) and performing basic format validation and optimization (such as removing duplicate entries and sorting normalization) to ensure that the output results are readable, accurate, and conform to the established structural format specifications.

[0082] Furthermore, the system of this invention also includes a result review and iterative optimization mechanism, corresponding to step S6 above. Specifically, after the result output unit generates the preliminary structured result and the restored flowchart, the system provides a manual review interface for professionals to verify and edit the results. Reviewers can compare the generated standardized flowchart with the original flowchart image one by one to check whether each node and connection relationship extracted by the model is correct. For any discrepancies found, the structured content can be edited and corrected through the following operations: adding missing nodes or connections, deleting incorrectly identified redundant nodes, incorrect connections, or invalid conditions, modifying inaccurate node names or condition descriptions, and merging the same node or equivalent condition repeatedly identified by the model. Through editing operations such as "adding, deleting, modifying, and merging," the model output result is comprehensively improved, ultimately obtaining a structured flowchart data that is completely consistent with and correct of the original flowchart. Subsequently, the system stores the original model output and the final structured result after manual correction together in the background database and associates it with the corresponding original flowchart image data. As the cases processed by the system accumulate, the database will accumulate a large number of real flowcharts and their structured results, as well as information on model errors and manual corrections. This historical data can be used to periodically update the model training set, triggering a new round of model fine-tuning training. This allows the model to gradually learn from previous errors, optimize parameters to reduce the occurrence of similar errors in the future, and further improve the model's adaptability and accuracy to complex and diverse flowcharts. Through the above feedback and iteration mechanism, the system of this invention can continuously evolve in long-term use, achieving adaptive learning and significantly improving the practicality and robustness of the diagnostic flowchart structure reconstruction method.

[0083] This invention also provides a terminal device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for restoring the structure of a diagnosis and treatment flowchart based on a multimodal large model as described above. Therefore, using this terminal device, users can conveniently perform automatic parsing and structural reconstruction of diagnosis and treatment flowcharts, improving work efficiency.

[0084] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method for restoring the diagnostic and treatment flowchart structure based on a multimodal large model as described above.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for reconstructing the diagnostic and treatment flowchart structure based on a multimodal large model, comprising the following steps: Step S1: Input the diagnosis and treatment flowchart, and preprocess the input flowchart. Methods include: Step S11: Automatically detect and trim the effective process areas in the diagnosis and treatment flowchart; Step S12: Enhance the image quality of the cropped diagnosis and treatment flowchart, including denoising the cropped diagnosis and treatment flowchart and using image super-resolution technology to improve the clarity of the text areas in the diagnosis and treatment flowchart to ensure the recognizability of key text information. Step S13: Perform size normalization, distortion correction, and contrast adjustment on the enhanced diagnostic flowchart to form a high-quality standardized image after preprocessing; Step S2: Develop standardized formats and prompt templates for the diagnostic and treatment flowchart; The specified format for the diagnostic and treatment flowchart is custom-built according to user needs; the prompt statement template is used to explain the flowchart structure extraction rules in the prompt statements of the large model, and is customized by the user. Step S3: The user designs a graph-structured loss function for the multimodal large model based on the characteristics of the task, and fine-tunes the diagnosis and treatment flowchart structure to restore the multimodal large model; The fine-tuning process of the multimodal large model includes data labeling, loss design, and model training. Data annotation involves, after obtaining the original diagnostic and treatment flowchart data, performing model-assisted annotation or manual annotation on the structural content of these flowcharts. The annotation content includes the content of the flowchart nodes, the connection relationships and connection conditions between the flowchart nodes. Loss design: The graph structured loss function includes two parts: graph structure loss and text content loss. The loss functions and their weights for the two parts are specified according to the requirements. Model training uses an open-source or closed-source multimodal large model as a base, and fine-tunes it using the above annotations to enable the multimodal large model to acquire the capabilities for this task. Step S4: Use the pre-processed diagnosis and treatment flowchart and prompt statement template as image and text information respectively, input the fine-tuned diagnosis and treatment flowchart structure to restore the multimodal large model, and the multimodal large model outputs structured content that conforms to the diagnosis and treatment flowchart structure format specification; Step S5: Reconstruct the diagnosis and treatment flowchart structure of the structured content output by the multimodal large model, and re-render it into a standardized and clear flowchart for visualization results; Step S6: Compare the diagnosis and treatment flowchart with the restored visualization results, review and edit the structured content output by the multimodal large model, and store the structured content output by the multimodal large model, the edited structured content, and the corresponding diagnosis and treatment flowchart into the database; Review and edit the structured content output by the multimodal large model, including adding, deleting, modifying, and merging: Added: Manually add missing nodes, connections, or conditions that were not identified in the multimodal large model; Delete: Remove redundant nodes, erroneous connections, or invalid conditions that are incorrectly identified in the multimodal large model; Edit: Correct any errors or inaccuracies in the description of node content, connection relationships, or conditions; Merging: Merging nodes that are repeatedly identified in a multimodal large model or conditions with the same semantics. 2.A multi-modal large model-based diagnosis and treatment flowchart structure restoration system, characterized in that, It includes a data annotation unit, a model training unit, and a result output unit; The data annotation unit includes a model-assisted annotation module and a manual annotation module; the model training unit includes a loss design module and a training module; and the result output unit includes a model inference module and a format output module. The model-assisted annotation module is used to perform preliminary annotation of the structural reconstruction of several original flowcharts using a custom model, thus assisting the manual annotation process. The manual annotation module is used for manual review and modification of auxiliary annotations on the model; Loss Design Module: Used to design the loss function for multimodal large models, guiding the training objectives of multimodal large models; Training module: Used to feed the labeled flowchart content into the multimodal large model and train the model using the designed loss function; Model Inference Module: Used to feed unlabeled flowchart content into the multimodal large model and use the trained multimodal large model to perform inference output; Formatted output module: Used to convert the output results of multimodal large models into a user-defined output format.

3. A terminal device, characterized by comprising: The terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for restoring the diagnostic and treatment flowchart structure based on a multimodal large model as described in claim 1.

4. A computer-readable storage medium having stored thereon a computer program, characterized in that When the computer program is executed by the processor, it implements the method for restoring the diagnostic flowchart structure based on a multimodal large model as described in claim 1.

Citation Information

Patent Citations

  • Industrial field multi-mode large model fine tuning method, device, equipment and medium

    CN118070778A

  • Flow chart image analysis and structured reconstruction method and device, and storage medium

    CN120808375A