Multimodal medical data synthesis method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610805271.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-06-05
AI Technical Summary
[0004]本申请提供一种多模态医疗数据合成方法、装置、电子设备及存储介质,用以解决现有技术中问答类型单一、提问角度有限、与医疗实践脱节的缺陷
[0017]本申请还提供一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现如上述任一种所述多模态医疗数据合成方法。
Smart Images

Figure CN122337676B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device and storage medium for multimodal medical data synthesis. Background Technology
[0002] A key objective of building a large-scale medical industry model—high-quality data accumulation—is to develop multimodal capabilities within the medical industry. This includes building a medical image question-answering model with high-quality medical image question-answering capabilities. Therefore, it is necessary to acquire large-scale, high-quality Chinese medical image question-answering pairs that are closely aligned with clinical scenarios as training data to train the medical image question-answering model.
[0003] Currently, the method for obtaining medical image question-and-answer pairs for training medical image question-and-answer models is to directly construct medical image question-and-answer pairs using medical images and image descriptions in medical imaging literature. This method fails to fully utilize the image information in medical imaging literature to generate rich medical image question-and-answer pairs, resulting in problems such as limited question and answer types, limited question angles, and disconnect from medical practice. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, and storage medium for synthesizing multimodal medical data, in order to address the shortcomings of existing technologies, such as limited question-and-answer types, limited question angles, and disconnect from medical practice.
[0005] This application provides a method for synthesizing multimodal medical data, including: Retrieve medical image text pairs; The first prompt word is determined based on the medical role type of the first intelligent agent, the meta-attribute of the medical question, and the medical image-text pair; The first prompt word is input into the first intelligent agent to obtain the medical problem data output by the first intelligent agent; Based on the medical problem data and the medical image text pair, a second prompt word is determined; The second prompt word is input into the second agent to obtain the medical image question-and-answer pair output by the second agent.
[0006] According to the multimodal medical data synthesis method provided in this application, the step of determining a first prompt word based on the medical role type of the first intelligent agent, the medical question meta-attribute, and the medical image-text pair includes: determining role preference data based on the medical role type; determining user question data based on the medical question meta-attribute; and combining the role preference data, the user question data, and the medical image-text pair according to the first prompt word template to obtain the first prompt word.
[0007] According to the multimodal medical data synthesis method provided in this application, the second intelligent agent is a medical expert intelligent agent; the step of determining the second prompt word based on the medical question data and the medical image text pair includes: determining the respondent preference data based on the medical expert intelligent agent; determining the questioner preference data based on the medical role type; and combining the respondent preference data, the questioner preference data, the medical question data, and the medical image text pair according to the second prompt word template to obtain the second prompt word.
[0008] According to the multimodal medical data synthesis method provided in this application, the medical problem meta-attributes include medical problem type and medical scenario.
[0009] According to a multimodal medical data synthesis method provided in this application, the step of obtaining medical image-text pairs includes: obtaining composite medical image-text pairs; the composite medical image-text pairs include composite medical images and associated text data; the composite medical images include multiple medical sub-images; the associated text data includes a contextual description of the composite medical images and a chapter title of a medical imaging document; performing sub-image segmentation on the composite medical images in the composite medical image-text pairs to obtain multiple medical sub-images; performing semantic segmentation on the associated text data in the composite medical image-text pairs to obtain multiple sub-title text descriptions and multiple image body text descriptions; performing image-text alignment on the multiple medical sub-images, the multiple sub-title text descriptions, and the multiple image body text descriptions to obtain multiple aligned single image-text pairs; each aligned single image-text pair includes one medical sub-image and its corresponding aligned text data; and determining the medical image-text pairs based on the aligned single image-text pairs.
[0010] According to a multimodal medical data synthesis method provided in this application, the step of aligning multiple medical sub-images, multiple sub-title text descriptions, and multiple image text descriptions to obtain multiple aligned single image-text pairs includes: inputting the multiple medical sub-images, multiple sub-title text descriptions, and multiple image text descriptions into an image-text alignment model to obtain the initial similarity of the multiple single image-text pairs output by the image-text alignment model; each single image-text pair includes any one of the medical sub-images, any one of the sub-title text descriptions, and any one of the image text descriptions; for each single image-text pair... Yes, based on the first sequential index of the medical sub-image in the single image-text pair and the second sequential index of the subtitle text description in the single image-text pair, a penalty factor is determined, and the initial similarity of the single image-text pair is adjusted based on the penalty factor to obtain the corrected similarity of the single image-text pair, thereby determining the corrected similarity of all single image-text pairs; for each medical sub-image, a set of single image-text pairs containing the medical sub-image is determined, and based on the single image-text pair with the highest corrected similarity in the set of single image-text pairs, the aligned single image-text pairs of the medical sub-image are determined, thereby determining the aligned single image-text pairs of all medical sub-images.
[0011] According to a multimodal medical data synthesis method provided in this application, the semantic segmentation of the associated text data in the composite medical image-text pair to obtain multiple subheading text descriptions and multiple image body descriptions includes: performing preliminary segmentation of the associated text data based on a preset regular expression to obtain preliminary segmented text; inputting the preliminary segmented text and its contextual validation content into a lightweight language model to obtain the segmentation boundary confidence score output by the lightweight language model, and determining semantically reasonable segmented text based on the preliminary segmented text whose segmentation boundary confidence score is greater than a preset segmentation confidence score threshold; determining a third prompt word based on the semantically reasonable segmented text and a third prompt word template, and inputting the third prompt word into a large model arbitration layer to obtain multiple subheading text descriptions and multiple image body descriptions output by the large model arbitration layer.
[0012] According to the multimodal medical data synthesis method provided in this application, the step of performing sub-image segmentation on the composite medical image in the composite medical image-text pair to obtain multiple medical sub-images includes: inputting the composite medical image in the composite medical image-text pair into a sub-image segmentation model to obtain multiple medical sub-images output by the sub-image segmentation model; wherein, the sub-image segmentation model is implemented based on an improved YOLO model; the feature extraction network of the improved YOLO model introduces a self-attention mechanism; the prediction head of the improved YOLO model introduces a multi-scale feature fusion module; the loss function of the improved YOLO model is determined based on a bounding box coordinate loss term, an object presence confidence loss term, a classification loss term, and a self-attention regularization term.
[0013] According to a multimodal medical data synthesis method provided in this application, the step of obtaining medical image-text pairs includes: obtaining a structured data file of medical imaging literature; locating a visual image in the structured data file and obtaining a contextual description of the visual image; the visual image includes at least one sub-image; using a large language model to perform image-text association processing on the visual image, the contextual description of the visual image, and the chapter title of the medical imaging literature to obtain image-text pairs; filtering out the medical image-text pairs from the image-text pairs; wherein, the medical image-text pairs include single medical image-text pairs and composite medical image-text pairs; the composite medical image-text pairs include composite medical images and associated text information; the composite medical images include multiple medical sub-images; the associated text information includes the contextual description and the chapter title.
[0014] This application also provides a multimodal medical data synthesis device, comprising: The image-text pair acquisition module is used to acquire medical image-text pairs; The first prompt word determination module is used to determine the first prompt word based on the medical role type of the first intelligent agent, the medical problem meta-attribute, and the medical image text pair; The medical problem acquisition module is used to input the first prompt word into the first intelligent agent and obtain the medical problem data output by the first intelligent agent; The second prompt word determination module is used to determine the second prompt word based on the medical problem data and the medical image text pair; The question-and-answer pair acquisition module is used to input the second prompt word into the second agent and obtain the medical image question-and-answer pair output by the second agent.
[0015] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multimodal medical data synthesis method described above.
[0016] This application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multimodal medical data synthesis method as described above.
[0017] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the multimodal medical data synthesis method described above.
[0018] The multimodal medical data synthesis method, apparatus, electronic device, and storage medium provided in this application, through the design of a multi-agent data synthesis collaborative architecture, firstly utilizes a first agent to expand the semantic range of image understanding of medical image-text pairs by constructing a multi-dimensional hybrid matrix according to different medical role types and different medical question meta-attributes, thereby generating medical question data. Then, a second agent is used to generate text from the medical image text and medical question data to output medical image question-and-answer pairs. This can automatically, scalably, and systematically generate medical image question-and-answer pairs with multiple roles, multiple scenarios, and multiple dimensions of professional questions based on the same medical image-text pairs under different medical roles and different medical question meta-attributes. It realizes the generation of medical image question-and-answer pairs with diverse question-and-answer types and questioning angles, which can meet the needs of high-quality training data for medical image question-and-answer model training, significantly reduce the data threshold and cost of model training, and make the model trained using medical image question-and-answer pairs more accurate and reliable in understanding complex clinical scenarios and answering professional questions. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating the multimodal medical data synthesis method provided in this application.
[0021] Figure 2 This is one of the example images provided in this application for generating the first prompt word.
[0022] Figure 3 This is the second example image provided in this application for generating the first prompt word.
[0023] Figure 4 This is an example image provided in this application for generating a second prompt word.
[0024] Figure 5 This is a schematic diagram of the process for obtaining image-text pairs provided in this application.
[0025] Figure 6 This is a schematic diagram of the process for obtaining medical image text pairs provided in this application.
[0026] Figure 7 This is a flowchart illustrating the subgraph segmentation process provided in this application.
[0027] Figure 8 This is a schematic diagram of the multimodal medical data synthesis device provided in this application.
[0028] Figure 9 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0030] The following is combined with Figures 1 to 9 This application describes a method, apparatus, electronic device, and storage medium for synthesizing multimodal medical data.
[0031] Multimodal data synthesis is a key technology in the fields of artificial intelligence and machine learning. Through algorithmic models, it can simultaneously process and integrate data types from different sources (such as images, text, audio, video, and sensor signals), and learn the complex relationships and joint distributions among them to generate artificial data that integrates multiple modal information.
[0032] A multi-agent system (MAS) is a distributed system composed of multiple autonomous agents (autonomous software entities) capable of acting autonomously, communicating and cooperating with each other. Each agent possesses specific expertise and solves complex and comprehensive problems through division of labor and cooperation. Its core characteristics lie in the autonomy and cooperation of the agents.
[0033] Figure 1 This is a flowchart illustrating the multimodal medical data synthesis method provided in this application, as shown below. Figure 1 As shown, the multimodal medical data synthesis method includes, but is not limited to, steps 101 to 105.
[0034] It should be noted that the execution subject of the multimodal medical data synthesis method provided in this application can be a server, computer equipment, such as mobile phone, tablet computer, laptop computer, handheld computer, vehicle electronic equipment, wearable device, ultra-mobile personal computer (UMPC), netbook or personal digital assistant (PDA), etc.
[0035] Step 101: Obtain medical image text pairs.
[0036] Medical image text pairs include medical images and their corresponding image descriptions; the image descriptions are textual information used to describe the medical images.
[0037] Medical image-text pairs can be any type of image-text pair, such as composite medical image-text pairs or aligned single image-text pairs.
[0038] The medical image in a composite medical image text pair is a composite medical image that includes multiple medical sub-images. Correspondingly, the image description in the composite medical image text pair includes text information describing the multiple medical sub-images (such as the context description and chapter title of the medical sub-image).
[0039] The medical image in an aligned single image-text pair is a single medical sub-image, and the corresponding image description in the aligned single image-text pair only includes the text information of that single medical sub-image.
[0040] Medical images include, but are not limited to, MRI images of the effects on internal organs and tissues of the human body, CT image slices, and pathological slices observing pathological changes in cells and tissues.
[0041] Specifically, multi-source and heterogeneous medical imaging literature data are collected from professional journals and websites such as radiology and medical imaging. Medical images and their corresponding descriptive medical texts are extracted from the medical imaging literature data according to a predetermined medical image text extraction method. The resulting medical image text pairs are used as seed data for generating medical image question-and-answer pairs for multimodal medical data.
[0042] Aligning single-image-text pairs and composite medical image-text pairs containing richer context together constitutes a high-quality, highly reliable seed dataset, laying the foundation for generating rich medical image question-and-answer pairs under a multi-agent architecture.
[0043] Step 102: Determine the first prompt word based on the medical role type of the first intelligent agent, the medical problem meta-attribute, and the medical image text pair.
[0044] The first intelligent agent, also known as the role-based intelligent agent, is an intelligent agent used to generate medical problem data based on different medical roles.
[0045] The medical role type is a pre-defined type of medical role for the first intelligent agent.
[0046] For example, the medical role type of the first intelligent agent can be any one of the role types such as medical student, patient, and emergency responder; correspondingly, the first intelligent agent can be any one of the intelligent agents such as medical student agent, patient agent, and emergency responder agent.
[0047] Medical issue meta-attributes are pre-defined meta-attributes used to expand the semantic space when generating medical issue data. For example, medical issue meta-attributes include, but are not limited to, several attributes such as medical issue type, medical scenario, and treatment time.
[0048] As an optional implementation, the medical problem meta-attributes include medical problem type and medical scenario.
[0049] The medical question type is used to determine the type of medical question data generated by the first agent; the medical question type can be any of the following types: open-ended questions (such as "How to interpret CT results?"), closed-ended questions (such as "Do I need to see a doctor immediately?"), multiple-choice questions (such as choosing the most likely answer to the imaging results), case analysis (such as "Give a diagnosis based on medical records"), and multi-turn dialogues (such as asking for symptom details based on imaging results).
[0050] The medical scenario is used to determine the diagnosis and treatment scenario to which the medical problem data generated by the first intelligent agent belongs; the medical scenario can be any of the scenarios such as in-hospital diagnosis and treatment scenario and post-discharge follow-up scenario.
[0051] Optionally, the meta-attributes of medical issues are either pre-set or determined based on real-time user input.
[0052] Specifically, the medical role type and medical problem meta-attributes of the first intelligent agent used to generate medical problem data are extracted from the user's real-time input. When determining the first prompt word used to generate medical problem data, prompt word generation processing is performed based on the medical role type, medical problem meta-attributes, and the acquired medical image text pairs to obtain the first prompt word.
[0053] Understandably, the meta-attributes of medical issues and the medical role type of the first agent together constitute a combination based on different [roles, scenarios, problem types] hybrid matrices, which expands the semantic space, broadens the scope of synthetic data issues, and increases the degree of semantic generalization.
[0054] By scheduling first agents of different medical role types and constructing prompt words for different medical scenarios and different types of medical problems, medical image-text pairs corresponding to the questions can be generated as seed data.
[0055] By adding medical role types to expand the role of the first intelligent agent, increasing the types of medical scenarios, and adding medical question meta-attributes, the professional scope of question answering can be expanded.
[0056] Table 1 is an example table of the medical role type of the first intelligent agent provided in this application when it is a medical student. As shown in Table 1, when the medical scenario includes in-hospital diagnosis and treatment scenario or post-hospital follow-up scenario, and the medical question type includes open-ended questions or multiple-choice questions, the mixed matrix includes [medical student, scenario 1, question type 1], [medical student, scenario 2, question type 1], [medical student, scenario 1, question type 2] and [medical student, scenario 2, question type 2].
[0057] Table 1
[0058] Table 2 is an example table of the medical role type of the first intelligent agent provided in this application when it is a patient. As shown in Table 2, when the medical scenario includes in-hospital diagnosis and treatment scenario or post-hospital follow-up scenario, and the medical question type includes open-ended questions or multiple-choice questions, the mixed matrix includes [patient, scenario 1, question type 1], [patient, scenario 2, question type 1], [patient, scenario 1, question type 2] and [patient, scenario 2, question type 2].
[0059] Table 2
[0060] Optionally, the first prompt word is determined based on the medical role type of the first intelligent agent, the medical question meta-attribute, the first prompt word template, and the medical image text pair.
[0061] The first prompt template is a pre-determined template used to determine the components of the first prompt.
[0062] Optionally, the first prompt word is determined based on the medical role type of the first intelligent agent, the medical question meta-attribute, the first prompt word template, the medical image text pair, and the question output format.
[0063] The question output format is used to determine the output format of the medical question data generated by the first agent. For example, the question output format includes the medical question and the reason for asking the question.
[0064] Figure 2 This is one of the example images provided in this application for generating the first prompt word, such as... Figure 2 As shown, when the medical role type of the first intelligent agent is a medical student, the medical question meta-attributes include medical question type and medical scenario, the medical question type is an open question and answer, the medical scenario is an in-hospital diagnosis and treatment scenario, and the question output format includes the medical question and the reason for asking the question, the medical image-text pair includes a medical image and an image description, and the example A of the first prompt word obtained is as follows: system: I am a medical student, in my second year of doctoral studies, specializing in thoracic surgery. I enjoy discussing details and prefer detailed answers. User: Based on your persona and the images and descriptions below, construct an open-ended question about a hospital treatment scenario. The answer to the question must be derived from the images and descriptions. You only need to provide the question, not the answer, and give the reason for asking the question. Medical images: <image-a>; Image description: Focal fibrosis. a) Super-resolution CT shows the anterior segment of the right upper lobe...; Please output the following JSON format: {"Medical Issues": XXX,} Reason for asking the question: XXX}.
[0065] in, <image-a>This is a placeholder for a medical image.
[0066] Step 103: Input the first prompt word into the first intelligent agent to obtain the medical problem data output by the first intelligent agent.
[0067] Specifically, the first prompt word is input into the first intelligent agent, which then generates medical problem data for the medical image-text pair that conforms to the medical role type and medical problem meta-attributes of the first intelligent agent.
[0068] For example, if the first prompt word of Example A is input into the doctor agent, the doctor agent outputs the medical question data based on the combined hybrid matrix [medical student, scenario 1, question type 1] as shown below: Medical question: "What characteristics does the cell arrangement in the image possess, and what does it represent?" Reason for the question: "The image is a pathological image with obvious pathological features. Since the question is asked by a second-year doctoral student in medicine, and the question is to construct an open-ended question-and-answer question in a hospital diagnosis and treatment scenario, the question given is 'What characteristics does the cell arrangement in the image have, and what does it represent?'"
[0069] Figure 3 This is the second example image provided in this application for generating the first prompt word, such as... Figure 3 As shown, when the medical role type of the first intelligent agent is patient, the medical question meta-attributes include medical question type and medical scenario, the medical question type is multiple choice, the medical scenario is post-discharge follow-up, and the question output format includes medical question and reason for questioning, the medical image-text pair includes medical image and image description, and the example B of the first prompt word obtained is as follows: system: I am a 52-year-old woman with underlying COPD and no medical knowledge; User: Based on your persona and the images and descriptions below, construct a multiple-choice question depicting a post-discharge follow-up scenario. The answer must be derived from the images and descriptions. Provide the correct answer from the image description, and for the remaining three incorrect options, please supplement with choices that are somewhat related to the image. You only need to provide the question, not the answer, and give your reason for posing the question. Medical images: <image-c>; Image description: Atypical adenomatous hyperplasia. a) Super-resolution CT shows the posterior segment of the right upper lobe...; Please output the following JSON format: {"Medical Issues": XXX,} Reason for asking the question: XXX}.
[0070] in, <image-c>This is a placeholder for a medical image.
[0071] Correspondingly, the first prompt word of Example B is input into the patient agent. The patient agent outputs the medical question data based on the given combination of the hybrid matrix [medical role type = patient, medical scenario = post-discharge follow-up, medical question type = single choice] as shown below: Medical question: "What disease do I have based on the picture? Choose: (a) No abnormalities; (b) Atypical adenomatous hyperplasia; (c) Pulmonary nodules; (d) Atelectasis" Reason for the question: "The image is a lung CT scan with obvious pathological features. Since the question is about a patient and a single-choice question is to be constructed in a hospital follow-up scenario, the question is given as follows: 'What disease can I tell from the image? Choose: (a) No abnormalities; (b) Atypical adenomatous hyperplasia; (c) Pulmonary nodules; (d) Atelectasis'."
[0072] Step 104: Determine the second prompt word based on the medical problem data and the medical image text pair.
[0073] The second prompt word is the prompt word used by the second agent to generate medical image text pairs and corresponding medical image question-and-answer pairs.
[0074] The second agent is used to generate medical image question-and-answer pairs based on the medical question data output by the first agent. The medical role type of the second agent is different from that of the first agent.
[0075] For example, the medical role type of the second intelligent agent is medical expert, and the second intelligent agent is a medical expert intelligent agent.
[0076] Optionally, the first intelligent agent is a medical student intelligent agent, a patient intelligent agent, or a first aider intelligent agent; and / or, the second intelligent agent is a medical expert intelligent agent.
[0077] The first and second agents are built based on a large model.
[0078] For example, the first and second agents are built on the qwen3-vl-235B large model, which has powerful basic capabilities and is suitable as a base model for multimodal medical data synthesis.
[0079] Specifically, after constructing medical question data output by different role intelligent agents, when determining the second prompt word for generating medical image question-answer pairs, prompt word generation processing is performed based on the medical question data and medical image text pairs to obtain the second prompt word.
[0080] Optionally, the second prompt word is determined based on the medical problem data, the medical image text pair, and the second prompt word template.
[0081] The second prompt template is a pre-determined template used to determine the components of the second prompt.
[0082] Optionally, the second prompt word is determined based on the medical problem data, the medical image text pair, the second prompt word template, and the question-and-answer pair output format.
[0083] The question-and-answer pair output format is used to determine the output format of the medical image question-and-answer pairs generated by the second agent. For example, the question-and-answer pair output format includes the question and the answer.
[0084] Figure 4 This is an example image of generating second prompt words provided in this application, such as... Figure 4 As shown, in the case where the first agent is a medical student agent and the second agent is a medical expert agent, based on the medical question data generated from the first prompt word in the aforementioned example A, the second prompt word, example C, is determined as follows: System: You are a medical expert. Please answer questions based on user needs. User: I am a medical student, second-year PhD student, specializing in thoracic surgery. I enjoy discussing details and prefer thorough answers. Please answer my question based on the image and description below, and provide the radiological evidence supporting the image. If the answer cannot be derived from the image and description, please type "Refuse to answer". If the question is irrelevant to the image or unreasonable, please type "Unreasonable question". Medical images: <image-b>; Image description: Focal fibrosis. b) Microscopic examination shows moderate to severe localized fibrosis within the lesion; Medical data: What characteristics do the cell arrangements in the image possess, and what do they represent? Please output the following JSON format: Problem: XXX "Answer": XXX}
[0085] in, <image-b>This is a placeholder for a medical image.
[0086] Step 105: Input the second prompt word into the second agent to obtain the medical image question-and-answer pair output by the second agent.
[0087] The medical image question-and-answer pair includes a medical image question and a corresponding medical image question and answer. The medical image question includes the medical image in the medical image text pair and the medical image question description. The medical image answer corresponds to the image description in the medical image text pair.
[0088] Specifically, the second prompt word is input into the second intelligent agent, which generates a medical image question-and-answer pair for the medical image text pair based on the second prompt word, thus obtaining a medical image question-and-answer pair synthesized from multimodal medical data.
[0089] For example, if the second prompt word and the medical image text pair from Example C are input into a medical expert agent, the medical expert agent will output the following medical image question-and-answer pair: Question: "What characteristics does the cell arrangement in the image possess, and what does it represent?" Answer: "This pathological image shows focal fibrosis. Microscopic examination reveals moderate to severe localized fibrosis within the lesion, with visible..."
[0090] Optionally, after inputting the second prompt word into the second agent and obtaining the medical image question-and-answer pair output by the second agent, the medical image question-and-answer pair is input into the second agent, prompting the second agent to perform multi-dimensional scoring on the input medical image question-and-answer pair to obtain the credibility score of the medical image question-and-answer pair; if the credibility score is greater than or equal to the preset credibility threshold, the medical image question-and-answer pair is retained; if the credibility score is less than the preset credibility threshold, the medical image question-and-answer pair is deleted.
[0091] By using a second intelligent agent to repeatedly review the generated medical image question-and-answer pairs, it can be ensured that the synthesized multimodal data has the correctness of medical knowledge and the relevance of images and text; unreasonable questions and questions that cannot be answered will be rejected.
[0092] The multimodal medical data synthesis method provided in this application designs a multi-agent data synthesis collaborative architecture. First, a first agent expands the semantic range of image understanding of medical image-text pairs by constructing a multi-dimensional hybrid matrix according to different medical role types and different medical question meta-attributes to generate medical question data. Then, a second agent generates text from the medical image text and medical question data to output medical image question-and-answer pairs. This method can automatically, scalably, and systematically generate medical image question-and-answer pairs with multiple roles, scenarios, and dimensions of professional questions based on the same medical image-text pairs under different medical roles and different medical question meta-attributes. It realizes the generation of medical image question-and-answer pairs with diverse question-and-answer types and questioning angles, which can meet the needs of high-quality training data for medical image question-and-answer model training, significantly reduce the data threshold and cost of model training, and make the model trained with medical image question-and-answer pairs more accurate and reliable in understanding complex clinical scenarios and answering professional questions.
[0093] Furthermore, by using medical role types to determine prompt words when generating medical question data, it can be ensured that the generated medical question data is closely integrated with medical practice and clinical scenarios, effectively supporting the deep visual-linguistic semantic association of the medical image question answering model. In addition, by simultaneously inputting medical image text pairs into the agent to generate medical question data and medical image question answering pairs, it is ensured that the question answering pair generation process fully integrates visual and text information.
[0094] Based on the above embodiments, as an optional embodiment, determining the first prompt word according to the medical role type of the first intelligent agent, the medical question meta-attribute, and the medical image text pair includes: Based on the aforementioned medical role type, determine role preference data; Based on the aforementioned medical question meta-attributes, determine the user's question data; Based on the first prompt word template, the role preference data, the user question data, and the medical image text pair are combined to obtain the first prompt word.
[0095] Specifically, when determining the first prompt word, on the one hand, role preference data is determined based on the medical role type of the first intelligent agent, and on the other hand, user question data is determined based on the medical question meta-attributes. Then, according to the first prompt word template, the role preference data, user question data, and medical image text pairs are combined to obtain the first prompt word used to generate medical question data.
[0096] For example, in Example A of the first prompt word, the medical role type of the first agent is a medical student, and the determined role preference data is "I am a medical student, a second-year doctoral student specializing in thoracic surgery. I like to discuss details and prefer detailed answers." The medical question meta-attribute is the medical question type of open-ended questions and the medical scenario of in-hospital treatment scenarios. The determined user question data is "Based on your persona, construct an open-ended question for an in-hospital treatment scenario based on the following image and description. The answer to the question needs to be derived from the image and image description. You only need to give the question, not the answer, and give the reason for asking the question."
[0097] Optionally, the first prompt word is obtained by combining role preference data, user question data, medical image text pairs, and question output format according to the first prompt word template.
[0098] In one embodiment, the medical problem meta-attributes also include image type, anatomical location, lesion type, and imaging features to generate medical problem data involving professional perspectives such as image type recognition, anatomical location, lesion identification, and imaging feature analysis.
[0099] The multimodal medical data synthesis method provided in this application determines role preference data by utilizing the medical role type of the first intelligent agent, determines user question data by utilizing medical question meta-attributes, and combines the role preference data, user question data, and medical image text pairs according to the first prompt word template to obtain the first prompt word. It can simulate the interaction of multiple roles and multiple medical scenarios such as medical students, patients, and emergency responders, making the generated medical question data and medical image question-and-answer pairs closer to the real clinical workflow. It can effectively improve the practicality and ease of use of the synthesis of the medical image question-and-answer model trained on this multimodal medical data using medical image question-and-answer pairs.
[0100] Based on the above embodiments, as an optional embodiment, the second intelligent agent is a medical expert intelligent agent; the step of determining the second prompt word based on the medical problem data and the medical image text pair includes: Based on the aforementioned medical expert AI agent, determine the respondent's preference data; Based on the aforementioned medical role type, determine the questioner's preference data; Based on the second prompt word template, the respondent preference data, the questioner preference data, the medical question data, and the medical image text pair are combined to obtain the second prompt word.
[0101] Specifically, when determining the second prompt word, on the one hand, the respondent preference data is determined based on the medical expert role of the medical expert agent; on the other hand, the questioner preference data is determined based on the medical role type of the first agent; then, according to the second prompt word template, the respondent preference data, the questioner preference data, the medical question data, and the medical image text pair are combined to obtain the second prompt word used to generate medical image question-and-answer pairs.
[0102] For example, in Example C of the second prompt, the respondent's preference data is determined based on the medical expert role of the medical expert agent as "You are a medical expert, please answer the question according to the user's needs"; the questioner's preference data is determined based on the medical role type as "I am a medical student, a second-year doctoral student, my specialty is thoracic surgery, I like to discuss details and prefer detailed answers. Please answer my question based on the following picture and description, and provide the imaging evidence based on the picture"; the medical question data generated by the first agent is "What characteristics does the cell arrangement in the picture have, and what does it represent?"
[0103] Optionally, based on the second prompt word template, the respondent preference data, questioner preference data, medical question data, medical image-text pairs, and question-answer pair output formats are combined to obtain the second prompt word.
[0104] The multimodal medical data synthesis method provided in this application determines the questioner's preference data by utilizing the medical role type of the first intelligent agent, determines the respondent's preference data by utilizing the medical expert intelligent agent, and combines the respondent's preference data, questioner's preference data, medical question data, and medical image text pairs according to the second prompt word template. This method can improve the scene realism and question-and-answer professionalism of the medical image question-and-answer pairs generated by the medical expert intelligent agent based on the second prompt word.
[0105] Furthermore, by adopting a multi-agent collaborative architecture consisting of role-based intelligent agents and medical expert intelligent agents, and constructing a hybrid matrix through multi-dimensional combinations of role personas, medical scenarios, and medical question types, it is possible to simultaneously integrate image understanding intelligent agents and text generation intelligent agents, ensuring that the question-and-answer pair generation process fully integrates visual and textual information, making the generation of medical image question-and-answer pairs more flexible and information-rich.
[0106] Based on the above embodiments, as an optional embodiment, the acquisition of medical image text pairs includes: Obtain structured data files of medical imaging literature; Locate the visualization image in the structured data file and obtain the contextual description of the visualization image; the visualization image includes at least one sub-image. The Large Language Model (LLM) is used to perform image-text association processing on the visualized image, the contextual description of the visualized image, and the chapter titles of the medical imaging literature to obtain image-text pairs; Filter the medical image-text pairs from the image-text pairs; The medical image-text pairs include single medical image-text pairs and composite medical image-text pairs; the composite medical image-text pairs include composite medical images and associated text information; the composite medical images include multiple medical sub-images; the associated text information includes the context description and the chapter title.
[0107] Visualized images include medical images and non-medical images; medical images include single medical images involving only one medical subplot and composite medical images involving multiple medical subplots.
[0108] Specifically, Figure 5 This is a schematic diagram of the process for obtaining image-text pairs provided in this application, such as... Figure 5 As shown, structured data files are extracted from medical imaging documents in different document data formats (such as PDF, CAJ, HTML, etc.). Image range search is used to locate the visualized image in the structured data file. Based on the position of the visualized image in the structured data file, several contextual descriptions of the visualized image and the chapter titles of the medical imaging documents are obtained.
[0109] Further leveraging a large language model for deep semantic understanding of structured data, the model inputs visual images, their contextual descriptions, and chapter titles from medical imaging literature for image-text association processing. The large language model intelligently identifies all visual elements such as figures and tables within the text and accurately associates each visual element with its corresponding contextual description and chapter title, extracting high-quality image-text pairs. Each image-text pair includes a document image and its corresponding associated text information, including the contextual description and chapter title. This precise image-text pair extraction based on the large language model overcomes the limitations of traditional rule-based matching, significantly improving extraction accuracy and recall.
[0110] The image-text pairs initially extracted from medical imaging literature include not only medical image-text pairs, but also non-medical image-text pairs such as advertisements, icons, and logos. Therefore, single medical image-text pairs and composite medical image-text pairs are selected as medical image-text pairs, while non-medical image-text pairs are eliminated.
[0111] Optionally, combined Figure 5 As shown, when medical imaging documents are in PDF format, high-precision optical character recognition technology is used to identify the text information and layout of the medical imaging documents, and the recognition results are obtained; the recognition results are then reconstructed into structured data files in Markdown and JSON formats.
[0112] In another embodiment, combined Figure 5 As shown, when medical imaging documents are in HTML format, the medical imaging documents themselves are structured data files. The step of obtaining the structured data file of the medical imaging documents can be skipped. Using an HTML parser, the document object model of the HTML format medical imaging document webpage can be directly analyzed to quickly locate and extract embedded visual images, contextual descriptions of visual images, and chapter titles of medical imaging documents.
[0113] By adopting a dual-channel parallel processing architecture, the document is intelligently split into two channels according to its format. A differentiated strategy is used for in-depth content extraction. For scanned medical image documents in PDF format, OCR intelligent recognition and structured reconstruction technology are used for parsing, realizing the conversion from unstructured image-based PDFs to structured data. For medical image documents in HTML format, direct extraction based on DOM parsing is adopted, which utilizes its inherent structured features for parsing, fast location and extraction, balancing processing efficiency and extraction depth.
[0114] The multimodal medical data synthesis method provided in this application, through structured processing of multi-source, heterogeneous medical image documents and image-text association processing based on a large model, enables medical image text pairs, regardless of their source format, to be uniformly encapsulated into standardized composite image and title pairs within the system, ensuring consistency of downstream processing interfaces. By innovatively applying a large language model to fine-grained information extraction from medical image documents, semantic understanding is expanded, which can effectively improve the accuracy of image-text association.
[0115] Optionally, medical image-text pairs are filtered from image-text pairs, including: The document image in the image-text pair is input into a pre-trained medical image classification model to obtain the image category result output by the medical image classification model; Image-text pairs corresponding to document images whose image category results are medical images are identified as medical image-text pairs; The medical image classification model is built on a deep residual network (ResNet); it is pre-trained using multiple training samples; each training sample includes an image sample and its corresponding image category label.
[0116] Image category results include, but are not limited to, categories such as X-ray images, CT scans, MRI imaging, ultrasound images, pathological sections, and non-medical images.
[0117] By applying a fine-grained image element classification method based on deep residual networks in the critical quality filtering stage of the seed data preparation pipeline for medical image-text pairs, the original image-text pairs obtained from medical literature can be accurately classified, effectively eliminating non-medical related images and ensuring the purity and professionalism of the dataset for subsequent processing.
[0118] Optionally, the medical image classification model uses the weighted focus loss function shown in the following expression: ; in, This represents the total loss value. This represents the number of training samples; Image category The class weight coefficient is determined by the number of samples in each image class; the smaller the number of samples in each class, the larger the class weight coefficient. For training samples Belongs to the image category The predicted probability; These are the modulation parameters.
[0119] By setting the modulation factor This can reduce the number of easily classified samples (i.e. The loss contribution of training samples close to 1 enables the medical image classification model to focus on training samples that are difficult to classify; by assigning greater weight to medical image categories with fewer samples, the class imbalance problem is effectively solved.
[0120] Optionally, the medical image classification model is trained using the following method: In terms of training data construction, a phased automatic and manual collaborative annotation process is adopted. First, the model is pre-trained using a large-scale image classification dataset, and the model performs initial screening and pre-annotation of image samples to generate candidate image category labels. Then, for difficult image samples with low prediction confidence and randomly sampled image samples, experts in fields such as radiologists perform fine manual annotation to form a high-quality golden test set. Simultaneously, targeted data augmentation techniques such as rotation, color jitter, and elastic deformation are used to optimize the class imbalance problem, ultimately constructing a high-quality training set containing 15,000 training samples. In terms of model training optimization, a transfer learning strategy is adopted: first, a ResNet model pre-trained on a very large-scale dataset is used to obtain general feature representations, and then the medical image classification dataset collected earlier is used for end-to-end fine-tuning. The results show that after only 8 training epochs, the model achieves an accuracy of 95% on the validation set, fully demonstrating the effectiveness of the training data construction, loss function design, and transfer learning strategy.
[0121] By combining task-specific automatic pre-annotation with expert-precise annotation in training data construction, domain-specific datasets can be efficiently built. Designing a focus loss function with class weight adjustment can effectively improve the model's ability to distinguish difficult samples. Establishing an efficient transfer learning strategy can achieve rapid model optimization on small-scale professional datasets, providing high-quality medical image data support for subsequent image segmentation, alignment, and question-answer pair construction.
[0122] Based on the above embodiments, as an optional embodiment, the acquisition of medical image text pairs includes: Obtain composite medical image-text pairs; the composite medical image-text pairs include composite medical images and associated text data; the composite medical images include multiple medical sub-images; the associated text data includes a contextual description of the composite medical images and chapter titles of medical imaging literature; The composite medical image in the composite medical image-text pair is segmented into multiple medical sub-images. Semantic segmentation is performed on the associated text data in the composite medical image-text pair to obtain multiple subtitle text descriptions and multiple image body text descriptions; Multiple medical sub-images, multiple sub-title text descriptions, and multiple image body text descriptions are aligned to obtain multiple aligned single image-text pairs; each aligned single image-text pair includes one medical sub-image and its corresponding aligned text data. Based on the aligned single image-text pair, the medical image-text pair is determined.
[0123] The subtitle text description is a highly generalized description of the image content of the medical sub-image, such as "focal fibrosis" or "atypical adenomatous hyperplasia", providing preliminary semantic guidance for the image content.
[0124] The image text is supplementary explanatory text embedded in the main text of the medical imaging literature. It contains more detailed and specific clinical observation and descriptive information such as the specific location, morphological characteristics, and relationship with surrounding tissues of the lesion in the image, which can further enrich the semantic information of the image.
[0125] Many medical images in medical imaging literature are composite images comprising multiple medical sub-images. Therefore, medical image-text pairs often include composite medical image-text pairs. While composite medical image-text pairs can be directly used to construct prompts for generating medical image question-and-answer pairs, the inclusion of multiple medical sub-images and their corresponding text descriptions interferes with the agent's generation of question-and-answer pairs for individual medical sub-images. Therefore, to improve the effectiveness of role-based agents and medical expert agents in generating medical image question-and-answer pairs during multimodal medical data synthesis, it is necessary to split composite medical image texts into multiple aligned single-image-text pairs, allowing role-based agents and medical expert agents to generate medical image question-and-answer pairs based on these individual aligned single-image-text pairs.
[0126] Specifically, Figure 6 This is a schematic diagram of the process for obtaining medical image text pairs provided in this application, such as... Figure 6 As shown, several composite medical image-text pairs are selected from the image-text pairs. Each composite medical image-text pair includes a composite medical image and a corresponding text description (i.e., associated text data).
[0127] On the one hand, composite medical images containing multiple medical sub-images in composite medical image-text pairs are segmented to obtain multiple medical sub-images.
[0128] On the other hand, semantic segmentation is performed on the associated text data in the composite medical image text pairs to obtain multiple subtitle text descriptions and multiple image text descriptions. The subtitle text descriptions correspond one-to-one with the chapter titles of medical imaging documents, while the image text descriptions are some or all of the text descriptions extracted from the context descriptions of the composite medical images.
[0129] Align the multiple medical sub-images, multiple sub-title text descriptions, and multiple image text descriptions corresponding to the composite medical image text pair to obtain multiple aligned single image text pairs. Each aligned single image text pair includes a medical sub-image and its corresponding aligned text data. Each aligned text data includes a sub-title text description and / or the image text description.
[0130] It should be noted that, for the associated text data corresponding to composite medical images, provided that the associated text data includes a contextual description with a preset amount of text information, the associated text data includes all or part of the text description of the medical imaging literature.
[0131] The multimodal medical data synthesis method provided in this application, when preparing the seed data of medical image text pairs required for intelligent agents to synthesize medical image question-and-answer pairs, decomposes the composite medical image text pairs into multiple more refined aligned single image-text pairs by performing sub-image segmentation on the composite medical image, semantic segmentation on the associated text data, and image-text alignment on the medical sub-image, sub-title text description, and image body text description. This allows the role intelligent agent and the medical expert intelligent agent to synthesize medical image question-and-answer pairs based on the aligned single image-text pairs, thereby improving the effect of multimodal medical data synthesis and generating richer medical image question-and-answer pairs.
[0132] Based on the above embodiments, as an optional embodiment, the step of performing sub-image segmentation on the composite medical image in the composite medical image-text pair to obtain multiple medical sub-images includes: The composite medical image in the composite medical image-text pair is input into the sub-graph segmentation model to obtain multiple medical sub-graphs output by the sub-graph segmentation model; The subgraph segmentation model is based on an improved YOLO model. The feature extraction network of the improved YOLO model introduces a self-attention mechanism (Self-Attn). The prediction head of the improved YOLO model incorporates a multi-scale feature fusion module. The loss function of the improved YOLO model is determined based on the bounding box coordinate loss term, the object presence confidence loss term, the classification loss term, and the self-attention regularization term.
[0133] Specifically, Figure 7 This is a flowchart illustrating the subgraph segmentation process provided in this application, such as... Figure 7 As shown, after selecting composite medical image-text pairs from structured data files of medical imaging literature, a sub-image segmentation model is used to segment the composite medical images within these pairs. This requires introducing a self-attention mechanism into the feature extraction network of the original YOLO model and a multi-scale feature fusion module into the prediction head of the original YOLO model to obtain an improved YOLO model. The sub-image segmentation model is then implemented based on this improved YOLO model. By introducing the self-attention mechanism, the sub-image segmentation model can better capture the correlation between different regions in the image, especially demonstrating stronger recognition capabilities for complex structures and small objects commonly found in medical images. Furthermore, the introduction of a multi-scale feature fusion module in the prediction head effectively improves the detection accuracy of medical sub-images of different sizes.
[0134] Furthermore, based on the improved YOLO model, a composite loss function for the subgraph segmentation model is designed, including a bounding box coordinate loss term, an object presence confidence loss term, a classification loss term, and a self-attention regularization term. By introducing a new self-attention regularization term into the composite loss function, the model's attention to important regions can be enhanced.
[0135] After the subgraph segmentation model is pre-trained, when performing subgraph segmentation on a composite medical image, the composite medical image is input into the subgraph segmentation model, which outputs multiple medical subgraphs.
[0136] For a composite medical image containing six sub-images, the sub-image segmentation model can output six medical sub-images after processing; for a composite medical image containing three sub-images, the sub-image segmentation model can output three medical sub-images after processing.
[0137] Optionally, the expression for the composite loss function of the subgraph segmentation model is as follows: ; in, This is the composite loss, or the total loss value. Loss for bounding box coordinates; These are the weighting coefficients for the bounding box coordinate loss; The object exhibits confidence loss. The weighting coefficients for the confidence loss of the object; For classification loss; These are the weighting coefficients for the classification loss; For self-attention regularization; represents the weight coefficients of the self-attention regularization term.
[0138] By balancing and adjusting the weight coefficients, the subgraph segmentation model can improve localization accuracy while maintaining high recall.
[0139] Optionally, the subgraph segmentation model is trained using a two-stage training strategy. Specifically, it is first pre-trained using large-scale open-source image segmentation data to enable the subgraph segmentation model to acquire general object detection capabilities; then, it is fine-tuned using medical-specific data to focus on optimizing the subgraph segmentation model's ability to perceive medical image features. This training strategy ensures both the generalization of the subgraph segmentation model and its excellent performance in a specific domain.
[0140] The multimodal medical data synthesis method provided in this application significantly improves the accuracy of subgraph segmentation by introducing a self-attention mechanism into a subgraph segmentation model based on the YOLO model, particularly for complex medical images. Through an improved loss function, the model maintains high detection recall while improving bounding box localization accuracy by approximately 15%. Effective composite medical image segmentation increases the amount of usable image data by about three times, significantly improving the efficiency of subsequent processing. This provides reliable technical support for the refined processing of medical image data and lays a solid foundation for subsequent image-text alignment and question-answering pair construction.
[0141] Based on the above embodiments, as an optional embodiment, the semantic segmentation of the associated text data in the composite medical image-text pair to obtain multiple subtitle text descriptions and multiple image body text descriptions includes: Based on a preset regular expression, the associated text data is initially segmented to obtain the initially segmented text; The initial segmented text and its contextual verification content are input into a lightweight language model to obtain the segmentation boundary confidence score output by the lightweight language model. Based on the initial segmented text whose segmentation boundary confidence score is greater than a preset segmentation confidence score threshold, a semantically reasonable segmented text is determined. Based on the semantically reasonable text segmentation and the third prompt word template, the third prompt word is determined and input into the large model arbitration layer to obtain multiple subtitle text descriptions and multiple image text descriptions output by the large model arbitration layer.
[0142] In addition to sub-segmenting composite medical images, it is also necessary to perform title segmentation and text description extraction on the associated text data, including composite subheadings and inline captions in the main text of medical imaging documents.
[0143] To address this, a hybrid semantic unit segmentation method for medical image literature is employed. This method constructs a three-level dynamic arbitration architecture consisting of a rule engine, a lightweight language model, and a large language model. It adopts a hierarchical decision-making mechanism and achieves semantic unit segmentation through a three-level processing pipeline, thereby solving the problem of balancing accuracy and efficiency in the segmentation of compound titles and the extraction of image description text in medical literature.
[0144] Specifically, at the rule engine layer, a multi-pattern matching rule base is constructed based on the structured features of medical literature. This rule base includes pre-defined regular expressions such as those for image title patterns, case description patterns, and section marker patterns. Using these pre-defined regular expressions, semantic segmentation is performed on associated text data, including composite subheadings and image descriptions within the main text, resulting in multiple preliminary segmented texts and achieving rule-based multi-level text segmentation.
[0145] In the small model validation layer, a context window with a preset data volume (e.g., 512 tokens) is first used to determine the context validation content corresponding to each preliminary segmented text output by the rule engine layer. Then, the multiple preliminary segmented texts output by the rule engine layer and their corresponding context validation content are input into the lightweight language model, which performs semantic rationality verification and obtains the segmentation boundary confidence score output by the lightweight language model. The preliminary segmented texts with segmentation boundary confidence scores greater than the preset segmentation confidence score threshold are determined as semantically reasonable segmented texts, thereby optimizing medical entity recognition and boundary ambiguity resolution.
[0146] In the large model arbitration layer, which serves as the final decision-making module, the thinking chain prompting engineering technique is used to process the semantically reasonable segmented text according to the third prompt word template, determine the third prompt word, and input the third prompt word into the large model arbitration layer. This guides the large model arbitration layer to complete tasks such as subtitle boundary recognition, image description text extraction, and preservation of the integrity of medical terminology. The segmentation results are output in a structured JSON format, resulting in multiple subtitle text descriptions and multiple image text descriptions output by the large model arbitration layer.
[0147] Optionally, the lightweight language model adopts a 7B-parameter medical professional language model; the large model arbitration layer is built based on the large language model, which adopts the DeepSeek-32B language model.
[0148] Optionally, a deep-level prompting fine-tuning strategy and parameter-efficient prompting fine-tuning technology are adopted to integrate medical professional knowledge such as anatomical terminology, pathological descriptions, and case reports into the large language model of the large model arbitration layer, thereby injecting medical prior knowledge. Among them, pathological descriptions include, but are not limited to, lesion characteristics, relationships with surrounding tissues, measurement data, etc.; case reports include, but are not limited to, clinical history, imaging manifestations, pathological diagnosis, etc.; imaging manifestations include, but are not limited to, location, size, shape, boundary, signal characteristics, etc.
[0149] In one embodiment, the third prompt word template preset using mind chain prompting engineering technology is shown below: prompt_template="" You are a medical literature analysis expert. Please perform semantic unit segmentation on the following text: Text: {text} Requirements: 1. Identify all subheading boundaries; 2. Extract image-related descriptive text (inline caption); 3. Maintain the integrity of medical terminology; 4. Provide a confidence score for ambiguous boundaries. Please output in JSON format: {{ "segments":[ {{ "text":"segmented text", "type":"subtitle / caption / content", "confidence": 0.95 }} ] }}.
[0150] The multimodal medical data synthesis method provided in this application performs preliminary segmentation of related text data based on preset regular expressions at the rule engine layer, performs semantic rationality verification of the preliminary segmented text using a lightweight language model at the small model verification layer, and makes a final decision on the semantically reasonable segmented text based on a large language model at the large model arbitration layer, thereby obtaining segmented subheading text descriptions and multiple image text descriptions. This solves the problem of balancing accuracy and efficiency in composite title segmentation and image description text extraction in medical literature.
[0151] Optionally, the large model arbitration layer also outputs the confidence scores corresponding to multiple subtitle text descriptions and multiple image text descriptions. The confidence scores are determined based on an accuracy-adaptive weight allocation mechanism, which includes the accuracy scores of the three module layers—the real-time calculation rule engine layer, the small model validation layer, and the large model arbitration layer—on the validation set. By combining considerations of efficiency weights and accuracy weights, dynamic weight allocation for normalization processing is achieved.
[0152] Optionally, considering the long text characteristics of medical image documents, a hierarchical attention mechanism based on sliding window attention is designed. By using sliding window attention technology and setting parameters for window size and overlapping area, combined with a context caching mechanism, the boundary effect problem in long text processing can be effectively avoided.
[0153] Optionally, since multiple tasks such as title segmentation and long text processing are involved, a three-level dynamic arbitration architecture consisting of a rule engine layer, a small model verification layer, and a large model arbitration layer is constructed to build a unified training objective function for the multi-task situation. The text segmentation loss, image description extraction loss, and medical entity recognition loss are organically combined through weight coefficients to achieve collaborative optimization of multiple tasks.
[0154] The calculation formula for the training objective function of the three-level dynamic arbitration architecture (rule engine layer, small model validation layer, and large model arbitration layer) is as follows: ; in, The objective function value; For text segmentation loss; The weights for the text segmentation loss; Extract loss for image description; Extract the weights of the loss for image description; Identify losses for medical entities; Weights for identifying losses for medical entities.
[0155] Optionally, the pre-training process of the rule engine layer, the small model verification layer, and the large model arbitration layer adopts a course learning strategy, which gradually increases the processing difficulty in three stages. The first stage is to train the basic segmentation task of the first text of the first length, the second stage is to train the multi-task learning of the second text of the second length, and the third stage is to train the complex structure processing problem of the third text of the third length. The first length is less than the second length, and the second length is less than the third length.
[0156] For example, the first text of the first length is a short text of 500 characters or less, the second text of the second length is a medium-length text of 500-1500 characters, and the third text of the third length is a long text of more than 1500 characters.
[0157] Table 3 presents the multi-dimensional evaluation results of the hybrid semantic unit segmentation method provided in this application. As shown in Table 3, during the comparative experiment, a multi-dimensional evaluation standard system was established, and a systematic evaluation was conducted on a test set containing 10,000 medical documents. The experimental results show that the hybrid semantic unit segmentation method achieves a subheading segmentation accuracy of 94.2%, an InlineCaption extraction F1 score of 92.7%, a long text processing accuracy of 89.5%, and an average inference speed of 3.2 seconds per document. All indicators are significantly better than the single rule baseline, the single small model, and the single DeepSeek-32B model. Error analysis based on the confusion matrix shows that adding medical dictionary matching improves the processing effect of cases with ambiguous boundaries by 5.3%, adopting a recursive segmentation strategy improves the processing of nested structures by 8.7%, and the entity recognition module reduces errors related to professional terminology recognition by 42%.
[0158] Table 3
[0159] As can be seen, this hybrid semantic unit segmentation method has four significant advantages: First, in terms of balancing accuracy and efficiency, it maintains a segmentation accuracy of 94.2% while improving inference speed by 2.7 times compared to a single large-scale model; second, in terms of adaptability to the medical field, the specially optimized medical text processing capabilities enable a professional terminology recognition F1 score of 93.5%; third, in terms of system robustness, the dynamic arbitration mechanism ensures stable performance under various document structures; and finally, the scalable architecture design supports the rapid integration of new processing rules and model components. This method provides a complete solution for the intelligent processing of medical documents, establishing a high-quality structured data foundation for subsequent image-text alignment and knowledge extraction.
[0160] In another embodiment, the associated text data in the composite medical image text pair is input into a pre-trained semantic segmentation model to obtain multiple subtitle text descriptions and multiple image body text descriptions output by the semantic segmentation model.
[0161] Based on the above embodiments, as an optional embodiment, the step of aligning the multiple medical sub-images, multiple sub-title text descriptions, and multiple image body text descriptions to obtain multiple aligned single image-text pairs includes: The plurality of medical sub-images, the plurality of subtitle text descriptions, and the plurality of image text descriptions are input into the image-text alignment model to obtain the initial similarity of the plurality of single image-text pairs output by the image-text alignment model; the single image-text pair includes any of the medical sub-images, any of the subtitle text descriptions, and any of the image text descriptions. For each single image-text pair, a penalty factor is determined based on the first sequential index of the medical sub-image in the single image-text pair and the second sequential index of the sub-title text description in the single image-text pair. The initial similarity of the single image-text pair is adjusted based on the penalty factor to obtain the corrected similarity of the single image-text pair, so as to determine the corrected similarity of all single image-text pairs. For each medical sub-image, a set of single-image-text pairs containing the medical sub-image is determined. Based on the single-image-text pairs with the highest similarity in the set of single-image-text pairs, the aligned single-image-text pairs of the medical sub-image are determined, so as to determine the aligned single-image-text pairs of all medical sub-images.
[0162] The first sequential index is the sequential index of the medical sub-image among all medical images in the medical imaging literature; the second sequential index is the sequential index of the chapter title corresponding to the subtitle text description among all chapter titles in the medical imaging literature.
[0163] The image-text alignment model is used to perform preliminary alignment of medical sub-images and subtitle text descriptions, as well as image body text descriptions.
[0164] Optionally, the image-text alignment model is implemented based on the CLIP model.
[0165] Specifically, multiple medical sub-images, multiple subtitle text descriptions, and multiple image text descriptions are input into the image-text alignment model. The image-text alignment model determines the initial similarity of a single image-text pair formed by any medical sub-image, any subtitle text description, and any image text description, thereby determining the initial similarity of all combined single image-text pairs.
[0166] Furthermore, leveraging the prior knowledge that images and titles typically appear sequentially in medical literature, a sequence-related penalty factor is introduced. This involves determining the penalty factor for each single image-text pair based on the first sequence index of the medical sub-image and the second sequence index of the sub-title text description within that pair. The initial similarity of the single image-text pairs output by the image-text alignment model is then adjusted based on this penalty factor to obtain the corrected similarity. This process is repeated for all single image-text pairs output by the image-text alignment model, calculating both the penalty factor and the corrected similarity to determine the corrected similarity for all pairs.
[0167] For each medical subgraph appearing in medical literature, identify several single-image-text pairs that include that subgraph, construct a set of single-image-text pairs for that subgraph, and then determine the single-image-text pair with the highest similarity in the set as the aligned single-image-text pair for that subgraph. Repeat the steps of determining the set of single-image-text pairs and aligning single-image-text pairs for all medical subgraphs appearing in the medical literature to determine the aligned single-image-text pairs for all medical subgraphs.
[0168] In one embodiment, the formula for calculating the corrected similarity of any single image-text pair is as follows: ; ; in, To correct the similarity; Image features Text features The initial similarity; Image features The first-order index and text features of the corresponding medical subgraph The ordinal distance between the second ordinal indices of the corresponding subheading text descriptions; For the first Image features identified by a medical sub-image; For the first The text features identified by the subheading text description; Image features The first sequential index of the corresponding medical subgraph; Text features The second sequential index of the corresponding subheading text description; The modulation function is a function that varies with the order distance. A weighting function that increases and decreases; It is a cosine function.
[0169] Alternatively, the expression for the modulation function is as follows: ; in, It is a decay factor between 0 and 1.
[0170] For example, =0.8.
[0171] When the first sequential index of the medical subplot and the second sequential index of the subtitle text description are the same or adjacent, i.e., sequential distance... , At this time, the initial similarity output by the image-text alignment model is basically unaffected. When the order distance When it is very large, This will significantly reduce the similarity score of single image-text pairs that are not matched in order.
[0172] By using the sequential index of medical sub-images and subtitle text descriptions to determine the penalty factor, the initial similarity of the preliminary image-text alignment is adjusted. The adjusted corrected similarity is then used to determine the aligned single image-text pairs, effectively avoiding the error of matching the image of Figure A with the title of Figure B with high similarity, and guiding the model to prioritize finding matches within the local neighborhood.
[0173] The multimodal medical data synthesis method provided in this application adjusts the initial similarity of single image-text pairs initially aligned by the image-text alignment model based on the sequential index of medical subgraphs and subtitle text descriptions in medical literature. This reduces the corrected similarity of single image-text pairs that do not conform to the sequence rules and increases the corrected similarity of single image-text pairs that conform to the sequence rules, so that the finally determined aligned single image-text pairs of medical subgraphs match the order of medical subgraphs and subtitle text descriptions in medical literature.
[0174] Optionally, in order to achieve accurate alignment of medical sub-images and subtitle text descriptions and image body descriptions, a two-stage alignment architecture is adopted. First, an improved CLIP model is used for rapid pre-alignment to select high-probability matching pairs. Then, a multimodal large model is introduced for refined semantic-level verification and final alignment, thereby constructing high-quality seed image-text pair data.
[0175] In the first stage of the image-text alignment model pre-alignment, the alignment of single image-text pairs for all medical sub-images is obtained.
[0176] In the second stage of fine alignment of the multimodal large model, the system inputs the aligned single image-text pairs generated by CLIP pre-alignment, along with their contextual information, into the multimodal large model for deep semantic verification. Through a secondary verification mechanism based on prompt word engineering, prompt templates are used to guide the multimodal model to perform a specialized image-text relevance discrimination task. For example, the prompt word could be, "Please determine whether this medical image is highly relevant to the text description below, and provide a confidence level."
[0177] Multimodal large model fine alignment no longer relies solely on shallow feature similarity, but is able to deeply understand the complex semantic relationships between medical entities and lesion descriptions in text and visual features in images, thereby correcting and confirming the preliminary results of CLIP, and finally outputting highly confident aligned single-text and image pairs.
[0178] The two-stage progressive processing effectively balances computational efficiency and alignment accuracy, significantly improving the automation level and result quality of image and text data processing in the medical field.
[0179] Furthermore, when constructing an image-text alignment model based on the CLIP model, matching medical sub-images with all subtitle text descriptions leads to a larger search range, reduced search efficiency, and decreased accuracy in image-text matching.
[0180] Therefore, a constraint mechanism based on image sequence and text structure is introduced into the image-text alignment model. On the one hand, an image order constraint is introduced. When calculating the initial similarity, the image-text alignment model will prioritize the order in which images appear in the document and their proximity to the corresponding subheadings to avoid mismatches between images and text in unrelated chapters. On the other hand, a strategy of calculating similarity in batches is adopted. Instead of comparing a single medical sub-image with all subheading text descriptions in the entire document globally, it performs local initial similarity calculations with several candidate subheading text descriptions that are most likely to be related (such as subheadings belonging to the same compound title). This greatly reduces the computational complexity and improves the efficiency of the image-text alignment model in processing long documents. At the same time, by focusing on the local context, the accuracy of the initial matching is enhanced.
[0181] When calculating similarity in batches, individual medical sub-images will no longer be compared with all images. Instead of aligning and comparing individual subheading text descriptions with images, the entire medical imaging literature is compiled based on the institution's textual medical imaging documents. The subheading text descriptions are divided into multiple batches. For a given medical subgraph First, determine the medical sub-map. Batch to which it belongs For example, with the medical subgraph All subheadings under the same compound heading constitute a batch.
[0182] Assuming batch Inside Each subheading text description, then the medical sub-figure The final match target is limited to this batch, as shown in the following expression: ; in, Medical subgraph In batch The title index of the subtitle text description that is finally matched.
[0183] By calculating similarity in batches, the computational complexity and noise issues of global matching of long texts in medical imaging literature are resolved. Furthermore, batch similarity calculation reduces the computational complexity from O(N) to O(M). It also reduces interference from irrelevant options in local contexts, such as matching different descriptions a and b for the same symptom "atypical adenomatous hyperplasia" instead of matching all texts, thus enhancing the accuracy of the initial matching.
[0184] The multimodal medical data synthesis method provided in this application utilizes cross-modal semantic understanding technology to accurately align images and text, enabling automated screening and construction of high-quality image-text related data. By providing accurate and rich image-text related data, it plays a crucial supporting role in downstream tasks such as medical knowledge graph construction, auxiliary diagnostic model training, and medical education content generation, thereby improving the efficiency and accuracy of medical data processing.
[0185] Overall, the multimodal medical data synthesis method provided in this application, centered on the core objective of "constructing high-quality medical multimodal data," has built a complete technical system that integrates large and small model collaboration and multi-agent architecture to systematically solve the challenges of extracting knowledge from original medical literature and generating high-quality training data.
[0186] In the seed data preparation stage, which involves collaboration between large and small models, high-quality seed image-text pairs are automatically constructed from multi-source heterogeneous medical image literature. This process includes: an automated corpus acquisition engine based on hybrid perception to collect literature from well-known domestic and international medical journals; filtering non-medical images using a fine-grained image classification model based on deep residual networks; precise segmentation of composite images using an object detection model that integrates self-attention mechanisms; subheading segmentation and text information extraction using a hybrid semantic unit segmentation method (combining a rule engine and dynamic arbitration between small and large models); and finally, a two-stage cross-modal alignment strategy. First, an improved CLIP model is used for rapid pre-alignment, and then a multimodal large model is introduced for refined semantic verification. This process ultimately produces high-confidence aligned single and composite image-text pairs, laying a solid foundation for subsequent data synthesis.
[0187] In the multi-agent data synthesis stage based on a hybrid matrix, the high-quality seed data produced in the first stage is used to generate diverse medical question-answer pairs on a large scale through an innovative multi-agent framework. The core of this framework includes: constructing a hybrid matrix composed of roles (such as medical students and patients), scenarios (such as in-hospital diagnosis and treatment and post-discharge follow-up), and question types (such as open-ended questions and multiple-choice questions), which greatly expands the semantic space and coverage of the generated questions; using carefully designed prompt word templates, the large model is driven to play the roles of "medical student agent" and "patient agent" respectively to generate questions that conform to specific matrix combinations; and then the "medical expert agent" combines seed images and text information to generate professional and accurate answers, thus forming the final instruction-answer pair; the entire process can also include a data review scheme, which uses a second set of multi-modal large models to perform multi-dimensional scoring and filtering to ensure the quality and reliability of the synthesized data.
[0188] Through two major stages of seamless integration and meticulous design, the system achieves automated synthesis of high-quality, multi-dimensional, and Chinese-language medical practice-relevant multimodal training data from original literature, providing crucial data support for training high-performance medical multimodal large-scale models.
[0189] Figure 8 This is a schematic diagram of the multimodal medical data synthesis device provided in this application, as shown below. Figure 8 As shown, the multimodal medical data synthesis device includes, but is not limited to, a text-image pair acquisition module 210, a first prompt word determination module 220, a medical question acquisition module 230, a second prompt word determination module 240, and a question-and-answer pair acquisition module 250.
[0190] The image-text pair acquisition module 210 is used to acquire medical image-text pairs.
[0191] The first prompt word determination module 220 is used to determine the first prompt word based on the medical role type of the first intelligent agent, the medical problem meta-attribute, and the medical image text pair.
[0192] The medical problem acquisition module 230 is used to input the first prompt word into the first intelligent agent and obtain the medical problem data output by the first intelligent agent.
[0193] The second prompt word determination module 240 is used to determine a second prompt word based on the medical problem data and the medical image text pair.
[0194] The question-and-answer pair acquisition module 250 is used to input the second prompt word into the second intelligent agent to obtain the medical image question-and-answer pair output by the second intelligent agent.
[0195] It should be noted that the multimodal medical data synthesis device provided in this application can execute the multimodal medical data synthesis method described in any of the above embodiments during specific operation, which will not be elaborated in this embodiment.
[0196] The multimodal medical data synthesis device provided in this application, through the design of a multi-agent data synthesis collaborative architecture, first utilizes a first agent to expand the semantic range of image understanding of medical image-text pairs by constructing a multi-dimensional hybrid matrix according to different medical role types and different medical question meta-attributes, thereby generating medical question data. Then, a second agent is used to generate text from the medical image text and medical question data to output medical image question-and-answer pairs. It can automatically, scalably, and systematically generate medical image question-and-answer pairs with multiple roles, multiple scenarios, and multiple dimensions of professional questions based on the same medical image-text pairs under different medical roles and different medical question meta-attributes. It realizes the generation of medical image question-and-answer pairs with diverse question-and-answer types and questioning angles, which can meet the needs of high-quality training data for medical image question-and-answer model training, significantly reduce the data threshold and cost of model training, and make the model trained using medical image question-and-answer pairs more accurate and reliable in understanding complex clinical scenarios and answering professional questions.
[0197] Furthermore, by using medical role types to determine prompt words when generating medical question data, it can be ensured that the generated medical question data is closely integrated with medical practice and clinical scenarios, effectively supporting the deep visual-linguistic semantic association of the medical image question answering model. In addition, by simultaneously inputting medical image text pairs into the agent to generate medical question data and medical image question answering pairs, it is ensured that the question answering pair generation process fully integrates visual and text information.
[0198] As an optional embodiment, the prompt word determination module is further configured to: determine role preference data based on the medical role type; determine user question data based on the medical question meta-attribute; and combine the role preference data, the user question data, and the medical image text pair according to the first prompt word template to obtain the first prompt word.
[0199] As an optional embodiment, the second intelligent agent is a medical expert intelligent agent; the second prompt word determination module is further configured to: determine respondent preference data based on the medical expert intelligent agent; determine questioner preference data based on the medical role type; and combine the respondent preference data, the questioner preference data, the medical question data, and the medical image text pair according to the second prompt word template to obtain the second prompt word.
[0200] As an optional embodiment, the image-text pair acquisition module further includes a composite image-text pair acquisition module, a sub-image segmentation module, a semantic segmentation module, an image-text alignment module, and an image-text pair determination module; The composite image-text pair acquisition module is used to acquire composite medical image-text pairs; the composite medical image-text pair includes composite medical images and associated text data; the composite medical image includes multiple medical sub-images; the associated text data includes a context description of the composite medical image and a chapter title of the medical imaging literature; The sub-image segmentation module is used to perform sub-image segmentation on the composite medical image in the composite medical image-text pair to obtain multiple medical sub-images; The semantic segmentation module is used to perform semantic segmentation on the associated text data in the composite medical image text pair to obtain multiple subtitle text descriptions and multiple image body text descriptions; The image-text alignment module is used to align multiple medical sub-images, multiple sub-title text descriptions, and multiple image body text descriptions to obtain multiple aligned single image-text pairs; each aligned single image-text pair includes one medical sub-image and its corresponding aligned text data. The image-text pair determination module is used to determine the medical image text pair based on the aligned single image-text pair.
[0201] As an optional embodiment, the image-text alignment module is further configured to input the plurality of medical sub-images, the plurality of sub-title text descriptions, and the plurality of image body text descriptions into the image-text alignment model to obtain the initial similarity of the plurality of single image-text pairs output by the image-text alignment model; the single image-text pair includes any one of the medical sub-images, any one of the sub-title text descriptions, and any one of the image body text descriptions; for each single image-text pair, a penalty factor is determined according to the first order index of the medical sub-image in the single image-text pair and the second order index of the sub-title text description in the single image-text pair, and the initial similarity of the single image-text pair is adjusted based on the penalty factor to obtain the corrected similarity of the single image-text pair, so as to determine the corrected similarity of all single image-text pairs; for each medical sub-image, a set of single image-text pairs containing the medical sub-image is determined, and the aligned single image-text pairs of the medical sub-image are determined based on the single image-text pair with the highest corrected similarity in the set of single image-text pairs, so as to determine the aligned single image-text pairs of all medical sub-images.
[0202] As an optional embodiment, the semantic segmentation module is further configured to perform preliminary segmentation on the associated text data based on a preset regular expression to obtain preliminary segmented text; input the preliminary segmented text and its contextual validation content into a lightweight language model to obtain the segmentation boundary confidence score output by the lightweight language model, and determine semantically reasonable segmented text based on the preliminary segmented text whose segmentation boundary confidence score is greater than a preset segmentation confidence score threshold; determine a third prompt word based on the semantically reasonable segmented text and a third prompt word template, and input the third prompt word into a large model arbitration layer to obtain multiple subtitle text descriptions and multiple image text descriptions output by the large model arbitration layer.
[0203] As an optional embodiment, the sub-graph segmentation module is further configured to input the composite medical image in the composite medical image-text pair into the sub-graph segmentation model to obtain multiple medical sub-graphs output by the sub-graph segmentation model; wherein, the sub-graph segmentation model is implemented based on an improved YOLO model; the feature extraction network of the improved YOLO model introduces a self-attention mechanism; the prediction head of the improved YOLO model introduces a multi-scale feature fusion module; the loss function of the improved YOLO model is determined based on the bounding box coordinate loss term, the object presence confidence loss term, the classification loss term, and the self-attention regularization term.
[0204] As an optional embodiment, the image-text pair acquisition module is further configured to acquire a structured data file of medical imaging literature; locate a visualized image in the structured data file and acquire the contextual description of the visualized image; the visualized image includes at least one sub-image; perform image-text association processing on the visualized image, the contextual description of the visualized image, and the chapter title of the medical imaging literature using a large language model to obtain image-text pairs; filter out the medical image-text pairs from the image-text pairs; wherein, the medical image-text pairs include single medical image-text pairs and composite medical image-text pairs; the composite medical image-text pairs include composite medical images and associated text information; the composite medical images include multiple medical sub-images; the associated text information includes the contextual description and the chapter title.
[0205] Figure 9 The structural schematic diagram of the electronic device provided in this application is as follows: Figure 9 As shown, the electronic device may include a processor 310, a communications interface 320, a memory 330, and a communication bus 340, wherein the processor 310, the communications interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 can call logical instructions in the memory 330 to execute the multimodal medical data synthesis method provided in any of the above embodiments. The multimodal medical data synthesis method includes, but is not limited to, the following steps: acquiring medical image-text pairs; determining a first prompt word based on the medical role type of the first intelligent agent, the medical question meta-attribute, and the medical image-text pairs; inputting the first prompt word into the first intelligent agent to obtain medical question data output by the first intelligent agent; determining a second prompt word based on the medical question data and the medical image-text pairs; and inputting the second prompt word into a second intelligent agent to obtain medical image-question-answer pairs output by the second intelligent agent.
[0206] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0207] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multimodal medical data synthesis method provided in any of the above embodiments. The multimodal medical data synthesis method includes, but is not limited to, the following steps: acquiring medical image-text pairs; determining a first prompt word based on the medical role type of a first intelligent agent, the medical question meta-attribute, and the medical image-text pairs; inputting the first prompt word into the first intelligent agent to obtain medical question data output by the first intelligent agent; determining a second prompt word based on the medical question data and the medical image-text pairs; and inputting the second prompt word into a second intelligent agent to obtain medical image question-and-answer pairs output by the second intelligent agent.
[0208] In another aspect, this application also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the multimodal medical data synthesis method provided in any of the above embodiments. The multimodal medical data synthesis method includes, but is not limited to, the following steps: acquiring medical image-text pairs; determining a first prompt word based on the medical role type of a first intelligent agent, the medical question meta-attribute, and the medical image-text pairs; inputting the first prompt word into the first intelligent agent to obtain medical question data output by the first intelligent agent; determining a second prompt word based on the medical question data and the medical image-text pairs; and inputting the second prompt word into a second intelligent agent to obtain medical image question-and-answer pairs output by the second intelligent agent.
[0209] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0210] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0211] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for synthesizing multimodal medical data, characterized in that, include: Retrieve medical image text pairs; The first prompt word is determined based on the medical role type of the first intelligent agent, the meta-attribute of the medical question, and the medical image-text pair; The first prompt word is input into the first intelligent agent to obtain the medical problem data output by the first intelligent agent; Based on the medical problem data and the medical image text pair, a second prompt word is determined; The second prompt word is input into the second agent to obtain the medical image question-and-answer pair output by the second agent; The acquisition of medical image text pairs includes: Obtain composite medical image-text pairs; the composite medical image-text pairs include composite medical images and associated text data; the composite medical images include multiple medical sub-images; the associated text data includes a contextual description of the composite medical images and chapter titles of medical imaging literature; The composite medical image in the composite medical image-text pair is input into the sub-graph segmentation model to obtain multiple medical sub-graphs output by the sub-graph segmentation model; wherein, the sub-graph segmentation model is implemented based on the improved YOLO model; the feature extraction network of the improved YOLO model introduces a self-attention mechanism; the prediction head of the improved YOLO model introduces a multi-scale feature fusion module; A three-level dynamic arbitration architecture consisting of a rule engine, a lightweight language model, and a large language model is constructed to perform semantic segmentation on the associated text data in the composite medical image text pair, resulting in multiple subheading text descriptions and multiple image body text descriptions. A two-stage alignment architecture is adopted to perform image-text alignment on multiple medical sub-images, multiple sub-title text descriptions, and multiple image body text descriptions to obtain multiple aligned single image-text pairs. Each aligned single image-text pair includes a medical sub-image and its corresponding aligned text data. The two-stage alignment architecture includes pre-alignment based on an improved CLIP model, semantic-level verification and final alignment based on a multimodal large model, and adjusting the initial similarity of the preliminary image-text alignment by using the order index of the medical sub-images and sub-title text descriptions, and using the adjusted corrected similarity to determine the aligned single image-text pairs. Based on the aligned single image-text pair, the medical image-text pair is determined.
2. The multimodal medical data synthesis method according to claim 1, characterized in that, The step of determining the first prompt word based on the medical role type of the first intelligent agent, the medical question meta-attribute, and the medical image text pair includes: Based on the aforementioned medical role type, determine role preference data; Based on the aforementioned medical question meta-attributes, determine the user's question data; Based on the first prompt word template, the role preference data, the user question data, and the medical image text pair are combined to obtain the first prompt word.
3. The multimodal medical data synthesis method according to claim 2, characterized in that, The second intelligent agent is a medical expert intelligent agent; the step of determining the second prompt word based on the medical problem data and the medical image text pair includes: Based on the aforementioned medical expert AI agent, determine the respondent's preference data; Based on the aforementioned medical role type, determine the questioner's preference data; Based on the second prompt word template, the respondent preference data, the questioner preference data, the medical question data, and the medical image text pair are combined to obtain the second prompt word.
4. The multimodal medical data synthesis method according to any one of claims 1-3, characterized in that, The meta-attributes of the medical issue include the type of medical issue and the medical scenario.
5. The multimodal medical data synthesis method according to claim 1, characterized in that, The step involves aligning multiple medical sub-images, multiple sub-title text descriptions, and multiple image body text descriptions to obtain multiple aligned single image-text pairs, including: The plurality of medical sub-images, the plurality of subtitle text descriptions, and the plurality of image text descriptions are input into the image-text alignment model to obtain the initial similarity of the plurality of single image-text pairs output by the image-text alignment model; the single image-text pair includes any of the medical sub-images, any of the subtitle text descriptions, and any of the image text descriptions. For each single image-text pair, a penalty factor is determined based on the first sequential index of the medical sub-image in the single image-text pair and the second sequential index of the sub-title text description in the single image-text pair. The initial similarity of the single image-text pair is adjusted based on the penalty factor to obtain the corrected similarity of the single image-text pair, so as to determine the corrected similarity of all single image-text pairs. For each medical sub-image, a set of single-image-text pairs containing the medical sub-image is determined. Based on the single-image-text pairs with the highest similarity in the set of single-image-text pairs, the aligned single-image-text pairs of the medical sub-image are determined, so as to determine the aligned single-image-text pairs of all medical sub-images.
6. The multimodal medical data synthesis method according to claim 1, characterized in that, The semantic segmentation of the associated text data in the composite medical image-text pair yields multiple subtitle text descriptions and multiple image body text descriptions, including: Based on a preset regular expression, the associated text data is initially segmented to obtain the initially segmented text; The initial segmented text and its contextual verification content are input into a lightweight language model to obtain the segmentation boundary confidence score output by the lightweight language model. Based on the initial segmented text whose segmentation boundary confidence score is greater than a preset segmentation confidence score threshold, a semantically reasonable segmented text is determined. Based on the semantically reasonable text segmentation and the third prompt word template, the third prompt word is determined and input into the large model arbitration layer to obtain multiple subtitle text descriptions and multiple image text descriptions output by the large model arbitration layer.
7. The multimodal medical data synthesis method according to claim 1, characterized in that, The loss function of the improved YOLO model is determined based on the bounding box coordinate loss term, the object presence confidence loss term, the classification loss term, and the self-attention regularization term.
8. A multi-modality medical data synthesis apparatus characterized by comprising: include: The image-text pair acquisition module is used to acquire medical image-text pairs; The first prompt word determination module is used to determine the first prompt word based on the medical role type of the first intelligent agent, the medical problem meta-attribute, and the medical image text pair; The medical problem acquisition module is used to input the first prompt word into the first intelligent agent and obtain the medical problem data output by the first intelligent agent; The second prompt word determination module is used to determine the second prompt word based on the medical problem data and the medical image text pair; The question-and-answer pair acquisition module is used to input the second prompt word into the second agent and obtain the medical image question-and-answer pair output by the second agent; The acquisition of medical image text pairs includes: Obtain composite medical image-text pairs; the composite medical image-text pairs include composite medical images and associated text data; the composite medical images include multiple medical sub-images; the associated text data includes a contextual description of the composite medical images and chapter titles of medical imaging literature; The composite medical image in the composite medical image-text pair is input into the sub-graph segmentation model to obtain multiple medical sub-graphs output by the sub-graph segmentation model; wherein, the sub-graph segmentation model is implemented based on the improved YOLO model; the feature extraction network of the improved YOLO model introduces a self-attention mechanism; the prediction head of the improved YOLO model introduces a multi-scale feature fusion module; A three-level dynamic arbitration architecture consisting of a rule engine, a lightweight language model, and a large language model is constructed to perform semantic segmentation on the associated text data in the composite medical image text pair, resulting in multiple subheading text descriptions and multiple image body text descriptions. A two-stage alignment architecture is adopted to perform image-text alignment on multiple medical sub-images, multiple sub-title text descriptions, and multiple image body text descriptions to obtain multiple aligned single image-text pairs. Each aligned single image-text pair includes a medical sub-image and its corresponding aligned text data. The two-stage alignment architecture includes pre-alignment based on an improved CLIP model, semantic-level verification and final alignment based on a multimodal large model, and adjusting the initial similarity of the preliminary image-text alignment by using the order index of the medical sub-images and sub-title text descriptions, and using the adjusted corrected similarity to determine the aligned single image-text pairs. Based on the aligned single image-text pair, the medical image-text pair is determined.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the multimodal medical data synthesis method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by the processor, it implements the multimodal medical data synthesis method as described in any one of claims 1 to 7.
11. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the multimodal medical data synthesis method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Medical text data synthesis method based on role driving
CN119646164A
Question and answer data construction method and device
CN121708413A