A Multimodal Medical Information Processing Method, Device, and Medium
By introducing adaptive adjustment encoder, multi-scale self-prompt generation module and cross-contrast learning method in the medical information processing model, the problem of multimodal medical data fusion and adaptation to different diagnosis and treatment styles is solved, and the model is powerful adaptability and efficient data fusion effect is achieved.
Patent Information
- Application Number
- CN202510237750.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The high-dimensional, heterogeneous and complex semantic correlation of multimodal medical data leads to the challenges of effective fusion and in-depth analysis of data. Digital human doctors need to learn and adapt to the diagnosis and treatment styles and decision-making processes of different doctors, requiring the model to have strong learning and generalization capabilities.
Adaptive adjustment encoder (AT-encoder), multi-scale self-prompt generation module and feature pyramid network are used, and combined with cross-contrast learning methods, medical information processing models are trained to extract and fuse visual features and text embedding in multimodal data.
The model adaptability and promotion ability to diversified data is realized, the effect of multimodal data fusion and the ability to understand complex tasks is improved, and the application effect in fields such as multimodal medical care is significantly improved.
Smart Images

Figure CN119722685B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medical technology, and particularly relates to a multi-modal medical information processing method, device and medium. Background Art
[0002] With the rapid development of artificial intelligence and big data technologies, intelligent medical care has gradually become an important direction in the medical industry. As one of the core applications of intelligent medical care, digital human doctors can simulate the diagnosis and treatment processes of human doctors and provide convenient and intelligent health management and diagnosis and treatment suggestions for patients. However, multi-modal data has high dimensionality, heterogeneity, and complex semantic associations, which pose challenges to the effective fusion and in-depth analysis of the data. Each doctor has different diagnosis and treatment styles, decision-making processes, and key focuses. Digital human doctors need to have the ability to learn and adapt to these unique styles, which places high requirements on the learning and generalization capabilities of the model.
[0003] In summary, there is an urgent need for a multi-modal medical information processing method, device and medium to solve the problems in the prior art. Summary of the Invention
[0004] The object of the present invention is to provide a multi-modal medical information processing method, device and medium, and the specific technical solutions are as follows:
[0005] A multi-modal medical information processing method includes the following steps:
[0006] S1: Construct a medical information processing model, where the medical information processing model includes an adaptive adjustment encoder, a multi-scale self-prompt generation module, and a feature pyramid network;
[0007] S2: Train the medical information processing model using cross-contrast learning;
[0008] S3: Obtain the medical information of the patient, input the medical information into the trained medical information processing model, and obtain the processing result;
[0009] The adaptive adjustment encoder is used to extract the visual features of medical images in medical information. The adaptive adjustment encoder adopts a VIT architecture and performs domain bypass processing on the image embedding after the multi-head self-attention operation.
[0010] Optionally, the process of performing domain bypass processing on the image embedding after the multi-head self-attention operation is as follows:
[0011] Preprocess the image embedding using a domain-general embedding to obtain a first embedding containing general knowledge;
[0012] Process the first embedding using a domain-specific embedding to obtain a second embedding containing domain-specific knowledge of a certain domain;
[0013] Call the multi-layer perceptron to process the second embedding to obtain the final image embedding.
[0014] Optionally, the multi-scale self-prompt generation module is used to generate self-prompt tokens with multi-scale features, and the process is as follows:
[0015] Take the image embeddings of different layers in the adaptive adjustment encoder as inputs, fuse the image embeddings of different layers to obtain multi-scale image embeddings;
[0016] Use convolutional layer heads with different strides to perform foreground segmentation on the multi-scale image embeddings to obtain foreground features;
[0017] Filter out the low-confidence tokens in the foreground features to obtain self-prompt tokens.
[0018] Optionally, in the multi-scale self-prompt generation module, binary cross-entropy loss is used to optimize foreground segmentation.
[0019] Optionally, specifically fusing the image embeddings of different layers is: call the Feature Pyramid Network to integrate the image embeddings of different layers from top to bottom.
[0020] Optionally, the medical information processing model is trained using cross-contrast learning, and the process is as follows:
[0021] Extract medical information using block tokens to obtain image foreground tokens, image background tokens, text foreground features, and text background features;
[0022] Use the text foreground features and text background features as positive and negative sample pairs of the image foreground tokens to obtain the first granularity information and the second granularity information;
[0023] Use the text background features and text foreground features as positive and negative sample pairs of the image background tokens to obtain the third granularity information and the fourth granularity information;
[0024] Based on the first granularity information, the second granularity information, the third granularity information, and the fourth granularity information, perform cross-modal contrast learning to complete model training.
[0025] Optionally, contrast learning loss is used as the loss for training the medical information processing model.
[0026] In addition, the present invention also includes a computer device, including a memory and a processor;
[0027] The memory is used to store a computer program that can run on the processor;
[0028] The processor is used to implement the steps of the multi-modal medical information processing method as described above when executing the computer program.
[0029] In addition, the present invention further includes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the multi-modal medical information processing method as described above are implemented.
[0030] Applying the technical solution of the present invention has the following beneficial effects:
[0031] (1) The present invention proposes an Adaptive Tuning Encoder (AT-Encoder) for effectively coordinating the fusion of visual features with domain-general knowledge and domain-specific knowledge. While capturing image features, the AT-Encoder combines information from different knowledge domains, enhancing the model's adaptability to diverse data. In addition, in the present invention, the query decoder uses domain query features for fusion decoding, and by dynamically adjusting the feature selection during the decoding process, flexible adaptation to different knowledge domains is achieved. This joint design enables the model to exhibit excellent generalization ability in diverse knowledge domains, contributing to the wide application of the model in complex scenarios such as multi-modal medical treatment and industrial inspection.
[0032] (2) The multi-scale self-prompt generation module in the present invention is used to automatically generate high-quality self-prompts containing multi-scale knowledge to improve the effect of multi-modal data fusion. Through a multi-scale processing mechanism, the multi-scale self-prompt generation module enables the model to extract and utilize information at different scales and provides effective support for the multi-modal fusion task in the framework of self-supervised learning. By generating self-prompts with multi-level details and context, the multi-scale self-prompt generation module not only enhances the model's understanding ability for complex tasks but also improves its performance in the case of scarce multi-modal data, contributing to the promotion of wide applications in fields such as multi-modal medical treatment.
[0033] (3) The present invention proposes a cross-contrast learning method, which applies contrast learning at different granularity levels of images to achieve efficient alignment of visual features and text embeddings. The present invention conducts fine-grained contrast between the foreground and background of the image, and by training the model only using the target category and background labels, significantly improves the focusing ability on relevant regions while suppressing the interference of irrelevant objects. The contrast learning strategy of the present invention can effectively activate key regions, enabling the model to have higher accuracy and robustness when processing complex visual-text fusion tasks.
[0034] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The following will refer to the drawings for a further detailed description of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] To more clearly illustrate the technical solutions of the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0036] Figure 1 It is a flowchart of the steps of the multi-modal medical information processing method in the preferred embodiment of the present invention. Detailed implementation manners
[0037] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the following will further elaborate on the present invention in conjunction with the accompanying drawings and specific implementation manners. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0038] As Figure 1 shown, this embodiment provides a multi-modal medical information processing method, including the following steps:
[0039] A multi-modal medical information processing method, including the following steps:
[0040] S1: Construct a medical information processing model, where the medical information processing model includes an adaptive adjustment encoder, a multi-scale self-prompt generation module, and a feature pyramid network;
[0041] S2: Train the medical information processing model using cross-contrast learning;
[0042] S3: Obtain the medical information of the patient, input the medical information into the trained medical information processing model, and obtain the processing result.
[0043] It should be noted that the medical information includes medical images, case text information, gene data, etc. By applying the multi-modal medical information processing method proposed in this embodiment, the multi-modal medical information of the patient can be fused, and finally a processing result can be generated. The processing result can be a personalized diagnosis and treatment suggestion, which helps the patient recover.
[0044] Optionally, the adaptive adjustment encoder is used to extract the visual features of the medical images in the medical information. The adaptive adjustment encoder adopts a VIT architecture (composed of multiple Transformer layers), and domain bypass processing is adopted for the image embeddings after the multi-head self-attention operation.
[0045] Optionally, the image embedding after the domain bypass processing of the multi-head self-attention operation is as follows:
[0046] Use the domain-general embedding Preprocess the image embedding to obtain the first embedding containing general knowledge, with the expression as follows:
[0047] ;
[0048] where, represents the first embedding.
[0049] Use the domain-specific embedding (the domain-specific embedding of the th medical image) to process the first embedding, and obtain the second embedding containing a certain domain-specific knowledge, with the expression as follows:
[0050] ;
[0051] where, represents the second embedding, and on the basis of the general knowledge, the second embedding incorporates the th domain-specific knowledge.
[0052] Call a multi-layer perceptron to process the second embedding and obtain the final image embedding:
[0053] ;
[0054] where, represents the final image embedding.
[0055] Optionally, the multi-scale self-prompt generation module is used to generate self-prompt tokens with multi-scale features, and the process is as follows:
[0056] Take the image embeddings of different layers in the adaptive adjustment encoder as inputs, fuse the image embeddings of different layers, and obtain the multi-scale image embedding;
[0057] Use convolutional layer heads with different strides to perform foreground segmentation on the multi-scale image embedding to obtain foreground features;
[0058] Filter out the low-confidence tokens in the foreground features to obtain the self-prompt tokens, with the expression as follows:
[0059] ;
[0060] where, represents the self-prompt tokens, represents the activation function sigmoid, Indicates the threshold value used to determine high-quality foreground markers.
[0061] Optionally, in the multi-scale self-prompt generation module, binary cross-entropy loss is used to optimize foreground segmentation, and the binary cross-entropy loss is expressed as follows:
[0062] ;
[0063] represents the number of tokens in the image embedding, represents the th ground-truth token.
[0064] Optionally, fusing image embeddings from different layers specifically involves: calling a Feature Pyramid Network to integrate image embeddings from different layers from top to bottom.
[0065] Optionally, the medical information processing model is trained using cross-contrast learning, and the process is as follows:
[0066] Medical information is extracted using block tokens to obtain image foreground markers, image background markers, text foreground features, and text background features;
[0067] The text foreground feature and the text background feature are used as the positive and negative sample pairs for the image foreground marker to obtain the first granularity information and the second granularity information ; the expression is as follows:
[0068] ;
[0069] where represents the image-level label of the th class, represents the number of label classes.
[0070] The text background feature and the text foreground feature are used as the positive and negative sample pairs for the image background marker to obtain the third granularity information and the fourth granularity information ; the expression is as follows:
[0071] ;
[0072] Based on the first granularity information the second granularity information the third granularity information and the fourth granularity information , perform cross-modal contrastive learning to complete model training.
[0073] Optionally, use the contrastive learning loss as the loss for training the medical information processing model. The contrastive learning loss has the following expression:
[0074] ;
[0075] where represents the number of tokens of the image embedding, represents the hyperparameter for balancing, represents the temperature factor.
[0076] In addition, this embodiment also discloses a computer device, including a memory and a processor;
[0077] The memory is used to store a computer program that can run on the processor;
[0078] The processor is used to implement the steps of the above multi-modal medical information processing method when executing the computer program.
[0079] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device.
[0080] The computer device can be a computing device such as a mobile phone, a desktop computer, a notebook, a palm computer, and a cloud server. The computer device may include, but is not limited to, a processor and a memory. For example, the computer device may further include input / output devices, network access devices, a bus, etc.
[0081] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the computer device and connects various parts of the entire computer device through various interfaces and lines.
[0082] The memory can be used to store the computer program and / or modules. The processor realizes the computer program by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store the data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disks, memory, plug-in hard disks, Smart Media Cards (SMCs), Secure Digital (SD) cards, Flash Cards, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0083] Among them, if the modules / units integrated in the computer device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0084] In addition, an embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned multi-modal medical information processing method are implemented.
[0085] The present invention provides a multi-modal medical information processing method, device and medium based on a temporal knowledge graph. The method of the present invention includes an adaptive adjustment encoder (AT-encoder), a multi-scale self-prompt generation module and a feature pyramid network. While capturing image features, the AT-encoder combines information from different knowledge domains, enhancing the model's adaptability to diverse data. The multi-scale self-prompt generation module enables the model to extract and utilize information at different scales through a multi-scale processing mechanism and provides effective support for the multi-modal fusion task in the framework of self-supervised learning. In addition, the present invention adopts cross-contrast learning, applying contrast learning at different granularity levels of the image to achieve efficient alignment of visual features and text embeddings.
[0086] It should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0087] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multimodal medical information processing method, characterized in that: The following steps are involved: S1: constructing a medical information processing model, wherein the medical information processing model includes an adaptive adjustment encoder, a multi-scale self-prompt generation module and a feature pyramid network; S2: The medical information processing model is trained by cross-contrast learning, and the process is as follows: Medical information is extracted by using block marking to obtain image foreground marking, image background marking, text foreground features and text background features; Using the text foreground features and the text background features as a positive sample pair and a negative sample pair of the image foreground marker, to obtain first granularity information and second granularity information; Using the text background features and the text foreground features as a positive sample pair and a negative sample pair of image background labels to obtain third granularity information and fourth granularity information; Based on the first granularity information Second granularity information Third granularity information and the fourth granularity information Perform cross-modal contrastive learning to complete model training; The contrastive learning loss is used as the loss for training the medical information processing model, and the expression is as follows: Where N represents the number of labels for image embedding, λ represents the hyperparameter used for balancing, and τ represents the temperature factor; S3: Obtain the patient's medical information, input the medical information into the trained medical information processing model, and obtain the processing result; The adaptive adjustment encoder is used to extract visual features of medical images in medical information. The adaptive adjustment encoder adopts the VIT architecture and uses domain bypass to process image embedding after multi-head self-attention operation.
2. The multimodal medical information processing method according to claim 1, characterized in that: The domain bypass is used to process the image embedding after the multi-head self-attention operation. The process is as follows: Preprocess the image embedding using domain-universal embedding to obtain the first embedding containing universal knowledge; Processing the first embedding with a domain-specific embedding to obtain a second embedding that contains some domain-specific knowledge; Call the multi-layer perceptron to process the second embedding and get the final image embedding.
3. The multimodal medical information processing method according to claim 2, characterized in that: The multi-scale self-hint generation module is used to generate self-hint tags with multi-scale features. The process is as follows: Taking the image embeddings of different layers in the adaptively adjusted encoder as input, fusing the image embeddings of different layers to obtain multi-scale image embeddings; The convolutional layers with different step sizes are used to perform foreground segmentation on multi-scale image embeddings to obtain foreground features. Low-confidence tags in foreground features are filtered out to obtain self-prompted tags.
4. The multimodal medical information processing method according to claim 3, characterized in that: In the multi-scale self-hint generation module, a binary cross-entropy loss is adopted to optimize the foreground segmentation.
5. The multimodal medical information processing method according to claim 4, characterized in that: The specific method of fusing image embeddings at different layers is to call the feature pyramid network to integrate image embeddings at different layers from top to bottom.
6. A computer device, characterized in that: including memory and processor; The memory is used to store a computer program executable on the processor; The processor is used to implement the steps of the multimodal medical information processing method according to any one of claims 1 to 5 when executing the computer program.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the multimodal medical information processing method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Transform and multi-scale feature fusion-based medical image segmentation method and system
CN116977348A
Method for reconstructing three-dimensional scene based on variation fraction distillation and electroencephalogram encoder
CN118262045A
Training method of medical image segmentation model, image segmentation method and related device
CN119131522A