Multi-modal blood disease diagnosis and treatment data processing method, electronic equipment and program product

Through multimodal blood disease diagnosis and treatment data processing methods, cross-modal fusion strategies are used to extract features and generate reference plans, which solves the problem of insufficient accuracy in single-modal data processing and achieves more accurate pathology prediction and identification.

CN120600329APending Publication Date: 2025-09-05THE SECOND AFFILIATED HOSPITAL ARMY MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510532057.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In existing technologies, disease diagnosis and treatment methods based on single-modal data have characteristic limitations, resulting in insufficient accuracy in pathology prediction and identification.

Method used

A multimodal blood disease diagnosis and treatment data processing method is adopted. By acquiring multimodal diagnosis and treatment data, feature extraction is performed using a cross-modal fusion strategy, and target data is screened from a preset database to generate a reference plan.

Benefits of technology

It improves the completeness of feature extraction of disease diagnosis and treatment data, enhances the predictive accuracy of subsequent diagnosis and treatment plans, and improves the problem of insufficient accuracy in pathology prediction and recognition caused by single features in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600329A_ABST
    Figure CN120600329A_ABST
Patent Text Reader

Abstract

The invention provides a multi-mode blood disease diagnosis and treatment data processing method, electronic equipment and a program product, and relates to the technical field of artificial intelligence. The method comprises the following steps: acquiring multi-modal diagnosis and treatment data of a target patient in a blood disease treatment process; performing feature extraction on the multi-modal diagnosis and treatment data through a preset cross-modal fusion strategy to obtain fusion features corresponding to the multi-modal diagnosis and treatment data; screening target data from a preset database according to the fusion features; and generating a reference scheme corresponding to the target data through a preset reference scheme generation strategy according to the fusion features and the target data. Therefore, by taking the multi-modal diagnosis and treatment data as a data basis, feature extraction and fusion are performed respectively, so that the completeness of feature extraction of the patient diagnosis and treatment data is improved, the prediction of a subsequent diagnosis and treatment scheme is more accurate, and the defects of single feature and high accuracy of a traditional disease diagnosis and treatment data processing mode are improved. And the accuracy of subsequent pathology prediction and identification is insufficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multimodal blood disease diagnosis and treatment data processing method, electronic equipment, and program product. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, intelligent life, production, manufacturing and services have become popular in all walks of life and have become the goal pursued by producers, service providers and the public.

[0003] In existing technologies, disease diagnosis and treatment data processing is typically based on a single-modal data source. Existing technologies typically perform feature extraction on this single-modal data source and provide corresponding diagnosis and treatment recommendations. For example, a patient's blood sampling image is used to extract features, extract abnormal blood cells, and generate feature vectors of these abnormal cells. This allows for pathology identification and treatment recommendations, providing medical staff with reference data for disease diagnosis and treatment.

[0004] However, this traditional single-modal data processing method usually has limitations in the characteristics presented by single-modal data, and pathology prediction and identification based on single-modal data have the problem of insufficient accuracy. Summary of the Invention

[0005] In view of this, the purpose of the embodiments of the present application is to provide a multimodal blood disease diagnosis and treatment data processing method, electronic equipment and program product, which can improve the problem of single characteristics in traditional disease diagnosis and treatment data processing methods, which leads to insufficient accuracy in subsequent pathology prediction and identification.

[0006] To achieve the above technical objectives, the technical solutions adopted in this application are as follows:

[0007] In a first aspect, an embodiment of the present application provides a method for processing multimodal blood disease diagnosis and treatment data, the method comprising:

[0008] Obtain multimodal diagnostic and treatment data of target patients during their hematological disease treatment process;

[0009] By presetting a cross-modal fusion strategy, feature extraction is performed on the multimodal diagnosis and treatment data to obtain fusion features corresponding to the multimodal diagnosis and treatment data;

[0010] Filtering target data from a preset database based on the fusion feature, where the target data represents data in the preset database that has a high similarity to the fusion feature;

[0011] According to the fusion features and the target data, a reference solution corresponding to the target data is generated through a preset reference solution generation strategy.

[0012] In conjunction with the first aspect, in some optional embodiments, the multimodal diagnosis and treatment data includes a plurality of single-modal sub-data of different modalities;

[0013] By presetting a cross-modal fusion strategy, feature extraction is performed on the multimodal diagnosis and treatment data to obtain fusion features corresponding to the multimodal diagnosis and treatment data, including:

[0014] For each of the unimodal sub-data, according to the type of the unimodal sub-data, a feature extraction network corresponding to the type of the unimodal sub-data is used to perform feature extraction on the unimodal sub-data to obtain unimodal features corresponding to each of the unimodal sub-data;

[0015] Obtaining a stage parameter representing the diagnosis and treatment stage of the target patient;

[0016] Converting the phase parameters into a numerical vector to obtain a phase vector;

[0017] Determine the modal weight corresponding to each of the unimodal features according to the unimodal features and the stage parameters:

[0018] α m =Sfotmax(W attn Concat(F1,…,F M )+W stage OneHot(s)

[0019] Where, α m Represents the modal weight corresponding to the single modal feature with mode m, W attn 、W stage is a trainable weight matrix, Concat represents the concatenation operation, and OneHot(s) represents the phase vector;

[0020] According to the modality weights, the single modality features are weightedly fused to obtain the fused features:

[0021]

[0022] In the formula, H represents the fusion feature, M represents the number of modalities, and F m Represents a unimodal feature with mode m.

[0023] In conjunction with the first aspect, in some optional implementations, for each of the unimodal sub-data, based on the type of the unimodal sub-data, feature extraction is performed on the unimodal sub-data using a feature extraction network corresponding to the type of the unimodal sub-data to obtain unimodal features corresponding to each of the unimodal sub-data, including:

[0024] For each of the unimodal sub-data, when the type of the unimodal sub-data is the first type representing text, performing feature extraction on the unimodal sub-data by using a preset bidirectional transformer feature extraction network to obtain the unimodal feature;

[0025] When the type of the unimodal sub-data is the second type representing an image, performing feature extraction on the unimodal sub-data by using a preset residual neural feature extraction network to obtain the unimodal feature;

[0026] When the type of the unimodal sub-data is the third type representing the document, feature extraction is performed on the unimodal sub-data through a preset multi-layer perception feature extraction network to obtain the unimodal feature.

[0027] In conjunction with the first aspect, in some optional implementations, screening target data from a preset database based on the fusion feature includes:

[0028] Performing vector conversion on the data in the preset database to obtain multiple reference vectors;

[0029] Determine, based on the fused feature and the multiple reference vectors, a cosine similarity between the fused feature and each of the multiple reference vectors:

[0030]

[0031] Where S k represents cosine similarity, H represents fusion features, represents the reference vector;

[0032] A plurality of target vectors, among the plurality of reference vectors, whose cosine similarity with the fusion feature is greater than or equal to a preset threshold, are determined as the target data.

[0033] In conjunction with the first aspect, in some optional implementations, generating a reference solution corresponding to the target data by using a preset reference solution generation strategy based on the fusion feature and the target data includes:

[0034] Obtaining a risk factor representing the risk of the target patient's physical condition;

[0035] According to the fusion feature, the target data and the risk factor, the reference solution corresponding to the target data is generated by a preset large language model.

[0036] In conjunction with the first aspect, in some optional implementations, the loss function of the preset large language model is as follows:

[0037] L=λ1·CE(D raw ,D gt)+λ2·(β·R)

[0038] Where L represents the loss function, λ1 and λ2 are the balance coefficients between the diagnosis and treatment plan and the estimated risk, β represents the risk factor, R represents the risk prediction probability, CE(D raw ,D gt ) refers to the reference scheme D raw and the preset true label D gt The cross entropy loss between .

[0039] In conjunction with the first aspect, in some optional implementations, the method further includes:

[0040] According to the preset causal correction strategy, the reference plan is corrected to obtain a corrected reference plan as the target plan.

[0041] In conjunction with the first aspect, in some optional implementations, the reference solution is modified according to a preset causal modification strategy to obtain a modified reference solution as a target solution, including:

[0042] Constructing a causal graph, the causal graph including a cause node representing the pathological characteristics of the target patient, an outcome node representing the diagnosis and treatment reference plan, and a confounding factor representing an influencing factor of the outcome node;

[0043] According to the causal diagram, a causal intervention operation is performed on the reference solution to obtain a reference solution after causal intervention:

[0044]

[0045] Where D final represents the target solution, P(D final |do(H)) represents the causal intervention operation, D raw represents the result node, H represents the cause node, C represents the confounding factor, P prior (C) represents the probability distribution of confounding factors in the preset population.

[0046] In a second aspect, an embodiment of the present application further provides an electronic device, comprising a processor and a memory coupled to each other, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the electronic device executes the above-mentioned method.

[0047] In a third aspect, an embodiment of the present application further provides a computer program product, comprising a computer program, which implements the above method when executed by a processor.

[0048] The invention adopting the above technical solution has the following advantages:

[0049] In the technical solution provided in this application, the multimodal diagnosis and treatment data of the target patient during the treatment of blood diseases is first obtained, and the multimodal diagnosis and treatment data is feature extracted through a preset cross-modal fusion strategy to obtain the fusion features corresponding to the multimodal diagnosis and treatment data. Then, based on the fusion features, the target data is screened from the preset database. Finally, based on the fusion features and the target data, a reference solution corresponding to the target data is generated through a preset reference solution generation strategy. In this way, by taking the multimodal diagnosis and treatment data as the data basis, feature extraction and fusion are performed separately, the completeness of the feature extraction of the patient's diagnosis and treatment data is improved, making the prediction of subsequent diagnosis and treatment plans more accurate, and improving the problem of the traditional disease diagnosis and treatment data processing method having a single feature, which leads to insufficient accuracy in subsequent pathology prediction and recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The present application may be further illustrated by the non-limiting embodiments provided in the accompanying drawings. It should be understood that the following drawings illustrate only certain embodiments of the present application and are therefore not to be construed as limiting the scope of the present application. It is understood that a person skilled in the art can derive other relevant drawings from these drawings without inventive effort.

[0051] Figure 1 This is a structural block diagram of the electronic device provided in an embodiment of the present application.

[0052] Figure 2 A flowchart of a multimodal blood disease diagnosis and treatment data processing method provided in an embodiment of the present application.

[0053] Icon: 100-electronic device; 101-processor; 102-memory. DETAILED DESCRIPTION

[0054] The present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that similar or identical parts in the drawings or descriptions are numbered the same. Implementations not shown or described in the drawings are known to those of ordinary skill in the art. In the description of this application, the terms "first," "second," etc. are used solely to distinguish descriptions and are not to be construed as indicating or implying relative importance.

[0055] Please refer to Figure 1 In an embodiment of the present application, an electronic device 100 may include a processor 101 and a memory 102. The memory 102 stores a computer program. When the computer program is executed by the processor 101, the electronic device 100 can perform the corresponding steps in the following multimodal blood disease diagnosis and treatment data processing method.

[0056] In this embodiment, the processor 101 may be an integrated circuit chip having signal processing capabilities. The processor 101 may be a general-purpose processor. For example, the processor 101 may be a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, and may implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present application.

[0057] The memory 102 may be, but is not limited to, a random access memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, etc. In this embodiment, the memory 102 may be used to store multimodal diagnosis and treatment data, preset cross-modal fusion strategies, fusion features, preset databases, target data, preset reference plan generation strategies, reference plans, preset causal correction strategies, target plans, etc. Of course, the memory 102 may also be used to store programs, and the processor 101 executes the programs after receiving execution instructions.

[0058] It is understandable that Figure 1 The structure of the electronic device 100 shown in FIG is only a schematic diagram of a structure. The electronic device 100 may also include Figure 1 More components shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.

[0059] In this embodiment, the electronic device 100 can be a personal computer, cloud server, laptop computer, etc., and is used to obtain multimodal diagnostic and treatment data of a target patient during their hematological disease treatment process. The device then extracts features from the multimodal diagnostic and treatment data using a preset cross-modal fusion strategy to obtain fusion features corresponding to the multimodal diagnostic and treatment data. Based on the fusion features, target data is then filtered from a preset database. Finally, based on the fusion features and the target data, a reference protocol corresponding to the target data is generated using a preset reference protocol generation strategy.

[0060] Please refer to Figure 2 The present application also provides a multimodal blood disease diagnosis and treatment data processing method, which can be applied to an electronic device 100, and each step of the method is executed or implemented by the electronic device 100. The multimodal blood disease diagnosis and treatment data processing method can include the following steps:

[0061] Step 210, obtaining multimodal diagnosis and treatment data of the target patient during the process of hematological disease treatment;

[0062] Step 220: extracting features from the multimodal medical data using a preset cross-modal fusion strategy to obtain fusion features corresponding to the multimodal medical data;

[0063] Step 230: Filter target data from a preset database based on the fusion feature, where the target data represents data in the preset database that has a high similarity to the fusion feature;

[0064] Step 240 : Generate a reference solution corresponding to the target data according to the fusion features and the target data through a preset reference solution generation strategy.

[0065] In the above-mentioned implementation, the multimodal diagnosis and treatment data of the target patient during the treatment of hematological diseases is first obtained, and the multimodal diagnosis and treatment data is feature extracted through a preset cross-modal fusion strategy to obtain the fusion features corresponding to the multimodal diagnosis and treatment data. Then, based on the fusion features, the target data is screened from the preset database. Finally, based on the fusion features and the target data, a reference scheme corresponding to the target data is generated through a preset reference scheme generation strategy. In this way, by using the multimodal diagnosis and treatment data as the data basis, feature extraction and fusion are performed separately, the completeness of the feature extraction of the patient's diagnosis and treatment data is improved, making the prediction of subsequent diagnosis and treatment plans more accurate, and improving the problem of the traditional disease diagnosis and treatment data processing method having a single feature, which leads to insufficient accuracy in subsequent pathology prediction and recognition.

[0066] The following is a detailed description of the steps in the multimodal blood disease diagnosis and treatment data processing method:

[0067] In step 210, the multimodal diagnosis and treatment data may include but is not limited to basic patient information (such as age, gender, medical history, etc.), medical interview record text, test results (such as hemoglobin, white blood cell count, etc.), cell microscopic images (such as blood smears), and nursing records (such as body temperature, medication time series, etc.).

[0068] In this embodiment, the acquisition of multimodal diagnosis and treatment data can be in the method testing stage, where the user inputs the patient's basic information, medical interview record text, test results, cell micrographs, nursing records and other information in advance as multimodal diagnosis and treatment data, and pre-stores it in the memory 102 of the above-mentioned electronic device 100. During the subsequent data processing and reference scheme generation process, it is called based on the operation instructions issued by the user through the processor 101 of the above-mentioned electronic device 100; or, the acquisition method of multimodal diagnosis and treatment data can also be in the method use stage, where the user inputs the patient's multimodal diagnosis and treatment data in real time, and uploads it to the above-mentioned processor 101 in real time for subsequent data processing and reference scheme generation. The scheme for acquiring multimodal diagnosis and treatment data is not specifically limited here.

[0069] In step 220, the multimodal diagnosis and treatment data may include a plurality of single-modal sub-data of different modalities;

[0070] By presetting a cross-modal fusion strategy, feature extraction is performed on the multimodal diagnosis and treatment data to obtain fusion features corresponding to the multimodal diagnosis and treatment data, which may include:

[0071] For each of the unimodal sub-data, according to the type of the unimodal sub-data, a feature extraction network corresponding to the type of the unimodal sub-data is used to perform feature extraction on the unimodal sub-data to obtain unimodal features corresponding to each of the unimodal sub-data;

[0072] Obtaining a stage parameter representing the diagnosis and treatment stage of the target patient;

[0073] Converting the phase parameters into a numerical vector to obtain a phase vector;

[0074] The stage parameters may be obtained based on the test results in the multimodal diagnosis and treatment data:

[0075] s=argmax(W stage-cls ·D lab )

[0076] Where s represents the stage parameter, W stage-cls Denotes the trainable stage classification weight matrix, D lab Indicates test results;

[0077] Determine the modal weight corresponding to each of the unimodal features according to the unimodal features and the stage parameters:

[0078] α m =Softmax(W attn Concat(F1,…,F M )+W stage OneHot(s)

[0079] Where, α m Represents the modal weight corresponding to the single modal feature with mode m, W attn 、W stage is a trainable weight matrix, Concat represents the concatenation operation, and OneHot(s) represents the phase vector;

[0080] According to the modality weights, the single modality features are weightedly fused to obtain the fused features:

[0081]

[0082] In the formula, H represents the fusion feature, M represents the number of modalities, and F m Represents a unimodal feature with mode m.

[0083] In this way, this embodiment dynamically adjusts the weights of each modality according to the diagnosis and treatment stage parameters (i.e., stage vectors, such as acute stage, recovery stage, etc.), so that the fusion features are more relevant to the patient's current state and improve applicability.

[0084] In this embodiment, for each of the unimodal sub-data, according to the type of the unimodal sub-data, feature extraction is performed on the unimodal sub-data using a feature extraction network corresponding to the type of the unimodal sub-data to obtain unimodal features corresponding to each of the unimodal sub-data, which may include:

[0085] For each of the unimodal sub-data, when the type of the unimodal sub-data is the first type representing text, feature extraction is performed on the unimodal sub-data using a preset bidirectional transformer feature extraction network to obtain the unimodal feature, which is expressed as follows:

[0086] F text =BERT(D text )

[0087] Where, F text Denotes the unimodal features corresponding to the first type of unimodal sub-data, D text represents the first type of unimodal sub-data, and BERT represents a preset bidirectional transformer feature extraction network;

[0088] When the type of the unimodal sub-data is the second type representing an image, feature extraction is performed on the unimodal sub-data using a preset residual neural feature extraction network to obtain the unimodal feature, which is expressed as follows:

[0089] F image =ResNet(D image )

[0090] Where, F imageDenotes the unimodal features corresponding to the second type of unimodal sub-data, D image Represents the second type of single-modal sub-data, and ResNet represents the preset residual neural feature extraction network;

[0091] When the type of the unimodal sub-data is the third type representing the document, feature extraction is performed on the unimodal sub-data by using a preset multi-layer perceptual feature extraction network to obtain the unimodal feature, which is expressed as follows:

[0092] F m =MLP(D m )

[0093] Where, F m Denotes the unimodal features corresponding to the second type of unimodal sub-data, D m It represents the second type of single-modal sub-data (such as basic patient information, laboratory indicators, and nursing records), and MLP represents the preset multi-layer perceptual feature extraction network.

[0094] In this way, according to the semantics of text, the spatiality of images, and the structure of documents, dedicated networks are used to improve the discriminability of features, avoid feature degradation caused by general models, and improve the accuracy of feature extraction.

[0095] In step 230, based on the fusion features, filtering target data from a preset database may include:

[0096] Performing vector conversion on the data in the preset database to obtain multiple reference vectors;

[0097] Determine, based on the fused feature and the multiple reference vectors, a cosine similarity between the fused feature and each of the multiple reference vectors:

[0098]

[0099] Where S k represents cosine similarity, H represents fusion features, represents the reference vector;

[0100] A plurality of target vectors, among the plurality of reference vectors, whose cosine similarity with the fusion feature is greater than or equal to a preset threshold, are determined as the target data.

[0101] In this way, cosine similarity calculations can quickly locate highly relevant pathologies, reducing redundant calculations. At the same time, a preset threshold balances retrieval precision and recall, adapting to the needs of different clinical scenarios.

[0102] In step 240, based on the fusion features and the target data, a reference solution corresponding to the target data is generated by using a preset reference solution generation strategy, which may include:

[0103] Obtaining a risk factor representing the risk of the target patient's physical condition:

[0104] β=σ(W β ·D base )

[0105] Where β represents the risk factor, W β represents the risk weight, D base represents the basic information of the patient, σ(·) represents the activation function;

[0106] In this embodiment, the risk weight represents the changing relationship between the patient's basic information and the diagnosis and treatment risk. When a patient has characteristics such as advanced age or frailty, the patient is determined to have low risk tolerance;

[0107] According to the fusion feature, the target data and the risk factor, the reference solution corresponding to the target data is generated by a preset large language model.

[0108] In this embodiment, the loss function of the preset large language model can be as follows:

[0109] L=λ1·CE(D raw ,D gt )+λ2·(β·R)

[0110] Where L represents the loss function, λ1 and λ2 are risk balance coefficients, β represents the risk factor, and R represents the risk prediction probability, which can include the estimated probabilities of multiple risks, such as the probability of drug allergy, the probability of side effects, the probability of complications, etc. It can also represent the weighted sum of multiple risk probabilities to scalar the probability of multiple risks. CE(D raw ,D gt ) refers to the reference scheme D raw and the preset true label D gt The cross entropy loss.

[0111] In this way, this embodiment constrains the output of the large language model through risk factors (such as the risk of patient complications), reduces the probability of generating high-risk solutions, and combines cross-entropy loss with risk prediction to ensure that the generated reference solution has a reasonable balance between accuracy and safety.

[0112] As an optional implementation, the method may further include:

[0113] According to the preset causal correction strategy, the reference plan is corrected to obtain a corrected reference plan as the target plan.

[0114] In this embodiment, the reference solution is modified according to a preset causal modification strategy to obtain a modified reference solution as a target solution, including:

[0115] Constructing a causal graph, the causal graph including a cause node representing the pathological characteristics of the target patient, an outcome node representing the diagnosis and treatment reference plan, and a confounding factor representing an influencing factor of the outcome node;

[0116] According to the causal diagram, a causal intervention operation is performed on the reference solution to obtain a reference solution after causal intervention:

[0117]

[0118] Where D final represents the target solution, P(D final |do(H)) represents the causal intervention operation, D raw represents the result node, H represents the cause node, C represents the confounding factor, P prior (C) represents the probability distribution of confounding factors in the preset population.

[0119] In this embodiment, the fused features are used as the cause nodes of the causal graph, and the reference solution is used as the result node of the causal graph;

[0120] According to the basic information of the patient, the dense features of complications are extracted by presetting a multi-layer perceptron to obtain the confounding factors:

[0121] P=MLP comorbidity (D base )

[0122] Where, represents MLP comorbidity Preset the multi-layer perceptron. In this way, a causal graph is constructed.

[0123] In this embodiment, P prior (C) can represent the probability distribution of the confounding factor (the confounding factor in this embodiment represents the complications of the disease corresponding to the reference solution) in the preset group. The preset group can be all the people recorded when the user performs statistics on blood disease related data (i.e., statistics on blood disease complications). For example, if the probability of all recorded people having diabetes is 12%, then P prior (C1=1)=0.12.

[0124] It is understandable that in the aforementioned step 230, this embodiment obtains multiple target vectors as target data by screening through a preset threshold, that is, the target data contains multiple alternative solutions, and the reference solution generated in step 240 is also a collection of multiple solutions. This embodiment uses a causal correction strategy to correct the reference solution, and performs a causal intervention operation on each target vector in the target data to obtain the probability of each reference solution corresponding to each target vector, that is, P(D final |do(H)), and then select the reference solution with the highest probability as the final target solution.

[0125] In this way, through causal intervention operations, the interference of confounding factors on the results generated by the reference scheme is removed, so that the final target scheme generated depends only on pathological characteristics, enhancing the universality, objectivity and accuracy of the scheme.

[0126] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the electronic device 100 described above can refer to the corresponding processes of each step in the aforementioned method, and will not be elaborated here.

[0127] An embodiment of the present application further provides a computer program product, including a computer program, which implements the above-mentioned multimodal blood disease diagnosis and treatment data processing method when executed by the processor 101.

[0128] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present application can be implemented through hardware or by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each implementation scenario of the present application.

[0129] In summary, the embodiments of the present application provide a multimodal hematological disease diagnosis and treatment data processing method, electronic device and program product. In this technical solution, first, the multimodal diagnosis and treatment data of the target patient during the hematological disease treatment process is obtained, and through a preset cross-modal fusion strategy, feature extraction is performed on the multimodal diagnosis and treatment data to obtain the fusion features corresponding to the multimodal diagnosis and treatment data, and then according to the fusion features, the target data is screened from the preset database, and finally, according to the fusion features and the target data, a preset reference scheme generation strategy is used to generate a reference scheme corresponding to the target data. In this way, by taking the multimodal diagnosis and treatment data as the data basis, feature extraction and fusion are performed separately, the perfection of the feature extraction of the patient's diagnosis and treatment data is improved, the prediction of the subsequent diagnosis and treatment plan is made more accurate, and the problem of the single feature in the traditional disease diagnosis and treatment data processing method is improved, which leads to insufficient accuracy in subsequent pathology prediction and recognition.

[0130] In the embodiments provided in the present application, it should be understood that the disclosed method can also be implemented in other ways. The method embodiments described above are merely schematic. For example, the flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of code, and a part of the module, program segment or code includes one or more executable instructions for implementing the specified logical function. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. In addition, the functional modules in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0131] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A multimodal blood disease diagnosis and treatment data processing method, characterized in that: The method comprises: Obtain multimodal diagnostic and treatment data of target patients during their hematological disease treatment process; By presetting a cross-modal fusion strategy, feature extraction is performed on the multimodal diagnosis and treatment data to obtain fusion features corresponding to the multimodal diagnosis and treatment data; Filtering target data from a preset database based on the fusion feature, where the target data represents data in the preset database that has a high similarity to the fusion feature; According to the fusion features and the target data, a reference solution corresponding to the target data is generated through a preset reference solution generation strategy.

2. The method according to claim 1, characterized in that The multimodal diagnosis and treatment data includes a plurality of single-modal sub-data of different modalities; By presetting a cross-modal fusion strategy, feature extraction is performed on the multimodal diagnosis and treatment data to obtain fusion features corresponding to the multimodal diagnosis and treatment data, including: For each of the unimodal sub-data, according to the type of the unimodal sub-data, a feature extraction network corresponding to the type of the unimodal sub-data is used to perform feature extraction on the unimodal sub-data to obtain unimodal features corresponding to each of the unimodal sub-data; Obtaining a stage parameter representing the diagnosis and treatment stage of the target patient; Converting the phase parameters into a numerical vector to obtain a phase vector; Determine the modal weight corresponding to each of the unimodal features according to the unimodal features and the stage parameters: α m =Softmax(W attn ·Concat(F1,…,F M )+W stage ·OneHot(s)) Where, α m Represents the modal weight corresponding to the single modal feature with mode m, W attn 、W stage is a trainable weight matrix, Concat represents the concatenation operation, and OneHot(s) represents the phase vector; According to the modality weights, the single modality features are weightedly fused to obtain the fused features: In the formula, H represents the fusion feature, M represents the number of modalities, and F m Represents a unimodal feature with mode m.

3. The method according to claim 2, characterized in that For each of the unimodal sub-data, according to the type of the unimodal sub-data, feature extraction is performed on the unimodal sub-data using a feature extraction network corresponding to the type of the unimodal sub-data to obtain unimodal features corresponding to each of the unimodal sub-data, including: For each of the unimodal sub-data, when the type of the unimodal sub-data is the first type representing text, performing feature extraction on the unimodal sub-data by using a preset bidirectional transformer feature extraction network to obtain the unimodal feature; When the type of the unimodal sub-data is the second type representing an image, performing feature extraction on the unimodal sub-data by using a preset residual neural feature extraction network to obtain the unimodal feature; When the type of the unimodal sub-data is the third type representing the document, feature extraction is performed on the unimodal sub-data through a preset multi-layer perception feature extraction network to obtain the unimodal feature.

4. The method according to claim 1, wherein Filtering target data from a preset database based on the fusion features includes: Performing vector conversion on the data in the preset database to obtain multiple reference vectors; Determine, based on the fused feature and the multiple reference vectors, a cosine similarity between the fused feature and each of the multiple reference vectors: Where S k represents cosine similarity, H represents fusion features, represents the reference vector; A plurality of target vectors, among the plurality of reference vectors, whose cosine similarity with the fusion feature is greater than or equal to a preset threshold, are determined as the target data.

5. The method according to claim 1, characterized in that According to the fusion features and the target data, a reference solution corresponding to the target data is generated by using a preset reference solution generation strategy, including: Obtaining a risk factor representing the risk of the target patient's physical condition; According to the fusion feature, the target data and the risk factor, the reference solution corresponding to the target data is generated by a preset large language model.

6. The method according to claim 5, characterized in that The loss function of the preset large language model is as follows: L=λ1·CE(D raw ,D gt )+λ2·(β·R) Where L represents the loss function, λ1 and λ2 are risk balance coefficients, β represents the risk factor, R represents the risk prediction probability, CE(D raw ,D gt ) refers to the reference scheme D raw and the preset true label D gt The cross entropy loss.

7. The method according to claim 1, characterized in that The method further comprises: According to the preset causal correction strategy, the reference plan is corrected to obtain a corrected reference plan as the target plan.

8. The method according to claim 7, characterized in that According to the preset causal correction strategy, the reference solution is corrected to obtain a corrected reference solution as the target solution, including: Constructing a causal graph, the causal graph including a cause node representing the pathological characteristics of the target patient, an outcome node representing the diagnosis and treatment reference plan, and a confounding factor representing an influencing factor of the outcome node; According to the causal diagram, a causal intervention operation is performed on the reference solution to obtain a reference solution after causal intervention: Where D final represents the target solution, P(D final |do(H)) represents the causal intervention operation, D raw represents the result node, H represents the cause node, C represents the confounding factor, R prior (C) represents the probability distribution of confounding factors in the preset population.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory coupled to each other, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the electronic device executes the method according to any one of claims 1 to 8.

10. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 8.