Data analysis method based on multi-modal data fusion and related equipment
Through multimodal data fusion technology, combined with voice, image and medical record data, and using Transformer models and knowledge graphs to generate diagnostic suggestions, the problems of novice doctors' rapid diagnosis and patients' difficulty in understanding are solved, and the efficiency and accuracy of medical diagnosis are improved.
Patent Information
- Application Number
- CN202510651179.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-23
AI Technical Summary
Existing medical diagnosis is inefficient. Novice doctors find it difficult to make quick and accurate diagnoses in complex situations, and patients find it difficult to understand X-ray reports, resulting in large diagnostic errors.
A multimodal data fusion method is adopted to obtain the patient's voice information and medical imaging information through natural language processing technology, and the memory-driven Transformer model is used for analysis. Combined with the historical case knowledge graph and the domain expert knowledge base, explainable diagnostic recommendations are generated.
It improves the diagnostic efficiency and accuracy of novice doctors, provides explainable diagnostic reports, reduces diagnostic errors, and improves patients' understanding of diagnostic results.
Smart Images

Figure CN120690418A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of signal processing technology, and in particular to a data analysis method based on multimodal data fusion and related equipment. Background Art
[0002] With the rapid development of artificial intelligence and machine learning technologies, medical diagnosis based on intelligent assistance systems has shown great potential in improving medical efficiency and reducing diagnostic errors. For novice doctors, making a quick and accurate initial diagnosis amidst complex information and tight timelines is a pressing challenge. Furthermore, many patients struggle to understand the technical terminology contained in X-ray reports. Patients lacking a medical background often struggle to communicate effectively with their doctors, leading to misunderstandings of diagnostic information. Summary of the Invention
[0003] In view of this, the present invention provides a data analysis method and related equipment based on multimodal data fusion, which are used to solve the problems of low efficiency and large diagnostic errors in existing medical treatment.
[0004] In a first aspect, an embodiment of the present invention provides a data analysis method based on multimodal data fusion, the method comprising:
[0005] Based on the target person's voice information, obtain the target person's main complaint data through natural language processing technology;
[0006] Analyzing the medical imaging information of a target person using a preset deep learning model to obtain a radiology report of the target person, wherein the preset deep learning model is a memory-driven Transformer model, that is, a model obtained by incorporating relational memory into the Transformer model through memory-driven conditional layer normalization;
[0007] Inputting the chief complaint data and the radiology report into an initial diagnosis model based on a single-layer multi-head attention mechanism to obtain an initial diagnosis result;
[0008] Obtain the historical medical data of the target person and construct a historical case knowledge graph based on the historical medical data;
[0009] The initial diagnosis result and the historical case knowledge graph are embedded and learned using a medical consultation model to obtain a diagnosis suggestion, wherein the medical consultation model is a Transformer model;
[0010] The diagnosis suggestion is input into the domain expert knowledge base to verify the diagnosis suggestion. If the verification passes, the diagnosis suggestion is output; if the verification fails, the medical consultation model is updated based on the verification result.
[0011] Optionally, the step of obtaining the target person's main complaint data by using natural language processing technology based on the target person's voice information includes:
[0012] Collecting voice data of the target person and preprocessing the voice data to obtain the voice information, wherein the preprocessing includes at least one of denoising, enhancement, and framing;
[0013] The speech information of the target person is used to obtain the main complaint data in text format through the acoustic modeling method in natural language processing technology.
[0014] Optionally, the step of analyzing the medical imaging information of the target person by using a preset deep learning model to obtain a radiology report of the target person includes:
[0015] Extracting features from the medical imaging information of the target person to obtain image feature data, and inputting the image feature data into an encoder in the preset deep learning model to obtain encoded data;
[0016] The encoded data is input into a memory-driven decoder in the preset deep learning model to obtain a radiology report for the medical imaging information, where the data format of the radiology report is a text format.
[0017] Optionally, the step of inputting the encoded data into a memory-driven decoder in the preset deep learning model to obtain a radiology report for the medical imaging information includes:
[0018] controlling a memory-driven decoder to calculate a mean and a standard deviation according to the encoded data, and normalizing the mean and the standard deviation to obtain normalized data;
[0019] Controlling a memory-driven decoder to read a memory vector from the relational memory according to the normalized data, and calculating a parameter increment according to the memory vector through conditional layer normalization;
[0020] Controlling a memory-driven decoder to calculate a scaling parameter and an offset parameter according to the parameter increment, and updating the normalized data according to the scaling parameter and the offset parameter, so that the memory-driven decoder obtains a radiology report for the medical imaging information according to the updated normalized data.
[0021] Optionally, the step of inputting the chief complaint data and the radiology report into a preliminary diagnosis model based on a single-layer multi-head attention mechanism to obtain a preliminary diagnosis result includes:
[0022] The chief complaint data and the radiology report are cleaned, and the cleaned chief complaint data and the cleaned radiology report are input into the initial diagnosis model based on the single-layer multi-head attention mechanism to obtain the initial diagnosis result, wherein the mathematical representation of the initial diagnosis model is:
[0023] E=Transformer(X)=[e1,e2,…,e n ];
[0024]
[0025] Among them, E is the initial diagnosis result, Q i is the query vector of the i-th word, K j is the key vector of the jth word, V j is the value vector of the jth word, H is the additive attention score, α ij Indicates the strength of the relationship between the i-th word and the j-th word, e i Represents the embedding vector of each word, and the text data is X=[x1,x2,…,x n ],x i is the word embedding representation.
[0026] Optionally, the step of obtaining the historical medical data of the target person and constructing a historical case knowledge graph based on the historical medical data includes:
[0027] Performing entity recognition on the historical medical data to obtain entity data;
[0028] Extracting relationships from the historical medical data based on the entity data to obtain relationship data for the entity data;
[0029] Performing text splicing on the entity data and the relationship data through a Transformer model to obtain comprehensive representation data;
[0030] Based on the comprehensive representation data, using embedding vector learning technology to generate a disease vector representation;
[0031] The disease embedding vector and symptom embedding vector in the disease vector representation are used as nodes to construct a historical case knowledge graph.
[0032] On the other hand, an embodiment of the present application provides a data analysis system based on multimodal data fusion, characterized in that the system includes:
[0033] The first information collection module is used to obtain the target person's main complaint data through natural language processing technology based on the target person's voice information;
[0034] A second information acquisition module is configured to analyze the medical imaging information of the target person using a preset deep learning model to obtain a radiology report of the target person. The preset deep learning model is a memory-driven Transformer model, i.e., a model obtained by incorporating relational memory into the Transformer model through memory-driven conditional layer normalization;
[0035] A first prediction module is configured to input the chief complaint data and the radiology report into an initial diagnosis model based on a single-layer multi-head attention mechanism to obtain an initial diagnosis result;
[0036] The third information collection module is used to obtain the historical medical data of the target person and build a historical case knowledge graph based on the historical medical data;
[0037] The second prediction module is used to use the medical consultation model to embed the initial diagnosis results and the historical case knowledge graph to obtain diagnosis suggestions, and the medical consultation model is a Transformer model;
[0038] The verification module is used to input the diagnostic suggestion into the domain expert knowledge base to verify the diagnostic suggestion. If the verification passes, the diagnostic suggestion is output; if the verification fails, the medical consultation model is updated based on the verification result.
[0039] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the data analysis method based on multimodal data fusion as described above is implemented.
[0040] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the data analysis method based on multimodal data fusion as described above.
[0041] The technical solution of the embodiment of the present invention is to obtain the chief complaint data of the target person through natural language processing technology based on the voice information of the target person; analyze the medical imaging information of the target person through a preset deep learning model to obtain the radiology report of the target person, and the preset deep learning model is a memory-driven Transformer model, that is, a model obtained by merging the relational memory into the Transformer model through memory-driven conditional layer normalization; input the chief complaint data and the radiology report into the initial diagnosis model based on the single-layer multi-head attention mechanism to obtain the initial diagnosis result; obtain the historical medical data of the target person, and construct a historical case knowledge graph based on the historical medical data; use the medical model to embed the initial diagnosis result and the historical case knowledge graph to obtain a diagnosis suggestion, and the medical model is a Transformer model; input the diagnosis suggestion into the domain expert knowledge base to verify the diagnosis suggestion, and if the verification passes, output the diagnosis suggestion; if the verification fails, update the medical model based on the test result. It combines multimodal input, including the patient's real-time voice description, the patient's X-ray images provided by the hospital monitoring agency, and the patient's historical medical records. Through intelligent processing of three modal data of voice, image, and text and with the help of large models to assist model optimization, it helps novice doctors quickly understand the condition, provide patients with explainable report generation, and improve the accuracy and explainability of auxiliary diagnosis of images. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0043] in:
[0044] Figure 1 Schematic diagram of a flow chart of a data analysis method based on multimodal data fusion in an embodiment of the present invention;
[0045] Figure 2 4 is a logical framework diagram of a data analysis method based on multimodal data fusion in an embodiment of the present invention;
[0046] Figure 3 This is a flowchart of the operation of a preset deep learning model in a data analysis method based on multimodal data fusion in an embodiment of the present invention;
[0047] Figure 4Schematic diagram of a historical case knowledge graph in a data analysis method based on multimodal data fusion in an embodiment of the present invention;
[0048] Figure 5 1 is a flow chart of multimodal data fusion in a data analysis method based on multimodal data fusion in an embodiment of the present invention;
[0049] Figure 6 Flowchart of a verification process in a data analysis method based on multimodal data fusion according to an embodiment of the present invention;
[0050] Figure 7 Schematic diagram of the structure of a data analysis system based on multimodal data fusion in an embodiment of the present invention;
[0051] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0052] Figure 9 It is a structural diagram of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0054] like Figure 1 and Figure 2 As shown, a data analysis method based on multimodal data fusion according to an embodiment of the present invention specifically includes the following steps:
[0055] S110, obtaining the target person's main complaint data through natural language processing technology based on the target person's voice information;
[0056] In a possible implementation, the step of obtaining the target person's main complaint data by using natural language processing technology based on the target person's voice information includes:
[0057] Collecting voice data of the target person and preprocessing the voice data to obtain the voice information, wherein the preprocessing includes at least one of denoising, enhancement, and framing;
[0058] The speech information of the target person is used to obtain the main complaint data in text format through the acoustic modeling method in natural language processing technology.
[0059] For example, when a patient provides a voice description to a doctor, the patient's voice is analyzed in real time. Using speech recognition, processing, and acoustic modeling technologies, the voice is converted into text. Specifically, the text content includes the patient's description of symptoms, such as the location, nature, and duration of the pain.
[0060] S120. Analyze the medical imaging information of the target person using a preset deep learning model to obtain a radiology report of the target person, where the preset deep learning model is a memory-driven Transformer model, i.e., a model obtained by incorporating relational memory into the Transformer model through memory-driven conditional layer normalization;
[0061] For example, X-ray images manually uploaded by patients and doctors will be analyzed by a deep learning model (preset deep learning model) to identify potential lesion areas in the images and generate corresponding text reports (radiology reports), extracting key information from medical images to assist doctors in determining the location and severity of lesions and provide image analysis support for subsequent diagnosis.
[0062] S130, inputting the chief complaint data and the radiology report into an initial diagnosis model based on a single-layer multi-head attention mechanism to obtain an initial diagnosis result;
[0063] In one possible implementation, the step of inputting the chief complaint data and the radiology report into an initial diagnosis model based on a single-layer multi-head attention mechanism to obtain an initial diagnosis result includes:
[0064] The chief complaint data and the radiology report are cleaned, and the cleaned chief complaint data and the cleaned radiology report are input into the initial diagnosis model based on the single-layer multi-head attention mechanism to obtain the initial diagnosis result, wherein the mathematical representation of the initial diagnosis model is:
[0065] E=Transformer(X)=[e1,e2,…,e n ];
[0066]
[0067] Among them, E is the initial diagnosis result, Q i is the query vector of the i-th word, K j is the key vector of the jth word, V j is the value vector of the jth word, H is the additive attention score, α ij Indicates the strength of the relationship between the i-th word and the j-th word, e i Represents the embedding vector of each word, and the text data is X=[x1,x2,…,x n ],x i is the word embedding representation.
[0068] For example, a simplified Transformer model is used to process the fused text data using a single-layer multi-head attention mechanism. This attention mechanism enables the model to identify relationships between words and extract valuable information. Each word's embedding vector is generated through a multi-layer self-attention network, and the final disease vector representation serves as the core output of disease representation learning.
[0069] S140: Obtain historical medical data of the target person, and construct a historical case knowledge graph based on the historical medical data;
[0070] For example, the historical medical data includes information such as the patient's (the target person's) past health records, past treatment plans, and disease development history.
[0071] S150, using a medical consultation model to perform embedding learning on the initial diagnosis result and the historical case knowledge graph to obtain a diagnosis suggestion, wherein the medical consultation model is a Transformer model;
[0072] For example, Figure 5 As shown in the figure, the patient's voice description, X-ray analysis report, and historical medical records are cleaned and integrated into a unified dataset. First, the patient's voice and X-ray text information are concatenated to form a comprehensive symptom description. Next, the Transformer model is used to perform embedding learning on this fused text data and the knowledge graph of past cases, generating a vector representation of the condition. This is then inferred to provide a diagnostic recommendation.
[0073] S160. Input the diagnosis suggestion into the domain expert knowledge base to verify the diagnosis suggestion. If the verification passes, output the diagnosis suggestion; if the verification fails, update the medical consultation model based on the verification result.
[0074] For example, Figure 6 As shown, by integrating diagnostic recommendations with the domain expert knowledge base, the latest medical research results and professional recommendations can be dynamically obtained, which can continuously update and optimize diagnostic rules, provide more accurate and personalized diagnostic results, and feed back the generated diagnostic results to experts for review and correction to ensure that the output recommendations are consistent with the actual situation.
[0075] A comprehensive diagnostic report is generated based on the patient's voice description, X-ray analysis results, historical case data, and support from an expert knowledge base. The report includes a condition analysis, a diagnosis of potential illnesses, and corresponding treatment recommendations. The report is presented in concise and easy-to-understand language to help patients understand their health status. It also provides support for novice doctors to help them interpret diagnostic results and improve diagnostic efficiency.
[0076] The target person's chief complaint data is obtained through natural language processing technology based on the target person's voice information; the target person's medical imaging information is analyzed through a preset deep learning model to obtain the target person's radiology report, and the preset deep learning model is a memory-driven Transformer model, that is, a model obtained by merging the relational memory into the Transformer model through memory-driven conditional layer normalization; the chief complaint data and the radiology report are input into the initial diagnosis model based on the single-layer multi-head attention mechanism to obtain the initial diagnosis result; the target person's historical medical data is obtained, and a historical case knowledge graph is constructed based on the historical medical data; the initial diagnosis result and the historical case knowledge graph are embedded and learned using the medical model to obtain a diagnosis suggestion, and the medical model is a Transformer model; the diagnosis suggestion is input into the domain expert knowledge base to verify the diagnosis suggestion, and if the verification passes, the diagnosis suggestion is output; if the verification fails, the medical model is updated based on the test result. It combines multimodal input, including the patient's real-time voice description, the patient's X-ray images provided by the hospital monitoring agency, and the patient's historical medical records. Through intelligent processing of three modal data of voice, image, and text and with the help of large models to assist model optimization, it helps novice doctors quickly understand the condition, provide patients with explainable report generation, and improve the accuracy and explainability of auxiliary diagnosis of images.
[0077] In one possible implementation, the step of analyzing the medical imaging information of the target person using a preset deep learning model to obtain a radiology report of the target person includes:
[0078] Extracting features from the medical imaging information of the target person to obtain image feature data, and inputting the image feature data into an encoder in the preset deep learning model to obtain encoded data;
[0079] The encoded data is input into a memory-driven decoder in the preset deep learning model to obtain a radiology report for the medical imaging information, where the data format of the radiology report is a text format.
[0080] Specifically, the step of inputting the encoded data into a memory-driven decoder in the preset deep learning model to obtain a radiology report for the medical imaging information includes:
[0081] controlling a memory-driven decoder to calculate a mean and a standard deviation according to the encoded data, and normalizing the mean and the standard deviation to obtain normalized data;
[0082] Controlling a memory-driven decoder to read a memory vector from the relational memory according to the normalized data, and calculating a parameter increment according to the memory vector through conditional layer normalization;
[0083] Controlling a memory-driven decoder to calculate a scaling parameter and an offset parameter according to the parameter increment, and updating the normalized data according to the scaling parameter and the offset parameter, so that the memory-driven decoder obtains a radiology report for the medical imaging information according to the updated normalized data.
[0084] For example, Figure 3 As shown in Figure 1, the relational memory (RM) mechanism is used to record information from the previous generation process, and the memory-driven conditional layer normalization (MCLN) is used to incorporate the relational memory into the Transformer. The core idea of MCLN is: in the traditional Transformer, each decoding step is adjusted by conditional layer normalization (scaling parameter γ and offset parameter β), thereby improving generalization ability and performance. t The output of is fed to γ and β to merge the relational memory to enhance the decoding performance of Transformer, which can be expressed as:
[0085]
[0086] Predict Δγ through MLP network t Changes:
[0087] Δγ t =MLP(M t )
[0088] Finally, using the updated Δγ t To adjust the output results and improve the generation ability of the model.
[0089] In a possible implementation, the step of obtaining the historical medical data of the target person and constructing a historical case knowledge graph based on the historical medical data includes:
[0090] Performing entity recognition on the historical medical data to obtain entity data;
[0091] Extracting relationships from the historical medical data based on the entity data to obtain relationship data for the entity data;
[0092] Performing text splicing on the entity data and the relationship data through a Transformer model to obtain comprehensive representation data;
[0093] Based on the comprehensive representation data, using embedding vector learning technology to generate a disease vector representation;
[0094] The disease embedding vector and symptom embedding vector in the disease vector representation are used as nodes to construct a historical case knowledge graph.
[0095] For example, key entities (such as disease name, treatment method, drug information, medical history information, etc.) and the relationships between them are extracted from historical case data to construct a graph containing medical entities and their relationships.
[0096] A Transformer model is used to perform text concatenation on historical case data, fusing relevant data extracted from the cases (such as symptom descriptions and features) into a comprehensive representation. Embedded vector learning techniques are then used to generate a vector representation for each case, forming a symptom vector representation. This vector representation reflects the semantic information of the case data and serves as the basic input for subsequent auxiliary diagnosis.
[0097] The embedded vectors of diseases and symptoms learned from historical case data are used as nodes to construct a case knowledge graph. The edges in the graph represent the relationships between different diseases, such as comorbidity relationships, drug treatment relationships, disease progression, etc. The graph supports dynamic updates. As more case data is added, the graph content is continuously enriched, thereby improving the model's ability to infer future cases. The structure of the historical case knowledge graph is as follows: Figure 4 As shown, it contains the relationships between nodes such as different diseases, symptoms, and treatment methods.
[0098] In one possible implementation, Figure 7 As shown, an embodiment of the present application provides a data analysis system based on multimodal data fusion, characterized in that the system includes:
[0099] The first information collection module 201 is used to obtain the target person's main complaint data through natural language processing technology based on the target person's voice information;
[0100] A second information acquisition module 202 is configured to analyze the medical imaging information of the target person using a preset deep learning model to obtain a radiology report of the target person, wherein the preset deep learning model is a memory-driven Transformer model, i.e., a model obtained by incorporating relational memory into the Transformer model through memory-driven conditional layer normalization;
[0101] The first prediction module 203 is configured to input the chief complaint data and the radiology report into a preliminary diagnosis model based on a single-layer multi-head attention mechanism to obtain a preliminary diagnosis result;
[0102] The third information collection module 204 is used to obtain the historical medical data of the target person and construct a historical case knowledge graph based on the historical medical data;
[0103] The second prediction module 205 is used to use the medical consultation model to embed the initial diagnosis result and the historical case knowledge graph to obtain a diagnosis suggestion, and the medical consultation model is a Transformer model;
[0104] The verification module 206 is used to input the diagnosis suggestion into the domain expert knowledge base to verify the diagnosis suggestion. If the verification passes, the diagnosis suggestion is output; if the verification fails, the medical consultation model is updated based on the verification result.
[0105] In one possible implementation, Figure 8 As shown, the embodiment of the present application provides a terminal device 300, including: a memory 310, a processor 320 and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, the following are achieved: for the voice information of the target person, the main complaint data of the target person is obtained by natural language processing technology; the medical imaging information of the target person is analyzed by a preset deep learning model to obtain a radiology report of the target person, and the preset deep learning model is a memory-driven Transformer model, that is, the relational memory is merged into the Transformer model through memory-driven conditional layer normalization. The model obtained after the conversion of the chief complaint data and the radiology report into the initial diagnosis model based on the single-layer multi-head attention mechanism is constructed to obtain the initial diagnosis result; the historical medical data of the target person is obtained, and a historical case knowledge graph is constructed based on the historical medical data; the initial diagnosis result and the historical case knowledge graph are embedded in the medical model to obtain a diagnosis suggestion, and the medical model is a Transformer model; the diagnosis suggestion is input into the domain expert knowledge base to verify the diagnosis suggestion, and if the verification passes, the diagnosis suggestion is output; if the verification fails, the medical model is updated based on the test result.
[0106] In one possible implementation, Figure 9As shown, the embodiment of the present application provides a computer-readable storage medium 400, on which a computer program 411 is stored. When the computer program 411 is executed by a processor, the following are achieved: for the voice information of the target person, the main complaint data of the target person is obtained by natural language processing technology; the medical imaging information of the target person is analyzed by a preset deep learning model to obtain the radiology report of the target person, and the preset deep learning model is a memory-driven Transformer model, that is, the model obtained by merging the relational memory into the Transformer model through memory-driven conditional layer normalization type; input the chief complaint data and the radiology report into the initial diagnosis model based on the single-layer multi-head attention mechanism to obtain the initial diagnosis result; obtain the historical medical data of the target person, and construct a historical case knowledge graph based on the historical medical data; use the medical model to embed the initial diagnosis result and the historical case knowledge graph to obtain a diagnosis suggestion, and the medical model is a Transformer model; input the diagnosis suggestion into the domain expert knowledge base to verify the diagnosis suggestion, and if the verification passes, output the diagnosis suggestion; if the verification fails, update the medical model based on the test result.
[0107] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0108] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0109] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process of the above-mentioned method embodiment by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can at least include: any entity or device capable of carrying computer program code to the camera / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, mobile hard drive, magnetic disk, or optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunication signals.
[0110] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0111] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0112] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0113] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0114] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
[0115] The above disclosure is merely a preferred embodiment of the present invention and certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.
Claims
1. A data analysis method based on multimodal data fusion, characterized in that: include: Based on the target person's voice information, obtain the target person's main complaint data through natural language processing technology; Analyzing the medical imaging information of a target person using a preset deep learning model to obtain a radiology report of the target person, wherein the preset deep learning model is a memory-driven Transformer model, that is, a model obtained by incorporating relational memory into the Transformer model through memory-driven conditional layer normalization; Inputting the chief complaint data and the radiology report into an initial diagnosis model based on a single-layer multi-head attention mechanism to obtain an initial diagnosis result; Obtain the historical medical data of the target person and construct a historical case knowledge graph based on the historical medical data; The initial diagnosis result and the historical case knowledge graph are embedded and learned using a medical consultation model to obtain a diagnosis suggestion, wherein the medical consultation model is a Transformer model; Inputting the diagnostic suggestion into a domain expert knowledge base to verify the diagnostic suggestion, and outputting the diagnostic suggestion if the verification passes; If the verification fails, the medical consultation model is updated based on the verification result.
2. The data analysis method based on multimodal data fusion according to claim 1, characterized in that: The step of obtaining the target person's main complaint data by using natural language processing technology based on the target person's voice information includes: Collecting voice data of the target person and preprocessing the voice data to obtain the voice information, wherein the preprocessing includes at least one of denoising, enhancement, and framing; The speech information of the target person is used to obtain the main complaint data in text format through the acoustic modeling method in natural language processing technology.
3. The data analysis method based on multimodal data fusion according to claim 1, characterized in that: The step of analyzing the medical imaging information of the target person by using a preset deep learning model to obtain a radiology report of the target person includes: Extracting features from the medical imaging information of the target person to obtain image feature data, and inputting the image feature data into an encoder in the preset deep learning model to obtain encoded data; The encoded data is input into a memory-driven decoder in the preset deep learning model to obtain a radiology report for the medical imaging information, where the data format of the radiology report is a text format.
4. The data analysis method based on multimodal data fusion according to claim 3, characterized in that: The step of inputting the encoded data into a memory-driven decoder in the preset deep learning model to obtain a radiology report for the medical imaging information includes: controlling a memory-driven decoder to calculate a mean and a standard deviation according to the encoded data, and normalizing the mean and the standard deviation to obtain normalized data; Controlling a memory-driven decoder to read a memory vector from the relational memory according to the normalized data, and calculating a parameter increment according to the memory vector through conditional layer normalization; Controlling a memory-driven decoder to calculate a scaling parameter and an offset parameter according to the parameter increment, and updating the normalized data according to the scaling parameter and the offset parameter, so that the memory-driven decoder obtains a radiology report for the medical imaging information according to the updated normalized data.
5. The data analysis method based on multimodal data fusion according to claim 1, characterized in that: The step of inputting the chief complaint data and the radiology report into the initial diagnosis model based on the single-layer multi-head attention mechanism to obtain the initial diagnosis result includes: The chief complaint data and the radiology report are cleaned, and the cleaned chief complaint data and the cleaned radiology report are input into the initial diagnosis model based on the single-layer multi-head attention mechanism to obtain the initial diagnosis result, wherein the mathematical representation of the initial diagnosis model is: E=Transformer(X)=[e1,e2,…,e n ]; Among them, E is the initial diagnosis result, Q i is the query vector of the i-th word, K j is the key vector of the jth word, V j is the value vector of the jth word, H is the additive attention score, α ij Indicates the strength of the relationship between the i-th word and the j-th word, e i Represents the embedding vector of each word, and the text data is X=[x1,x2,…,x n ],x i is the word embedding representation.
6. The data analysis method based on multimodal data fusion according to claim 1, characterized in that: The step of obtaining the historical medical data of the target person and constructing a historical case knowledge graph based on the historical medical data includes: Performing entity recognition on the historical medical data to obtain entity data; Extracting relationships from the historical medical data based on the entity data to obtain relationship data for the entity data; Performing text splicing on the entity data and the relationship data through a Transformer model to obtain comprehensive representation data; Based on the comprehensive representation data, using embedding vector learning technology to generate a disease vector representation; The disease embedding vector and symptom embedding vector in the disease vector representation are used as nodes to construct a historical case knowledge graph.
7. A data analysis system based on multimodal data fusion, characterized in that: The system comprises: The first information collection module is used to obtain the target person's main complaint data through natural language processing technology based on the target person's voice information; A second information acquisition module is configured to analyze the medical imaging information of the target person using a preset deep learning model to obtain a radiology report of the target person. The preset deep learning model is a memory-driven Transformer model, i.e., a model obtained by incorporating relational memory into the Transformer model through memory-driven conditional layer normalization; A first prediction module is configured to input the chief complaint data and the radiology report into an initial diagnosis model based on a single-layer multi-head attention mechanism to obtain an initial diagnosis result; The third information collection module is used to obtain the historical medical data of the target person and build a historical case knowledge graph based on the historical medical data; The second prediction module is used to use the medical consultation model to embed the initial diagnosis results and the historical case knowledge graph to obtain diagnosis suggestions, and the medical consultation model is a Transformer model; The verification module is used to input the diagnostic suggestion into the domain expert knowledge base to verify the diagnostic suggestion. If the verification passes, the diagnostic suggestion is output; if the verification fails, the medical consultation model is updated based on the verification result.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the data analysis method based on multimodal data fusion according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the data analysis method based on multimodal data fusion according to any one of claims 1 to 6 is implemented.