Model training method, system and equipment based on multi-modal data fusion and medium
By preprocessing and extracting key data on multimodal data, combined with the deep learning model architecture of dynamic routing layer and quantum computing layer, the problems of large computing resource consumption and poor generalization capabilities in training multimodal data fusion model are solved, and more efficient data fusion and model performance improvement are achieved.
Patent Information
- Application Number
- CN202510355689.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-08-01
AI Technical Summary
The computing resources consume huge amounts of time during training of multimodal data fusion models, the model is prone to overfitting, poor generalization ability, and data redundancy and irrelevant features affect the model performance.
The multimodal data is preprocessed, key data is extracted and data fusion is carried out, and deep learning models are trained using deep learning models, including data cleaning, standardization, missing value processing and multiple dimensionality reduction operations, and a deep learning model architecture combining dynamic routing layer and quantum computing layer.
It improves the performance of the model on the training data, enhances generalization ability, reduces the risk of overfitting, and improves the feasibility and data interpretability of cross-modal data fusion.
Smart Images

Figure CN120408185A_ABST
Abstract
Description
Background Art
[0002] Multimodal data integrates multiple data types, such as images, texts, audios, and videos. The dimensions and scale of multimodal data are usually large. When directly performing model training, it will result in huge consumption of computing resources. Moreover, multimodal data contains a large number of features, and some of these features may be irrelevant or redundant to the task. Directly using all data for training may cause the model to be too complex and easily overfit to the noise or specific distribution of the training data, thereby reducing the generalization ability of the model and performing poorly in test data or actual application scenarios. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a model training method, system, device, and medium based on multimodal data fusion in view of the deficiencies of the prior art, specifically as follows:
[0004] 1) In the first aspect, the present invention provides a model training method based on multimodal data fusion, and the specific technical solution is as follows: [[ID=1>
[0005] Preprocess the collected multimodal data to obtain preprocessed multimodal data. The multimodal data includes clinical symptoms, physical sign data, and treatment diagnosis results, or the multimodal data includes industrial equipment operation failure data and industrial equipment failure cause data;
[0006] Obtain the key data in each modal data of the preprocessed multimodal data and perform data fusion to obtain fused data;
[0007] Use the fused data to train a preset deep learning model to obtain a trained deep learning model.
[0008] The beneficial effects of a model training method based on multimodal data fusion provided by the present invention are as follows:
[0009] On the one hand, key data usually contains the features most relevant to the task. After removing irrelevant and redundant information, the model can focus more on learning important patterns and relationships. This helps to improve the performance of the model on the training data and enhance its generalization ability, enabling it to perform better on unseen data and reducing the risk of overfitting. On the other hand, the extraction of key data helps to better align and fuse information from different modalities. By selecting complementary and relevant key features, different modality data can be more effectively fused together, improving the model's comprehensive understanding and utilization ability of multimodal data, and thus achieving better results in multimodal application scenarios. Moreover, the extracted key data is easier to understand and interpret, effectively enhancing the interpretability of the data.
[0010] Based on the above solution, an improvement can be made to the model training method based on multi-modal data fusion of the present invention as follows.
[0011] Further, preprocess the multi-modal data, including:
[0012] When the multi-modal data includes clinical symptoms, sign data, and treatment diagnosis results, perform data cleaning, standardization processing based on the Human Phenotype Ontology library, and missing value processing on the multi-modal data to obtain the preprocessed multi-modal data.
[0013] Further, preprocess the multi-modal data, including:
[0014] When the multi-modal data includes industrial equipment operation failure data and industrial equipment failure cause data, perform data cleaning and missing value processing on the multi-modal data to obtain the preprocessed multi-modal data.
[0015] Further, obtain the key data in each modal data of the preprocessed multi-modal data, including:
[0016] Perform at least two dimensionality reduction operations continuously on each modal data of the preprocessed multi-modal data to obtain the key data in each modal data of the preprocessed multi-modal data.
[0017] The beneficial effects of adopting the above further solution are: During the process of multiple dimensionality reduction, it is possible to more clearly understand the contribution of each feature to data representation and model performance at different stages. This helps to determine which features are truly important and which features can be ignored, so as to be more targeted in analyzing the role of key features during model interpretation, improving the accuracy and credibility of model interpretation. Moreover, multiple dimensionality reduction can gradually map data of different modalities to a closer low-dimensional space, reducing the differences between modalities and improving the feasibility and effect of cross-modal data fusion.
[0018] 2) On the second aspect, the present invention also provides a model training system based on multi-modal data fusion. The specific technical solution is as follows:
[0019] It includes a data preprocessing module, a data extraction and fusion module, and a model training module;
[0020] The data preprocessing module is used for: preprocessing the collected multi-modal data to obtain the preprocessed multi-modal data, where the multi-modal data includes clinical symptoms, sign data, and treatment diagnosis results, or the multi-modal data includes industrial equipment operation failure data and industrial equipment failure cause data;
[0021] The data extraction and fusion module is used for: obtaining the key data in each modal data of the preprocessed multi-modal data and performing data fusion to obtain the fused data;
[0022] The model training module is used to: train a preset deep learning model using the fused data to obtain a trained deep learning model.
[0023] Based on the above solution, a model training system based on multi-modal data fusion according to the present invention can be further improved as follows.
[0024] Further, the data preprocessing module is specifically used for:
[0025] When the multi-modal data includes clinical symptoms, physical signs data, and treatment diagnosis results, perform data cleaning, standardization processing based on the Human Phenotype Ontology library, and missing value processing on the multi-modal data to obtain preprocessed multi-modal data.
[0026] Further, the data preprocessing module is specifically used for:
[0027] When the multi-modal data includes industrial equipment operation failure data and industrial equipment failure cause data, perform data cleaning and missing value processing on the multi-modal data to obtain preprocessed multi-modal data.
[0028] Further, the data extraction and fusion module is also specifically used for:
[0029] Perform at least two dimensionality reduction operations continuously on each modal data in the preprocessed multi-modal data to obtain key data in each modal data in the preprocessed multi-modal data.
[0030] 3) Thirdly, the present invention also provides an electronic device, which includes a processor. The processor is coupled to a memory, and at least one computer program is stored in the memory. The at least one computer program is loaded and executed by the processor so that the electronic device implements any one of the above model training methods based on multi-modal data fusion.
[0031] 4) Fourthly, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above model training methods based on multi-modal data fusion.
[0032] It should be noted that for the beneficial effects obtained by the technical solutions and corresponding possible implementation manners of the second to fourth aspects of the present invention, reference can be made to the technical effects of the first aspect and its corresponding possible implementation manners above, which will not be elaborated here. Description of the Drawings
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments of the present invention:
[0034] Figure 1 Schematic flowchart of a model training method based on multi-modal data fusion according to an embodiment of the present invention;
[0035] Figure 2 Schematic structural diagram of a model training system based on multi-modal data fusion according to an embodiment of the present invention;
[0036] Figure 3 Schematic structural diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0037] The principles and features of the present invention are described below. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.
[0038] The technical solutions of the present invention and how the technical solutions of the present invention solve the above technical problems are described in detail below with specific embodiments. These several specific embodiments may be combined with each other. For the same or similar concepts or processes, they may not be repeated in some embodiments. The embodiments of the present invention will be described below with reference to the accompanying drawings.
[0039] As Figure 1 shown, a model training method based on multi-modal data fusion according to an embodiment of the present invention includes the following steps:
[0040] S1. Preprocess the collected multi-modal data to obtain preprocessed multi-modal data. The multi-modal data includes clinical symptoms, physical sign data, and treatment diagnosis results, or the multi-modal data includes industrial equipment operation failure data and industrial equipment failure cause data;
[0041] Among them, the modality of clinical symptoms can be X-ray images, CT images, MRI images, ultrasound images, PET images, electroencephalograms, electrocardiograms, etc. The physical sign data includes body temperature, respiratory rate, blood pressure, head and neck physical signs, chest physical signs, abdominal physical signs, spine and limb physical signs, and nervous system physical signs, etc. The corresponding modalities can include text and images. The modality of treatment diagnosis results can include text reports.
[0042] Among them, the collected multi-modal data includes: clinical symptoms, physical sign data, and treatment diagnosis results of multiple patients.
[0043] Among them, the industrial equipment operation fault data includes industrial equipment operation status fault data, industrial equipment control fault data, industrial equipment monitoring fault data, industrial equipment performance fault data, industrial equipment communication fault data and industrial equipment wear fault data. The industrial equipment operation status fault data includes abnormal changes in electrical parameters such as current, voltage, and power, such as overcurrent, undervoltage, and abnormal power, and also includes abnormal changes in vibration data, rotational speed data, swing amplitude data, etc., such as exceeding the vibration amplitude standard and abnormal fluctuations in rotational speed, and also includes equipment overheating system faults and cooling system faults, etc. The industrial equipment control fault data includes programmable logic controller (PLC) data and machine instruction execution error information, etc. The industrial equipment monitoring fault data includes image data and video data during equipment faults, abnormal sound data during equipment faults, and fault alarm information in the operation log. The industrial equipment performance fault data includes abnormal changes in performance indicators such as a decrease in the output power of industrial equipment, a decrease in production efficiency, and a decrease in product quality. The industrial equipment communication fault data includes communication-related fault data such as communication interruption between industrial equipment and data transmission errors.
[0044] The industrial equipment fault cause data includes: mechanical faults or electronic component failures caused by abnormal thermal expansion and contraction of industrial equipment components due to too high or too low ambient temperature, internal short circuits and corrosion caused by too high humidity, heat dissipation problems and poor contact caused by excessive dust accumulation, loosening, falling off or damage of industrial equipment components caused by external vibration sources, and also includes operation errors, improper maintenance, power supply problems, component wear and material aging.
[0045] Among them, the industrial equipment operation fault data includes industrial equipment operation status data, industrial equipment control data and industrial equipment monitoring data. The industrial equipment operation status data includes electrical parameters such as current, voltage, and power used to reflect the power consumption and electrical performance of the equipment, and also includes vibration data, rotational speed data, swing amplitude data and temperature data. The industrial equipment control data includes programmable logic controller data and machine instructions, etc. The industrial equipment monitoring data includes image data and video data of industrial equipment, and also includes sound data during the operation of industrial equipment and text data in the operation log. The text data in the operation log records information such as the operation process and maintenance records of the equipment, which helps to trace the operation history and problem root cause of the equipment.
[0046] Among them, the industrial equipment can be an engine motor or a medical device, etc., which can be set according to the actual situation. The multi-modal data includes the industrial equipment operation fault data and the industrial equipment fault cause data of the same industrial equipment.
[0047] S2. Obtain the key data in each modal data of the preprocessed multi-modal data and perform data fusion to obtain the fused data;
[0048] S3. Use the fused data to train a preset deep learning model to obtain a trained deep learning model.
[0049] Optionally, preprocess the multimodal data, including:
[0050] 1) When the multimodal data includes clinical symptoms, physical examination data, and treatment diagnosis results, perform data cleaning, standardization processing based on the Human Phenotype Ontology (HPO), and missing value processing on the multimodal data to obtain preprocessed multimodal data.
[0051] The Human Phenotype Ontology (HPO) is a standardized vocabulary used to describe phenotype characteristics related to human diseases. During the standardization processing based on the Human Phenotype Ontology, calculate the semantic similarity between the technical terms used in the clinical symptoms, physical examination data, and treatment diagnosis results and the standardized vocabulary in the Human Phenotype Ontology. When the semantic similarity between any technical term used in the clinical symptoms, physical examination data, and treatment diagnosis results and any standardized term in the standardized vocabulary exceeds the preset semantic similarity threshold, replace any technical term used in the clinical symptoms, physical examination data, and treatment diagnosis results with the standardized term. When there are technical terms that have not been replaced by the standardized vocabulary, they can be directly retained, or an application for inclusion can be submitted to the Human Phenotype Ontology so that the unreplaced technical terms can be included in the Human Phenotype Ontology.
[0052] 2) When the multimodal data includes industrial equipment operation failure data and industrial equipment failure cause data, perform data cleaning and missing value processing on the multimodal data to obtain preprocessed multimodal data.
[0053] Taking two-dimensionality reduction operations as an example, the process of obtaining key data in each modal data of the preprocessed multimodal data is described as follows:
[0054] S20. The first data dimensionality reduction operation:
[0055] S200. For each modality data in the pre-processed multi-modal data, perform data standardization and feature extraction in sequence. Here, data standardization means: for text, split the text by words or characters to form a vocabulary. The word segmentation method can adopt word segmentation methods based on string matching, understanding, and statistics. Represent each word in the vocabulary as a dense vector of a fixed length through word embedding (specifically, Word2Vec or GloVe, etc.), and then normalize it to the interval [0, 1] for subsequent model training and optimization. For numerical values, use the normalization method of maximum and minimum values for data standardization. For images, adjust all images to the same size, and normalize the pixel values from the range of 0 - 255 to the interval [0, 1] or [-1, 1], usually achieved by dividing by 255 or performing a linear transformation, thereby realizing data standardization for each image. Combine the values of features from different modality data into a feature matrix. Each row in the feature matrix represents a subject, and each column represents the value of a feature. Specifically:
[0056] 1) When the collected multi-modal data includes clinical symptom, physical sign data, and treatment diagnosis results of multiple patients, each disease is a subject. The features of the subject include: body temperature, respiratory rate, blood pressure, head and neck physical signs, chest physical signs, abdominal physical signs, spine and limb physical signs, nervous system physical signs, X-ray images, CT images, MRI images, ultrasound images, PET images, electroencephalograms, electrocardiograms, and treatment diagnosis results. The values of the features of the subject are: body temperature, respiratory rate, blood pressure, head and neck physical signs, chest physical signs, abdominal physical signs, spine and limb physical signs, nervous system physical signs, X-ray images, CT images, MRI images, ultrasound images, PET images, electroencephalograms, electrocardiograms, and treatment diagnosis results after data standardization processing.
[0057] 2) When the collected multi-modal data includes industrial equipment operation failure data and industrial equipment failure cause data of the same industrial equipment, each failure cause is a subject. The features of the subject include: current, voltage, power, vibration data, rotational speed data, swing amplitude data, temperature data, programmable logic controller data, machine instructions, image data, video data, sound data during industrial equipment operation, and industrial operation logs, etc.
[0058] S201. Calculate the covariance matrix of the feature matrix, perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalues corresponding to each feature, sort all the eigenvalues in ascending order, select the features corresponding to the top k eigenvalues as the initial target features, construct a new feature space based on all the initial target features, then select the values of each initial target feature from the feature matrix and project them into the new feature space to obtain the data after the first dimensionality reduction. The data after the first dimensionality reduction is the refined feature matrix. Each row in this refined feature matrix represents a subject, and each column represents the value of an initial target feature.
[0059] S21. Second data dimensionality reduction operation:
[0060] S210. Based on the values of each initial target feature of every two subjects, calculate the similarity between every two subjects to obtain a similarity matrix. Any element (i.e., the similarity value) in the similarity matrix represents the similarity between the two subjects corresponding to this element.
[0061] S211. Construct a joint probability distribution based on the similarity matrix and establish a conditional probability distribution in a low-dimensional space. The low-dimensional space is in comparison to the refined feature matrix, and the number of dimensions in the low-dimensional space is lower than the number of initial target features of the refined feature matrix. Minimize the KL divergence (Kullback-Leibler divergence) to make the conditional probability distribution in the low-dimensional space as close as possible to the joint probability distribution. Specifically, use optimization algorithms such as gradient descent to iteratively update the positions of the data points in the low-dimensional space until the convergence condition or the preset number of iterations is reached. Plot the optimized low-dimensional data points (the values of all the final target features corresponding to each subject) in the low-dimensional space to obtain the feature matrix refined again. Each row in this feature matrix refined again represents a subject, and each column represents the value of a final target feature.
[0062] During the process of multiple data dimensionality reductions, it is possible to more clearly understand the contribution of each feature to data representation and model performance at different stages. This helps to determine which features are truly important and which features can be ignored, so as to more specifically analyze the role of key features during model interpretation, improving the accuracy and credibility of model interpretation. Moreover, multiple dimensionality reductions can gradually map data of different modalities into a closer low-dimensional space, reducing the differences between modalities and improving the feasibility and effect of cross-modal data fusion.
[0063] In S2, obtaining the key data in each modality data of the preprocessed multi-modal data is the value of each subject and all corresponding final target features. The process of fusing any subject and the values of all corresponding final target features includes: respectively forming an array for each subject and the values of each corresponding final target feature to obtain multiple arrays. The fused data includes all arrays, and each array is equivalent to a sample. Based on all samples, a preset deep learning model is trained to obtain a trained deep learning model. Specifically:
[0064] 1) When the collected multi-modal data includes the clinical symptoms, physical signs data, and treatment diagnosis results of multiple patients, the trained deep learning model is a disease recognition model. After performing the above-mentioned multiple data dimensionality reduction operations on the clinical symptoms, physical signs data, and treatment diagnosis results, the obtained data is input into the disease recognition model to obtain a disease recognition result.
[0065] 2) When the collected multi-modal data includes the industrial equipment operation failure data and industrial equipment failure cause data of the same industrial equipment, the trained deep learning model is a fault diagnosis model. After performing the above-mentioned multiple data dimensionality reduction operations on the industrial equipment operation failure data and industrial equipment failure cause data of the industrial equipment, the obtained data is input into the fault diagnosis model to obtain a fault diagnosis result.
[0066] Among them, the preset deep learning model includes an input preprocessing layer, a feature extraction layer, a dynamic routing layer, a quantum computing layer, and a task output layer set in sequence. Specifically:
[0067] 1) The input preprocessing layer includes an adversarial enhancement layer and a dynamic resampling layer set in sequence. The adversarial enhancement layer generates directional perturbation noise. After applying the directional perturbation noise to the input sample, a perturbed sample is obtained. The dynamic resampling layer uses a cubic spline interpolation algorithm to perform length normalization processing on the perturbed sample.
[0068] Among them, the process of the adversarial enhancement layer generating directional perturbation noise is as follows:
[0069] The input sample (array) is normalized by L2 norm to ensure that the perturbation amplitude in the perturbation noise is controllable. The data after L2 norm normalization is input into the sign function to generate an adversarial perturbation noise with the same dimension as the sample, which is the directional perturbation noise. The perturbed sample obtained after applying the directional perturbation noise to the input sample can enable the preset deep learning model to learn the distribution characteristics of the sample, which helps to improve the performance of the trained preset deep learning model. Among them, the sign function is a mathematical function that returns the corresponding sign for any real number.
[0070] Among them, the dynamic resampling layer uses the cubic spline interpolation algorithm to perform length normalization on the perturbed samples, which can effectively preserve the key information in the perturbed samples, help improve the performance of the trained preset deep learning model, and has high computational efficiency.
[0071] 2) The feature extraction layer includes: a multi-scale fusion layer and a channel adjustment layer. The multi-scale fusion layer includes multiple parallel branches. The multiple parallel branches include an average pooling branch, an original feature branch, and a max pooling branch. The multi-scale fusion layer also includes a convolutional layer and a dilated convolution group. The output of the dynamic resampling layer is respectively input into the average pooling branch, the original feature branch, and the max pooling branch. After the outputs of the average pooling branch, the original feature branch, and the max pooling branch pass through the convolutional layer and the dilated convolution group, a weighted summation operation is realized. The channel adjustment layer performs global average pooling on the output result of the multi-scale fusion layer, obtains channel statistics, and generates channel attention weights through a fully connected layer to realize the dynamic scaling of feature channels, which can effectively strengthen key information and help improve the performance of the trained preset deep learning model.
[0072] Among them, the average pooling branch includes an average pooling layer. The original feature branch is a layer without additional operations. That is to say, the original feature branch is an empty layer and does not perform any operations on the output of the dynamic resampling layer, that is, directly outputs the output of the dynamic resampling layer. The max pooling branch includes a max pooling layer. The average pooling branch plays a role in smoothing noise and retaining global semantics. The original feature branch plays a role in ensuring the integrity of information. The max pooling branch plays a role in feature enhancement and enhancing anti-interference ability.
[0073] 3) The dynamic routing layer connects multiple sub-networks in the mixture of experts model based on the differentiable top-k routing mechanism, realizes the dynamic allocation between the outputs of the average pooling branch, the original feature branch, and the max pooling branch and the corresponding sub-networks. By introducing the mixture of experts model and dynamically allocating sub-networks, the adaptability of the preset deep learning model to different tasks can be effectively improved.
[0074] 4) The quantum computing layer encodes the output of each sub-network into a quantum state, obtains the expected value of the quantum state through the Pauli-Z operator, obtains a classical feature vector according to the expected value of the quantum state, and inputs the classical feature vector into the task output layer, which can effectively improve the computational efficiency and effectively improve feature interaction, and help improve the performance of the trained preset deep learning model.
[0075] 5) The task output layer includes a fully connected layer. After processing the output of the quantum computing layer through the fully connected layer, the disease diagnosis result or the fault diagnosis result is obtained.
[0076] The network structure and data processing method of the preset deep learning model proposed by the present invention both have the ability of dynamic adjustment, can adapt to different tasks and data distributions, and do not require manual re-design of the architecture. Moreover, through the routing mechanism (dynamic routing layer), on-demand allocation of computing resources can be achieved, improving computing efficiency. The quantum computing layer introduces the computing characteristics of the quantum state space, which can effectively enhance the modeling ability for high-dimensional non-linear relationships.
[0077] The present invention will be further described through the following embodiments. Specifically:
[0078] Rare diseases refer to diseases with extremely low incidence rates, and due to atypical clinical symptoms and difficult diagnosis, they are often overlooked or misdiagnosed. Since the gene mutations involved in rare diseases are diverse and the symptom manifestations are complex, traditional diagnostic methods based on a single gene or clinical symptoms are difficult to comprehensively cover. In recent years, with the establishment and development of the Human Phenotype Ontology (HPO), new ideas have been provided for the diagnosis of rare diseases. HPO is a standardized vocabulary used to describe phenotypic characteristics related to human diseases. By associating the clinical manifestations of patients with HPO, the diagnostic scope can be effectively narrowed, the diagnostic efficiency can be improved, and the diagnostic accuracy can be enhanced.
[0079] The present invention collects the clinical phenotype data of patients, compares it with the known rare disease phenotype data in the phenotype ontology library, and combines deep learning algorithms to establish a prediction model for rare diseases, thereby realizing early prediction and accurate diagnosis of rare diseases. Specifically, it includes the following steps:
[0080] 1) Data collection: Collect the phenotype data of patients, including but not limited to clinical symptoms, signs, laboratory test results, etc., to form a phenotype feature library of patients.
[0081] 2) Phenotype data processing: Preprocess the collected patient phenotype data, including data cleaning, standardization, and missing value processing, etc., in order to interface with the Human Phenotype Ontology library.
[0082] 3) Matching with the Human Phenotype Ontology library: Compare the similarity between the patient phenotype data and the known phenotype data in the Human Phenotype Ontology library. The Human Phenotype Ontology library is a knowledge base containing various disease phenotype information, including the phenotype characteristics of rare diseases.
[0083] 4) Construction of a rare disease prediction model: Based on the matching results of the phenotype data and the phenotype ontology library, combine machine learning and deep learning algorithms (such as neural networks or support vector machines, etc.) to establish a rare disease prediction model.
[0084] 5) Prediction result output: Analyze the phenotypic data of the input patient through the trained rare disease prediction model, sort according to similarity, and output the most likely rare disease prediction results, including possible disease types, probabilities, and recommended further diagnostic measures.
[0085] Suppose patient A shows multiple phenotypic characteristics, and the clinical symptoms include limb weakness, heart abnormalities, muscle atrophy, etc. First, enter A's phenotypic data through the data collection module. Then, preprocess the data through the data processing module, standardize and fill in the missing values. Next, the phenotypic ontology library matching module matches the phenotypic data of patient A with the information in the phenotypic ontology library. Finally, the prediction model construction module performs data analysis based on machine learning algorithms and outputs the rare disease prediction results, indicating that A may suffer from a certain rare muscle disease.
[0086] The beneficial effects of the present invention are as follows:
[0087] 1) Improve the accuracy of rare disease prediction: By combining the rich phenotypic data in the Human Phenotype Ontology library with the clinical phenotypic data of patients and using deep learning algorithms, the prediction accuracy of rare diseases can be improved, avoiding misdiagnosis or missed diagnosis, and having high practical value and promotion prospects.
[0088] 2) Rapid screening: This method can quickly screen the phenotypic characteristics of patients, provide early prediction of rare diseases, and provide auxiliary decision-making support for clinicians.
[0089] 3) Automated diagnosis support: Using machine learning algorithms, potential disease information can be automatically extracted from a large amount of phenotypic data, reducing manual intervention and improving the diagnosis efficiency.
[0090] 4) Wide applicability: This method is not only applicable to the prediction of rare diseases, but also can be applied to the prediction and diagnosis of other genetic or complex diseases.
[0091] In the above embodiments, although the steps are numbered S1, S2, etc., they are only specific embodiments given by the present invention. Those skilled in the art can adjust the execution order of S1, S2, etc. according to the actual situation, and this is also within the protection scope of the present invention. It can be understood that in some embodiments, it may include some or all of the above embodiments.
[0092] As Figure 2 shown, a model training system 200 based on multi-modal data fusion according to an embodiment of the present invention includes a data preprocessing module 201, a data extraction and fusion module 202, and a model training module 203;
[0093] The data preprocessing module 201 is used for: preprocessing the collected multi-modal data to obtain the preprocessed multi-modal data, where the multi-modal data includes clinical symptoms, physical sign data, and treatment diagnosis results, or the multi-modal data includes industrial equipment operation failure data and industrial equipment failure cause data;
[0094] The data extraction and fusion module 202 is used for: obtaining the key data in each modal data of the preprocessed multi-modal data and performing data fusion to obtain the fused data;
[0095] The model training module 203 is used for: training a preset deep learning model using the fused data to obtain a trained deep learning model.
[0096] Optionally, in the above technical solution, the data preprocessing module 201 is specifically used for:
[0097] When the multi-modal data includes clinical symptoms, physical sign data, and treatment diagnosis results, performing data cleaning, standardization processing based on the Human Phenotype Ontology library, and missing value processing on the multi-modal data to obtain the preprocessed multi-modal data.
[0098] Optionally, in the above technical solution, the data preprocessing module 201 is specifically used for:
[0099] When the multi-modal data includes industrial equipment operation failure data and industrial equipment failure cause data, performing data cleaning and missing value processing on the multi-modal data to obtain the preprocessed multi-modal data.
[0100] Optionally, in the above technical solution, the data extraction and fusion module 202 is further specifically used for:
[0101] Performing at least two dimensionality reduction operations continuously on each modal data of the preprocessed multi-modal data to obtain the key data in each modal data of the preprocessed multi-modal data.
[0102] It should be noted that the beneficial effects of the model training system 200 based on multi-modal data fusion provided in the above embodiments are the same as those of the above model training method based on multi-modal data fusion, and will not be elaborated here. In addition, when the system provided in the above embodiments implements its functions, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be elaborated here.
[0103] Among them, the model training system based on multi-modal data fusion of the present invention can be a computer program (including program code) running on a computer device. For example, the model training system based on multi-modal data fusion of the present invention is an application software and can be used to execute the corresponding steps in the model training method based on multi-modal data fusion of the present invention.
[0104] In some embodiments, the model training system based on multi-modal data fusion of the present invention can be implemented in a combination of software and hardware. As an example, the model training system based on multi-modal data fusion of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the model training method based on multi-modal data fusion of the present invention. For example, the processor in the form of a hardware decoding processor can adopt one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.
[0105] Among them, the modules involved in the embodiments of the present invention can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the module itself in some cases.
[0106] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any one of the above-mentioned model training methods based on multi-modal data fusion. That is to say, an electronic device according to an embodiment of the present invention may include, but is not limited to: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the model training method based on multi-modal data fusion shown in any one of the embodiments of the present invention by calling the computer program.
[0107] In an alternative embodiment, an electronic device is provided, as Figure 3 shown, Figure 3The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between this electronic device and other electronic devices, such as data transmission and / or data reception, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not limit the embodiments of the present invention.
[0108] The processor 4001 can be a CPU (Central Processing Unit, central processing unit), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in connection with the disclosure of the present invention. The processor 4001 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0109] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 can be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard architecture) bus, etc. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 3 only a thick line is used to represent the bus 4002 in the figure, but it does not mean that there is only one bus or one type of bus.
[0110] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0111] The memory 4003 is used to store the application program code (computer program) for implementing the solution of the present invention and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the application program code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.
[0112] Among them, the electronic device can also be a terminal device, and the terminal device can be any device that can install applications, including at least one of a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a smart TV, and a smart vehicle device.
[0113] It should be noted that Figure 3 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0114] A computer-readable storage medium according to an embodiment of the present invention has a computer program stored thereon, and when the computer program is executed by a processor, it implements any one of the above model training methods based on multi-modal data fusion.
[0115] Optionally, the computer-readable storage medium can be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0116] In an exemplary embodiment, a computer program product or a computer program is further provided. The computer program product or the computer program includes computer instructions that are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the electronic device to execute any one of the above model training methods based on multimodal data fusion.
[0117] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0118] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0119] The computer-readable storage medium provided by the embodiments of the present invention may be, but is not limited to, a system, apparatus, or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EEPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0120] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.
[0121] The above description is only a preferred embodiment of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present invention.
[0122] It should be noted that the terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, and represent a limitation on a specific order or sequence. Under appropriate circumstances, the order of use of similar objects can be interchanged so that the embodiments of this application described here can be implemented in an order other than the illustrated or described order.
[0123] Those skilled in the art know that the present invention can be implemented as a system, method, or computer program product. Therefore, the present invention can be specifically implemented in the following forms: it can be completely hardware, can be completely software (including firmware, resident software, microcode, etc.), or can also be in the form of a combination of hardware and software, which is generally referred to as "circuit", "module", or "system" in this article. In addition, in some embodiments, the present invention can also be implemented in the form of a computer program product in one or more computer-readable media, and the computer-readable media contains computer-readable program code.
[0124] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A model training method based on multi-modal data fusion, characterized in that, Including: Preprocessing the collected multimodal data to obtain preprocessed multimodal data, where the multimodal data includes clinical symptoms, physical sign data, and treatment diagnosis results, or the multimodal data includes industrial equipment operation fault data and industrial equipment fault cause data; Obtaining key data in each modal data of the preprocessed multimodal data and performing data fusion to obtain fused data; Using the fused data to train a preset deep learning model to obtain a trained deep learning model.
2. The model training method based on multi-modal data fusion according to claim 1, wherein, Preprocessing the multimodal data includes: When the multimodal data includes clinical symptoms, physical sign data, and treatment diagnosis results, performing data cleaning, standardization processing based on the Human Phenotype Ontology library, and missing value processing on the multimodal data to obtain preprocessed multimodal data.
3. A model training method based on multimodal data fusion according to claim 1, characterized in that, Preprocessing the multimodal data includes: When the multimodal data includes industrial equipment operation fault data and industrial equipment fault cause data, performing data cleaning and missing value processing on the multimodal data to obtain preprocessed multimodal data.
4. A model training method based on multi-modal data fusion according to any one of claims 1 to 3, characterized in that Obtaining key data in each modal data of the preprocessed multimodal data includes: Continuously performing at least two dimensionality reduction operations on each modal data of the preprocessed multimodal data to obtain key data in each modal data of the preprocessed multimodal data.
5. A model training system based on multi-modal data fusion, characterized in that, Including a data preprocessing module, a data extraction and fusion module, and a model training module; The data preprocessing module is used to: preprocess the collected multimodal data to obtain preprocessed multimodal data, where the multimodal data includes clinical symptoms, physical sign data, and treatment diagnosis results, or the multimodal data includes industrial equipment operation fault data and industrial equipment fault cause data; The data extraction and fusion module is used to: obtain key data in each modal data of the preprocessed multimodal data and perform data fusion to obtain fused data; The model training module is used to: use the fused data to train a preset deep learning model to obtain a trained deep learning model.
6. The model training system based on multi-modal data fusion according to claim 5, wherein, The data preprocessing module is specifically used for: When the multimodal data includes clinical symptoms, physical sign data, and treatment diagnosis results, performing data cleaning, standardization processing based on the Human Phenotype Ontology library, and missing value processing on the multimodal data to obtain preprocessed multimodal data.
7. An model training system based on multimodal data fusion according to claim 5, characterized in that, The data preprocessing module is specifically used for: When the multimodal data includes industrial equipment operation fault data and industrial equipment fault cause data, performing data cleaning and missing value processing on the multimodal data to obtain preprocessed multimodal data.
8. A model training system based on multi-modal data fusion according to claim 5 or 6, characterized in that The data extraction and fusion module is also specifically used for: Continuously performing at least two dimensionality reduction operations on each modal data of the preprocessed multimodal data to obtain key data in each modal data of the preprocessed multimodal data.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements a model training method based on multi-modal data fusion according to any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements a model training method based on multi-modal data fusion according to any one of claims 1 to 4.
Citation Information
Patent Citations
Aero-engine fault diagnosis method and system based on multi-modal deep learning
CN116842423A
Forest resource analysis method and system based on forestry ecological big data
CN118350554A
Large model data fusion method and system based on multi-source heterogeneous data
CN118520418A
Knowledge fusion method and system for multi-source heterogeneous multi-modal data
CN118690838A
Multi-modal data feature processing method and system based on deep learning, and medium
CN119418142A