Multi-level classification method and system based on multi-modal data fusion
By introducing multi-level classification technology and self-attention mechanism fusion characteristics into the multi-modal data classification method, the problem that the existing technology is difficult to combine multi-modal data and hierarchical structure is solved, and a more accurate and refined classification effect in complex classification tasks is achieved.
Patent Information
- Application Number
- CN202510588804.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multimodal data classification methods are difficult to effectively combine multimodal data with hierarchical structures, resulting in poor performance in complex classification tasks and unable to meet the needs of fine classification in practical applications.
A multi-level classification method based on multi-modal data fusion is proposed. By obtaining image, text and tabular data for pre-processing, extracting features and fusion of self-attention mechanisms, building a hierarchical structure tree, and establishing a multi-level classification model for hierarchical classification training.
It has achieved a more accurate and refined classification scheme in multi-level complex classification scenarios, and can gradually refine the results from coarse-grained classification to meet the needs of fine classification in actual applications.
Smart Images

Figure CN120217110A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data analysis and intelligent information processing, and particularly relates to a multi-level classification method and system based on multi-modal data fusion. Background Art
[0002] Data classification in complex scenarios is one of the core challenges in many application fields. For example, in business, industry, and medical research, the complexity of massive data and the diversity of application scenarios usually make traditional single-modal and single-level classification methods inapplicable. Different data modalities often contain independent feature information, which cannot fully demonstrate its full potential in specific problems when analyzed separately. Traditional multi-modal classification models usually ignore the hierarchical relationships between classes and only regard all classes as equal and discrete individuals, making it difficult to effectively utilize the complementary information between modalities and the information of the hierarchical structure. Moreover, the contributions of different modalities to each level of classification are also different, and some modalities can provide key information for the classification of specific levels. Therefore, the research on combining multi-modal data fusion and multi-level classification technology has gradually attracted attention. By jointly modeling and fusing multi-modal features, the associated information between modalities can be fully exploited in the classification task, and the multi-level classification technology can gradually optimize the classification results and improve the classification accuracy from coarse-grained to fine-grained. The combination of this multi-modal fusion and multi-level classification provides a new idea for solving the data classification problem in complex scenarios, such as achieving fine classification of goods in the e-commerce field, completing accurate typing of diseases in the medical field, or optimizing abnormal behavior detection in traffic monitoring.
[0003] Existing technical solutions such as [Publication No.: CN112685565A, Invention Title: Text Classification Method Based on Multi-Modal Information Fusion and Related Devices] propose a text classification solution based on the fusion features of image and text modalities, but the extraction of its feature modalities is single, with only two modalities of image and text, and the text classification module is simple. It fails to optimize the classification for the data hierarchical structure and can only handle simple classification scenarios at a single level. At the same time, another technical solution [Publication No.: CN117056863A, Invention Title: A Big Data Processing Method Based on Multi-Modal Data Fusion] only proposes a fusion model for multi-modal data. Although it uses multi-modal data, it fails to propose a technical solution for complex hierarchical classification scenarios based on multi-modal data.
[0004] Most existing multimodal data classification methods are designed for specific single classification tasks. The models usually classify the data once and divide the results into coarse-grained categories, but fail to combine multimodal data with hierarchical structures, resulting in the classification system being difficult to effectively adapt to hierarchical requirements. The scalability and applicability of the models are poor and cannot be applied to complex classification tasks. For example, existing technologies may perform well in single-level scenarios, but they are clearly insufficient in actual scenarios that require hierarchical classification (such as from food classification to specific brands and ingredient classification). Summary of the invention
[0005] The purpose of the present invention is to address the shortcomings of single classification of multimodal data and propose a multi-level classification method and system based on multimodal data fusion, aiming to more comprehensively handle complex classification tasks in multimodal data scenarios through a hierarchical classification model.
[0006] The objective of the present invention is achieved through the following technical solution: A multi-level classification method based on multimodal data fusion, comprising the following steps:
[0007] Acquire image, text and table data to be classified, preprocess the image, text and table data to obtain preprocessed image data, text data and table features;
[0008] The preprocessed image data is input into the pretrained deep residual network model for feature extraction and structured vector representation to obtain vectorized image features; the preprocessed text data is input into the pretrained language model for feature extraction and structured vector representation to obtain vectorized text features;
[0009] Based on the self-attention mechanism, the image features, text features and table features are fused to obtain fused features;
[0010] Constructing a hierarchical structure tree, constructing a multi-level classification model based on the hierarchical tree, and inputting the fusion features into the multi-level classification model for hierarchical classification training;
[0011] The test set part of the multimodal fusion data is input into the trained multi-level classification model for test evaluation and output of the classification results.
[0012] Furthermore, the preprocessing of the image, text and table data includes: image enhancement, resizing and normalization of image data, removal of stop words, word segmentation and truncation and filling of text data, and processing of missing values, outliers and standardization of table data.
[0013] Furthermore, the preprocessed image data is input into a pre-trained deep residual network model for feature extraction and structured vector representation to obtain vectorized image features, including:
[0014] Select the ResNet50 model, input the preprocessed image data into the ResNet50 model, and output a 1×2048 feature vector through the global average pooling layer of the ResNet50 model. The feature vector contains the high-dimensional feature information of the image to represent the visual content of the image.
[0015] Further, the input of the preprocessed text data into the pre-trained language model for feature extraction and structured vector representation to obtain the vectorized text features includes:
[0016] Select the BERT model, input the preprocessed text data into the BERT model, obtain the token of each lexical unit through the BERT model, map each token to a vector representation in a high-dimensional space, and use the pooling layer in the BERT model to average and summarize all word embeddings to obtain a 768-dimensional numerical feature vector representation of the entire text.
[0017] Further, the feature fusion of the image features, text features, and table features based on the self-attention mechanism to obtain the fusion features includes:
[0018] Perform feature concatenation on the image features, text features, and table features to obtain a concatenated feature vector. Based on the self-attention mechanism, automatically assign different weights to different features, and weighted according to the weights to obtain a preliminary fusion feature, and input it into a fully connected layer for dimension compression to obtain the final fusion feature.
[0019] Further, the construction of a multi-level classification model based on the hierarchical tree and the input of the fusion features into the multi-level classification model for hierarchical classification training includes:
[0020] The multi-level classification model includes constructing a local classifier for each parent node, and setting the local base classifier for each parent node as logistic regression; representing the hierarchical labels as a true label matrix, with the first column being the parent node label and the second column being its corresponding child node label
[0021] Use the StratifiedShuffleSplit stratified sampler to split the dataset to ensure that the training and test sample ratios of different categories and levels are balanced. Input the training set into the logistic regression classifier corresponding to each parent node for top-down training, use the corresponding column in the label matrix as the training target, and perform hyperparameter optimization on the regularization strength of the logistic regression classifier of each parent node based on grid search to obtain the optimal local classifier for each parent node training.
[0022] Further, inputting the test set part of the multi-modal fusion data into the trained multi-level classification model for test evaluation and outputting the classification result includes:
[0023] Input the split test set into the trained multi-level classification model. Based on each local optimal logistic regression classifier, output the label predicted by this node to obtain a predicted label matrix, and output the node predicted label as the classification result;
[0024] Use hierarchical classification metrics such as hierarchical accuracy, hierarchical recall rate, and hierarchical F1-score to evaluate the classification result and output the evaluation result.
[0025] The present invention also provides a multi-level classification system based on multi-modal data fusion, including:
[0026] A multi-modal data preprocessing module for cleaning image, text, and table data;
[0027] A multi-modal feature extraction and vectorization module for respectively extracting features and performing structured vector representation on images and texts;
[0028] A self-attention mechanism multi-modal feature fusion module for fusing image features, text features, and table features;
[0029] A multi-level classification framework construction and training module for constructing a multi-level classification framework and performing hierarchical classification training;
[0030] A model test and evaluation module for testing and evaluating the multi-level classification result and outputting and displaying it.
[0031] The present invention also provides an electronic device, including a memory and a processor, and the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the multi-level classification method based on multi-modal data fusion as described above.
[0032] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the multi-level classification method based on multi-modal data fusion as described above.
[0033] The beneficial effect of the present invention is that: aiming at the limitation that the current multi-modal fusion data classification is only applicable to simple classification tasks at a single level, the present invention can, by combining the feature fusion of multi-modal data and a hierarchical classification model with gradual optimization, on the basis of fully mining and fusing the information of each modal data, give a more accurate and refined classification scheme for multi-level complex classification scenarios, gradually refine the classification result on the basis of coarse-grained classification, so as to meet the demand for fine classification in practical applications. Brief Description of the Drawings
[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0035] Figure 1 It is the system flowchart of the present invention;
[0036] Figure 2 It is the schematic diagram of image data feature extraction and vectorization in the embodiment of the present invention;
[0037] Figure 3 It is the schematic diagram of text data feature extraction and vectorization in the embodiment of the present invention;
[0038] Figure 4 It is the schematic diagram of self-attention mechanism multi-modal feature fusion in the embodiment of the present invention;
[0039] Figure 5 It is the schematic diagram of the multi-level structure tree of the disease to be investigated for fever in the embodiment of the present invention. Detailed Description of the Embodiments
[0040] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0041] As Figure 1 shown, the embodiment of the present invention provides a multi-level classification system based on multi-modal data fusion, including:
[0042] A multi-modal data preprocessing module for cleaning image, text, and table data.
[0043] A multi-modal feature extraction and vectorization module for respectively extracting features and performing structured vector representation on images, texts, and tables.
[0044] A self-attention mechanism multi-modal feature fusion module for fusing image features, text features, and table features.
[0045] A multi-level classification framework construction and training module for constructing a multi-level classification framework and performing hierarchical classification training and testing.
[0046] The model testing and evaluation module is used to test, evaluate and output the multi-level classification results for display.
[0047] Based on the above system, an embodiment of the present invention further provides a multi-level classification method based on multi-modal data fusion, including the following steps:
[0048] Step 1: Preprocess the image, text and tabular data based on the multi-modal data preprocessing module.
[0049] Taking the diagnosis of undiagnosed fever with rash in medicine as a specific implementation case, the preprocessing process specifically includes image enhancement, size adjustment and normalization of image data, removal of stop words, word segmentation and truncation padding of text data, and handling of missing values, outliers and standardization of tabular data.
[0050] The tabular data in the case of undiagnosed fever with rash includes a total of 144 indicators, including the patient's clinical basic characteristics such as gender and age, physiological characteristics such as respiration, heart rate, pulse, systolic blood pressure, diastolic blood pressure and body temperature, past medical history such as hypertension, diabetes, hepatitis B, heart disease, kidney disease and malignant tumor, 22 items of blood routine examination, 7 items of urine routine examination, 5 items of electrolyte examination, 16 items of liver function examination, 6 items of thyroid function examination, 6 items of coagulation function examination, 7 items of kidney function examination, 6 items of blood glucose and lipid examination, 6 items of APS-related antibody examination, 4 items of anti-human globulin test, 5 items of immunoglobulin examination, 3 items of urine protein quantitative examination, 7 items of cytokine examination, 5 items of myocardial enzyme spectrum examination, 7 items of tumor marker examination, 14 items of lymphocyte subset examination, 5 items of cerebrospinal fluid examination, 3 items of cerebrospinal fluid glucose, chlorine and protein determination, 5 items of serum protein electrophoresis examination, and 5 items of culture examination.
[0051] Among them, the preprocessing of the tabular data includes data cleaning such as handling missing values, outliers and standardization. Specifically, it includes removing outliers of all variables, eliminating variables with more than half of the missing values, filling missing values of categorical variables with the mode, filling missing values of continuous variables with the mean, encoding categorical variables using LabelEncoder (label encoding), and standardizing all variables using z-score.
[0052] The text data in the case of fever with rash to be investigated is the electronic medical record of the patient's rash text. Among them, the preprocessing of the text data includes removing stop words, word segmentation, and truncation from the electronic medical record of the rash text. The step of removing stop words includes removing common words that have no significant contribution to semantics in the process of text classification or modeling, such as "de", "le", "shi", etc., when processing the rash text data, so as to reduce the impact of noise on subsequent processing. The word segmentation step includes using the BERT (Bidirectional Encoder Representations from Transformers) tokenizer to segment the rash text of each patient, and splitting the original text into multiple lexical units (tokens) to convert the text into a form that can be processed by the BERT model. The truncation step includes performing padding or truncation operations on the rash text to ensure that all input texts have the same length. For texts with insufficient length, padding operations are used to supplement special padding symbols at the end. For overly long texts, truncation operations are used to maintain the maximum length of the text, so as to ensure that the model can accept fixed-length inputs during processing.
[0053] The image data in the case of fever with rash to be investigated is the medical CT image of the patient. Among them, the preprocessing of the image data includes image resampling, window width and window level adjustment, data augmentation, and size adjustment operations on the CT image. The resampling step includes using the B-spline interpolation algorithm to resample all image data to a unified isotropic voxel size of 1×1×1 cm³ to ensure consistent slice thickness and in-plane resolution of the image. The window width and window level adjustment step includes setting the image window level to 0 Hu and the window width to 400 Hu according to the gray values of the region of interest in the image, and at the same time using voxel value discretization to reduce image noise. The data augmentation step includes performing data augmentation operations such as rotating (90 degrees, 180 degrees), mirroring (left and right, up and down) on the image to increase the diversity of samples. The image size adjustment step includes selecting the middle frame of the CT image and taking one layer above and one layer below, a total of three grayscale images as the input of the RGB three channels, and using bilinear interpolation to scale all image data to a size of 224×224 to meet the input size requirements of the model.
[0054] Step 2: Vectorize the image, text, and table respectively based on the multi-modal feature extraction and vectorization module.
[0055] In the case of the implementation of fever with rash to be investigated in the present invention, it includes feature vectorization of the preprocessed image data and feature vectorization of the preprocessed text data. Since the table data is already a structured vector, no further processing is required.
[0056] In the extraction and vectorization of the image data features, the preprocessed image data (a three-channel RGB image with a size of 224×224) is taken as the input and fed into the ResNet50 model. For CT images, the three extracted grayscale images are used as three-channel data. The ResNet50 is a deep residual network model with 50 layers of depth and uses a residual learning strategy to avoid the problem of gradient vanishing. The network includes an initial convolutional layer, a max pooling layer, 4 residual block groups, and a global average pooling layer, and the model has been pre-trained on the ImageNet dataset and has powerful image feature extraction capabilities. After the input image undergoes the forward propagation process, ResNet50 generates a global feature representation. Specifically, a 2048×1×1 feature vector (i.e., a 2048-dimensional feature vector representation) is output through the global average pooling layer of the network. This feature vector contains the high-dimensional feature information of the image and represents the main visual content of the image, such as Figure 2 as shown
[0057] In the extraction and vectorization of the text data features, the preprocessed text data is input into the pre-trained BERT model. The BERT model is a pre-trained language model based on the Transformer (self-attention transformer) architecture, which models context information in the text through a bidirectional encoder. The tokenized rash electronic medical record diagnosis text is input into the BERT model to obtain the embedding representation (token) of each lexical unit. Each token will be mapped to a vector representation in a high-dimensional space. The pooling layer in the BERT model (using the output of the CLS token) is used to average and summarize all the word embeddings to obtain a 768-dimensional numerical feature vector representation of the entire text. The output 768-dimensional numerical feature vector reflects the semantic information learned from the rash electronic medical record diagnosis text of each patient in the BERT model, such as Figure 3 as shown
[0058] Step 3: Based on the self-attention mechanism multi-modal feature fusion module, perform self-attention mechanism feature fusion on the image, text, and table features.
[0059] In the implementation case of the present invention for patients with fever of unknown origin accompanied by rash, it includes fusing the vectorized image features, text features, and preprocessed table features to obtain multi-modal fusion features.
[0060] The self-attention mechanism feature fusion method mainly calculates the correlation between input features, enhances the representation of effective features, and suppresses the influence of irrelevant features, thereby avoiding the unique attributes and correlations between different modalities that may be ignored by simply concatenating multi-modal features. Specifically, the self-attention mechanism is a special attention mechanism in which each part of the input is calculated with other parts of the input to learn the internal correlation. The self-attention mechanism captures the relationships between features and assigns more weights to effective features. In the present invention, the fusion process of image, text, and table features is as Figure 4 shown, specifically including feature concatenation, calculating the weighted preliminary fusion features using the self-attention mechanism, and obtaining the final fusion features after dimension compression of the preliminary fusion features using a fully connected layer.
[0061] As Figure 4 shown, in the feature concatenation, the previously obtained image features, text features, and table features are concatenated. The 2048-dimensional feature vector obtained by processing the image through the ResNet50 model is represented as , where each represents a feature of the image; the 768-dimensional feature vector obtained by the text feature through the BERT model is represented as , where each represents a feature of the text; the 144-dimensional feature vector obtained by preprocessing the table feature is represented as , where each represents a feature of the table. The image features, text features, and table features are concatenated to obtain a -dimensional feature vector :
[0062] .
[0063] As Figure 4 shown, in the self-attention mechanism, the concatenated feature vector uses the self-attention mechanism to automatically assign different weights to different features. For each feature, a set of learned weight matrices are used to calculate (Query), (Key), and (Value) vectors respectively, where is the input vector representing the actual information of the feature; is the product of the input and the weight , representing the query vector of the feature; is the product of the input and the weight The product represents the relationship with other features. In particular, in the self-attention mechanism, Query, Key, and Value are the same and all come from the same input feature vector :
[0064] .
[0065] Use and to calculate the self-attention weights , and its formula is as follows:
[0066] .
[0067] Among them, is the transpose of Key, is the dimension of the Key vector. In this step, first calculate the dot product of Query and Key, then divide by the scaling factor , and finally normalize the result to a probability distribution through the Softmax function to obtain the attention weights of each part to other parts. The final output can be obtained by multiplying the Attention weights by the Value vector:
[0068] .
[0069] In the fusion feature compression, the output is input into the fully connected layer to obtain the finally fused feature :
[0070] .
[0071] Among them, and are the parameters of the fully connected layer. In the implementation case of the present invention for patients with fever to be investigated accompanied by rash, in order to obtain a highly abstract multi-modal feature representation and remove unnecessary redundant information, the output dimension of the fully connected layer is set to 512, and finally a 512-dimensional fused feature is obtained for input in subsequent multi-level classification.
[0072] Step 4, construct a hierarchical structure tree based on the multi-level classification framework construction and training module, construct a multi-level classification framework based on the hierarchical tree, and input training data for model training.
[0073] In the implementation case of the present invention for patients with fever to be investigated accompanied by rash, it includes constructing a hierarchical structure tree for the causes of fever to be investigated and a multi-level classification model, and inputting fused data for training.
[0074] The hierarchical structure tree of the diseases to be investigated for fever includes three levels, as Figure 5 shown. The first level is the root node; the second level is divided into four parent nodes, including infectious diseases, non-infectious inflammatory diseases (NIID), neoplastic diseases, and other diseases; the third level is further divided into multiple child nodes. Among them, infectious diseases are divided into four child nodes: bacterial infection, viral infection, fungal infection, and parasitic infection. NIID is divided into two child nodes: autoinflammatory diseases and autoimmune diseases. Neoplastic diseases are divided into three child nodes: lymphoma, hematological diseases, and malignant tumors. Others include one child node of drug fever.
[0075] The multi-level classification model constructed based on the hierarchical tree includes constructing a local classifier for each parent node (as Figure 5 ). For each parent node , in the present invention, the local base classifier is set as logistic regression, denoted as , where is the input feature vector, that is, the obtained fusion feature, is the activation function of logistic regression, and are the weights and biases of logistic regression; the hierarchical labels are represented as the true label matrix, with the first column being the parent node label and the second column being the corresponding child node label, denoted as , where is the th parent node label, is the th child node label corresponding to the parent node .
[0076] In the model classification training, the StratifiedShuffleSplit stratified sampler is used to split the dataset to ensure that the proportions of training and test samples for different categories and levels are balanced. Then, the training set is input into the logistic regression classifier corresponding to each parent node for training from top to bottom. The corresponding columns in the label matrix are used as the training targets. Based on grid search, hyperparameter optimization is performed on the regularization strength C of the logistic regression classifier for each parent node to obtain the optimal local classifier for each parent node training.
[0077] Step 5: Based on the model test and evaluation module, the test set part of the multi-modal fusion data is input into the trained multi-level classification framework for test evaluation and the classification result is output.
[0078] In the implementation case of the patient with fever of unknown origin accompanied by rash in the present invention, it includes inputting the split test set into a multi-level classification model for testing and outputting the multi-level labels of each sample.
[0079] The model testing includes inputting the test set split by the StratifiedShuffleSplit stratified sampler into the trained multi-level classification model, and based on each local optimal logistic regression classifier , outputting the label predicted by this node, and finally obtaining the predicted label matrix , where represents the predicted th parent node label, represents the predicted label of the th child node corresponding to the parent node , and finally taking the node predicted label as the classification result to output.
[0080] The model evaluation uses the hierarchical classification metrics hierarchical accuracy, hierarchical recall, and hierarchical F1 score (hF) to evaluate the prediction results of the multi-level disease etiology diagnosis model. The specific calculation methods of each evaluation metric are as follows, where is the set of the predicted sub-categories and their parent categories of the test samples, is the set of the true sub-categories and their parent categories of the test samples:
[0081] (1) Hierarchical accuracy (hP):
[0082] .
[0083] Among them, hP represents the proportion of the number of nodes correctly predicted by the model to the total number of nodes. For each parent node and its corresponding child node, if both are correctly predicted, they are regarded as correct. This metric can evaluate the accuracy of the multi-level disease etiology diagnosis model at each level.
[0084] (2) Hierarchical recall (hR):
[0085] .
[0086] Among them, hR represents the proportion of the number of positive examples correctly predicted by the model to the total number of positive examples. For each parent node and its corresponding child node, if both are correctly predicted, they are regarded as positive examples. This metric can evaluate the recall rate of the multi-level disease etiology diagnosis model at each level.
[0087] (3) Hierarchical F1 score (hF):
[0088] .
[0089] Among them, hF represents the harmonic mean of the hierarchical accuracy and hierarchical recall rate of the model. This metric comprehensively considers the accuracy and recall rate of the model and can comprehensively evaluate the performance of the multi-level disease etiology diagnosis model. Finally, the above metrics are used to evaluate the classification effect of the multi-level classification model.
[0090] The present invention deeply fuses multi-modal data (such as images, texts, and tables, etc.), uses the self-attention mechanism to fuse and model the relationships between different modal data, improves the expression ability of features, and combines a multi-level structure to gradually optimize the classification effect, realizing a classification process from coarse to fine. The multi-level classification framework proposed by the present invention can adapt to complex classification tasks in different fields and diverse application scenarios. Whether in medical diagnosis, product recommendation, or other complex environments, the present invention can provide accurate and efficient classification results.
[0091] An embodiment of the present invention also provides an electronic device, including a memory and a processor, the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the multi-level classification method based on multi-modal data fusion described above.
[0092] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the multi-level classification method based on multi-modal data fusion described above.
[0093] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or will be output.
[0094] Those skilled in the art will readily think of other embodiments of the present application after considering the specification and practicing the content disclosed herein. The present application aims to cover any variations, uses, or adaptive changes of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary.
[0095] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A multi-level classification method based on multimodal data fusion, characterized in that: The steps include: Acquire image, text and table data to be classified, preprocess the image, text and table data to obtain preprocessed image data, text data and table features; The preprocessed image data is input into the pre-trained deep residual network model for feature extraction and structured vector representation to obtain vectorized image features; The preprocessed text data is input into the pre-trained language model for feature extraction and structured vector representation to obtain vectorized text features; Based on the self-attention mechanism, the image features, text features and table features are fused to obtain fused features; Constructing a hierarchical structure tree, constructing a multi-level classification model based on the hierarchical tree, and inputting the fusion features into the multi-level classification model for hierarchical classification training; The test set part of the multimodal fusion data is input into the trained multi-level classification model for test evaluation and output of the classification results.
2. The multi-level classification method based on multimodal data fusion according to claim 1 is characterized in that: The preprocessing of the image, text and table data includes: image enhancement, size adjustment and normalization for image data, removal of stop words, word segmentation and truncation and filling for text data, and processing of missing values, outliers and standardization for table data.
3. The multi-level classification method based on multimodal data fusion according to claim 1 is characterized in that: The preprocessed image data is input into the pre-trained deep residual network model for feature extraction and structured vector representation to obtain vectorized image features including: A ResNet50 model is selected, and the preprocessed image data is input into the ResNet50 model. A 1×2048 feature vector is output through the global average pooling layer of the ResNet50 model. The feature vector contains high-dimensional feature information of the image to represent the visual content of the image.
4. The multi-level classification method based on multimodal data fusion according to claim 1 is characterized in that: The preprocessed text data is input into the pre-trained language model to perform feature extraction and structured vector representation, and the vectorized text features obtained include: The BERT model is selected, and the preprocessed text data is input into the BERT model. The token of each vocabulary unit is obtained through the BERT model, and each token is mapped to a vector representation in a high-dimensional space. The pooling layer in the BERT model is used to average all word embeddings to obtain a 768-dimensional numerical feature vector representation of the entire text.
5. The multi-level classification method based on multimodal data fusion according to claim 1, characterized in that: Based on the self-attention mechanism, the image features, text features and table features are fused to obtain the fused features including: The image features, text features and table features are concatenated to obtain a concatenated feature vector. Different weights are automatically assigned to different features based on the self-attention mechanism. According to the weights, preliminary fused features are weighted and input into the fully connected layer for dimensional compression to obtain the final fused features.
6. The multi-level classification method based on multimodal data fusion according to claim 1 is characterized in that: The constructing a multi-level classification model based on the hierarchical tree and inputting the fusion features into the multi-level classification model for hierarchical classification training comprises: The multi-level classification model includes constructing a local classifier for each parent node, setting the local base classifier for each parent node to be a logistic regression; representing the level label as a true label matrix, the first column is the parent node label, and the second column is the corresponding child node label; The StratifiedShuffleSplit stratified sampler is used to split the dataset to ensure a balanced ratio of training and test samples in different categories and levels. The training set is input into the logistic regression classifier corresponding to each parent node for top-down training. The corresponding columns in the label matrix are used as training targets. The regularization strength of the logistic regression classifier of each parent node is hyperparameter optimized based on grid search to obtain the optimal local classifier for each parent node training.
7. The multi-level classification method based on multimodal data fusion according to claim 6 is characterized in that: The step of inputting the test set portion of the multimodal fusion data into the trained multi-level classification model, performing test evaluation and outputting classification results includes: The split test set is input into the trained multi-level classification model. Based on each local optimal logistic regression classifier, the predicted label of the node is output to obtain the predicted label matrix, and the node predicted label is output as the classification result. The classification results are evaluated using hierarchical classification indicators hierarchical accuracy, hierarchical recall and hierarchical F1 score, and the evaluation results are output.
8. A multi-level classification system based on multimodal data fusion, characterized in that: include: Multimodal data preprocessing module, used for cleaning image, text and table data; Multimodal feature extraction and vectorization module, used to extract features and represent structured vectors for images and texts respectively; Self-attention mechanism multimodal feature fusion module, used to fuse image features, text features and table features; A multi-level classification framework construction and training module is used to construct a multi-level classification framework and conduct hierarchical classification training; The model testing and evaluation module is used to test and evaluate the multi-level classification results and output them for display.
9. An electronic device, comprising a memory and a processor, characterized in that: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement a multi-level classification method based on multimodal data fusion as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a multi-level classification method based on multimodal data fusion as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Cardiovascular disease risk prediction method based on multi-modal fusion
CN117153393A
Typing, grading and staging evaluation system and evaluation method for pancreaticobiliary tumor
CN117218419A
Brain disease classification method and device
CN118674972A
Ovarian cancer preoperative accurate prediction system and method based on multi-modal radiomics characteristics
CN119480090A