Data classification and grading processing method and device, medium and program product
By combining the multimodal large model with the vertical field small model, the problems of data classification and grading accuracy and privacy protection of the multimodal large model in the vertical field are solved, and the accuracy and reliability of data classification and grading are improved while protecting data privacy.
Patent Information
- Application Number
- CN202510827108.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-19
AI Technical Summary
Existing large multimodal models lack understanding of domain-specific semantics and context in vertical fields, resulting in low accuracy in data classification and grading, and difficulties in protecting data privacy.
By combining a large multimodal model with a small vertical field model, multi-source data features are extracted through a pre-trained multimodal feature extraction model, and processed using a pre-trained classification and grading model. Data classification and grading are achieved by combining voting method, weighting coefficient and decision fusion technology.
On the premise of protecting data privacy, the accuracy and reliability of data classification and grading are improved, and the accuracy and privacy protection issues of data classification and grading in vertical fields are solved.
Smart Images

Figure CN120671086A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data classification technology, and more specifically, to a data classification and grading processing method, device, medium, and program product. Background Art
[0002] The multimodal big model is a machine learning technology based on deep learning. It's a type of big model. Its core concept is to fuse data from different media (such as text, images, audio, and video) and learn the connections between these modalities to achieve more intelligent information processing. In a multimodal big model, data from different modalities is preprocessed and then fed into a deep neural network. After multiple layers of feature extraction and fusion, the resulting output is the corresponding result.
[0003] Large models, with their powerful semantic understanding, expressive power, and generalization capabilities, provide effective solutions for data classification and grading. However, these models are often pre-trained on general datasets and lack understanding of domain-specific knowledge. In some verticals (such as healthcare, law, and finance), these general-purpose large models may not accurately capture domain-specific semantics and context, resulting in low data classification and grading accuracy. Furthermore, because data in specific industry or enterprise scenarios often contains sensitive information such as personal privacy and commercial secrets, it is often difficult to train and process using open source large models.
[0004] In summary, there is an urgent need for a solution that can improve the accuracy of data classification and grading while protecting data privacy. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to provide a data classification and grading processing method, device, medium and program product to improve the accuracy of data classification and grading while protecting data privacy.
[0006] In a first aspect, an embodiment of the present application provides a data classification and grading processing method, comprising: Collecting multi-source data corresponding to the task to be classified and graded, and extracting multimodal features of the multi-source data using a pre-trained multimodal feature extraction model; The multimodal features are processed using a pre-trained classification and grading model to obtain data classification and grading results output by the classification and grading model; wherein the classification and grading model is composed of a fusion of a pre-trained multimodal large model and a pre-trained vertical domain small model; A final classification and grading decision strategy is determined based on the classification and grading results.
[0007] In the embodiments of the present application, by combining large models and small models to perform data classification and grading tasks in specific fields, not only can the versatility and powerful learning ability of the large model be fully utilized, but the advantages of efficiency, convenience and privacy protection of the small model can also be brought into play, thereby improving the accuracy of data classification and grading while protecting data privacy.
[0008] In some embodiments, the processing of the multimodal features using a pre-trained classification and grading model to obtain data classification and grading results output by the classification and grading model includes: Inputting the multimodal features into the multimodal large model to obtain a preliminary classification and grading result output by the multimodal large model; The preliminary classification and grading results are output to the vertical domain small model to obtain the data classification and grading results output by the vertical domain small model.
[0009] In an embodiment of the present application, a model integration method using a stacking method is adopted to fuse the prediction results of the large model and the vertical small model, thereby improving the accuracy and reliability of data classification and grading.
[0010] In some embodiments, the processing of the multimodal features using a pre-trained classification and grading model to obtain data classification and grading results output by the classification and grading model includes: Using the multimodal large model and the vertical domain small model to perform prediction based on the multimodal features, respectively, to obtain corresponding first classification and grading results and second classification and grading results; The first classification and grading result and the second classification and grading result are integrated based on a voting method to obtain the data classification and grading result.
[0011] In an embodiment of the present application, the accuracy and reliability of data classification and grading are improved by utilizing a large multimodal model and a small vertical domain model to output prediction results respectively, and fusing the outputs of the two models based on a voting method.
[0012] In some embodiments, the processing of the multimodal features using a pre-trained classification and grading model to obtain data classification and grading results output by the classification and grading model includes: Using the large multimodal model to make predictions based on the multimodal features to obtain a first classification and grading result and a corresponding first confidence score; and simultaneously using the small vertical domain model to make predictions based on the multimodal features to obtain a second classification and grading result and a corresponding second confidence score; A weighting coefficient is determined based on the first confidence score and the second confidence score, and the first classification and grading result and the second classification and grading result are weightedly fused based on the weighting coefficient to obtain the data classification and grading result.
[0013] In an embodiment of the present application, the accuracy and reliability of data classification and grading are improved by utilizing a large multimodal model and a small vertical domain model to output prediction results and their corresponding confidence scores respectively, and fusing the outputs of the two models in combination with the confidence scores.
[0014] In some embodiments, the processing of the multimodal features using a pre-trained classification and grading model to obtain data classification and grading results output by the classification and grading model includes: Processing the multimodal features using the vertical domain small model to obtain vertical domain feature data extracted by the vertical domain small model; Fusing the multimodal features with the vertical domain feature data to obtain a fused feature vector; The fused feature vector is input into a pre-trained classification and grading converter model to obtain the data classification and grading result.
[0015] In an embodiment of the present application, vertical feature data is obtained by processing multimodal features using a small vertical model, and then the vertical feature data is fused with the multimodal features. Finally, the result is predicted based on the fused features, thereby improving the accuracy and reliability of data classification and grading.
[0016] In some embodiments, collecting multi-source data corresponding to the task to be classified and graded, and extracting multi-modal features of the multi-source data using a pre-trained multi-modal feature extraction model, includes: For the text data in the multi-source data, extract features from the text data using a preset text feature extraction model to obtain text feature data; wherein the text feature extraction model includes a bag-of-words model, a word vector model, or a pre-trained transformer model; For the image data in the multi-source data, extract features of the image data using a pre-trained convolutional neural network model to obtain image feature data; For the audio data in the multi-source data, using a preset audio analysis toolkit to perform feature extraction on the image data to obtain acoustic feature data; A feature fusion operation is performed based on the text feature data, the image feature data and the acoustic feature data to obtain multimodal features of the multi-source data; wherein the feature fusion operation includes feature vector splicing, feature alignment and feature dimensionality reduction.
[0017] In the embodiment of the present application, the accuracy of feature extraction is effectively improved by selecting corresponding models according to the characteristics of different modalities for feature extraction and splicing and fusing the feature data of multiple modalities to obtain multimodal features.
[0018] In some embodiments, the method for constructing the vertical domain small model includes: Collect structured data and unstructured data in the vertical field of the target industry, and build a domain expertise database based on the structured data and the unstructured data; Based on the domain expertise database and preset annotated data, several preset initialization models are used for training to obtain corresponding small models in vertical fields; wherein the preset initialization models include the BioBERT model for processing text data, the DenseNet model for processing image data, and the CLIP model for processing multimodal data; The routing mechanism of each vertical field mini-model is configured to activate the corresponding type of vertical field mini-model according to the modality type of the input data.
[0019] In an embodiment of the present application, the accuracy of multimodal data processing is further improved by training corresponding models according to different modalities of the input data and setting the routing mechanism of the vertical domain small model according to the modality of the input data.
[0020] In a second aspect, an embodiment of the present application provides a data classification and grading processing device, comprising: A feature extraction module is used to collect multi-source data corresponding to the classification and grading tasks, and extract multimodal features of the multi-source data using a pre-trained multimodal feature extraction model; A fusion output module is used to process the multimodal features using a pre-trained classification and grading model to obtain data classification and grading results output by the classification and grading model; wherein the classification and grading model is formed by fusing a pre-trained multimodal large model and a pre-trained vertical domain small model; A strategy determination module is used to determine a final classification and grading decision strategy based on the classification and grading results.
[0021] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor can implement the method described in any embodiment of the first aspect when executing the program.
[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any embodiment of the first aspect can be implemented.
[0023] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program, wherein when the computer program is executed by a processor, it can implement the method described in any embodiment of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0025] Figure 1 A flowchart of a data classification and grading processing method provided in an embodiment of the present application; Figure 2 A schematic diagram of the structure of a data classification and grading processing device provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0027] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0028] It should be noted that the advantage of large multimodal models is their ability to fully leverage information from diverse media data to extract richer and more comprehensive features, thereby improving model performance and generalization. Furthermore, large multimodal models can further enhance the model's semantic understanding and expression capabilities by learning relationships between different modalities.
[0029] A vertical domain small model (vertical domain small model) refers to a small AI model trained and optimized specifically for a specific domain, industry, or task. In contrast to the concept of a general-purpose large model, a vertical domain small model focuses more on the knowledge and skills of a specific domain, offering a deeper understanding and more precise performance within that domain, and possessing higher domain expertise and practicality. Compared to general-purpose large models, vertical domain small models can more accurately meet the needs of specific user groups and industries, providing users with a rich factual foundation, sufficient citations, and detailed data analysis, greatly improving work efficiency and decision-making capabilities. Vertical domain models are typically developed through transfer learning or fine-tuning, leveraging existing large models as a foundation and training them for specific domain knowledge and tasks.
[0030] It should be noted that data classification and grading is an important and difficult task, especially for industry scenarios where multi-structured data, such as semi-structured and unstructured data, coexist. It is necessary to identify multi-source data and label data according to the specific requirements of the industry and the characteristics of the scenario. It often requires a lot of manpower and time to complete the classification and grading of all domain data. At the same time, data in specific industries or enterprise scenarios usually contains personal sensitive information and has compliance risks of varying magnitudes and sensitivities. This is especially true when it comes to important data, which may contain the company's commercial secrets and sensitive information. Therefore, the classification and grading tasks of these data are often not convenient to train or process through large open source models. As the needs for data classification and grading become more complex, a single model often cannot strike a balance between efficiency and accuracy.
[0031] Understandably, at this stage, AI technology can only solve practical problems and achieve greater impact and value if it is focused more closely on specific scenarios. In real-world industry applications, a "large model guides, small model executes" strategy can be adopted. Large models, with their powerful generalization capabilities and rich knowledge base, provide valuable resources for small models, such as foundational model architecture and pre-trained weights. For example, in intelligent assistant applications, large models can handle complex semantic understanding and contextual analysis, while small models can not only instantly execute simple user commands or perform local reasoning, but also fully safeguard data security within industries and enterprises. This combination of large and small models leverages the powerful learning capabilities of large models while allowing small models to deliver on their efficiency and convenience in real-world applications. It also addresses concerns among industries and enterprises regarding data leakage and privacy.
[0032] Since data in the medical industry is not only multi-source, complex, and dynamic, but also highly sensitive and professional, as an industry that is related to the national economy and people's livelihood, the medical industry especially needs to protect patient privacy, ensure the security of medical services, and promote data security governance. Only in this way can it effectively protect important and sensitive data, effectively support data sharing and exchange, and actively respond to data security threats.
[0033] In response to the problems existing in the prior art, the present invention provides a data classification and grading processing method, which aims to solve the following technical problems: 1. By focusing on the medical industry, we build small vertical domain models to address data identification, parsing, and labeling issues in data classification and grading scenarios within the medical industry. Specifically, we address the issue of inaccurate labeling and mining results due to limited understanding of industry terminology, abbreviations, and specialized knowledge. This reduces the reliance on professional analysis and significantly improves the efficiency and practicality of data classification and grading. Furthermore, these small vertical domain models facilitate the processing of sensitive data, preventing data leaks.
[0034] 2. Optimize specific tasks in vertical fields by building small vertical field models, and study and analyze loss functions, evaluation indicators and training strategies specific to the medical industry.
[0035] 3. Solve the black box problem of multimodal large models by locally deploying small models in vertical fields, especially when processing sensitive medical data, to solve the risk of privacy leakage caused by the opaque decision-making process.
[0036] 4. By studying the construction and integration of large and small models for data classification and grading scenarios in the medical industry, we can solve the problem that the vertical domain model has not yet formed a medical professional knowledge map and has not formed ideas and specific methods for analyzing the segmented scenarios.
[0037] 5. By introducing a multimodal large model, we address the challenges of processing and analyzing multi-source, heterogeneous data in the medical industry, resolving issues that can lead to insufficient information utilization, insufficient domain adaptability, data bias, modality alignment issues, poor interpretability, and a lack of domain knowledge. Through technical advantages such as unified feature representation, cross-modal information fusion, and powerful generalization capabilities, we effectively reduce reliance on manual analysis and ensure efficient and accurate analysis of semi-structured and unstructured data.
[0038] 6. By combining multimodal large models and vertical small models, an application method for the fusion of large and small models in data classification and grading is proposed. The professional knowledge base and vector library formed in the vertical model of the medical industry are mainly used to solve the problem that the existing technology has not proposed the specific implementation process of reasoning.
[0039] like Figure 1As shown, the embodiment of the present application provides a data classification and grading processing method, which may include the following steps: S1. Collect multi-source data corresponding to the classification and grading tasks, and use the pre-trained multi-modal feature extraction model to extract multi-modal features of the multi-source data.
[0040] For example, the classification and grading task can be a data classification and grading task for the medical industry. Data (multi-source data) containing text, images, audio, video, and other modalities is collected from various data sources in the medical industry, such as electronic medical records, paper medical record scans, X-rays, and CT images. For example, before feature extraction, this data can be preprocessed, including standardized preprocessing operations such as cleaning, denoising, and normalization, to ensure data quality and consistency.
[0041] For example, a multimodal feature extraction model can use advanced deep learning algorithms to extract features from data in different modalities, generating feature vectors. These feature vectors are then concatenated and fused to generate a cross-modal semantic vector representation, forming a unified multimodal feature representation.
[0042] S2. Use a pre-trained classification and grading model to process the multimodal features to obtain the data classification and grading results output by the classification and grading model; wherein the classification and grading model is composed of a fusion of a pre-trained multimodal large model and a pre-trained vertical domain small model.
[0043] Specifically, we first focus on the vertical field of the medical industry by collecting and organizing knowledge about concepts, entities, and relationships within the field to construct a medical expertise graph (knowledge base). This knowledge graph stores domain knowledge in a graph structure, where nodes represent concepts or entities and edges represent relationships between them. This knowledge graph enables semantic understanding and reasoning of medical industry knowledge and data, providing rich prior knowledge for training small models.
[0044] Then, based on the medical expertise graph and the corresponding annotated data, a suitable machine learning algorithm or deep learning architecture is used to train a small model for the medical industry's vertical domain (the vertical domain small model). This small model focuses on learning specific patterns and regularities within the domain, enabling more accurate classification and prediction of medical industry-related data.
[0045] It can be understood that the classification and grading model based on the fusion of the multimodal large model and the vertical domain small model can process the multimodal features and output the corresponding data classification and grading results.
[0046] For example, the multimodal features extracted by a multimodal feature extraction model (e.g., the multimodal data preprocessing module within a large multimodal model) can be fused with the features output by a small vertical domain model. Feature fusion can be performed using methods such as concatenation, addition, and multiplication to fully leverage multimodal information and medical industry expertise to produce richer and more discriminative feature representations.
[0047] For example, an attention mechanism can be introduced to automatically learn the weight distribution between features based on the different modalities of the data and the relevance to the medical industry. The weights are then used to fuse the multimodal features with the features output by the small vertical model. The attention mechanism can highlight important features and suppress irrelevant ones, further improving the effectiveness of feature fusion and the accuracy of classification and grading.
[0048] For example, the fused features are then fed into a classification and grading module (e.g., a pre-trained classification and grading transformer model) within a large multimodal model for joint classification and grading. The large multimodal model typically achieves final classification and grading results by applying nonlinear feature transformations and global information modeling, thereby enabling final classification and grading decisions for data in the healthcare vertical.
[0049] S3. Determine the final classification and grading decision strategy based on the classification and grading results.
[0050] For example, the classification and grading results of the two models (the multimodal large model and the vertical domain small model) are fused by formulating rules based on the knowledge of medical industry experts; for example, in the comprehensive decision-making of data classification, identification and grading judgment, according to the rules formulated by experts (such as when the order of magnitude of data of the same type exceeds a threshold, a decision of a higher data level should be determined); the identification results of the multimodal large model on the data type and the judgment results of the vertical domain small model on the specific magnitude are fused to obtain the final classification and grading decision strategy.
[0051] Based on this, by combining large models and small models to perform data classification and grading tasks in specific fields, we can not only fully utilize the versatility and powerful learning ability of large models, but also give full play to the advantages of efficiency, convenience and privacy protection of small models, thereby improving the accuracy of data classification and grading while protecting data privacy.
[0052] It should be noted that in some embodiments, the method of the present application mainly includes the following steps: (1) Multimodal data preprocessing; (2) Construction of small models in vertical fields; (3) Fusion of multimodal large models and vertical domain small models; (4) Joint data classification and grading decision-making; (5) Model optimization and evaluation.
[0053] In some embodiments, step S1, collecting multi-source data corresponding to the task to be classified and graded, and extracting multi-modal features of the multi-source data using a pre-trained multi-modal feature extraction model, may include: For text data in multi-source data, extract features from the text data using a preset text feature extraction model to obtain text feature data; wherein the text feature extraction model includes a bag-of-words model, a word vector model, or a pre-trained transformer model; For image data in multi-source data, a pre-trained convolutional neural network model is used to extract features from the image data to obtain image feature data; For audio data in multi-source data, a preset audio analysis toolkit is used to extract features from image data to obtain acoustic feature data; Feature fusion operations are performed based on text feature data, image feature data, and acoustic feature data to obtain multimodal features of multi-source data; wherein, the feature fusion operations include feature vector splicing, feature alignment, and feature dimensionality reduction.
[0054] Specifically, the multimodal data preprocessing process mainly includes data collection and integration, data cleaning and processing, multimodal feature extraction and feature fusion.
[0055] In some embodiments, the data collection and integration steps first identify the multimodal data types required for the current classification and grading task, such as medical record text, X-ray images, audio recordings of consultations, and videos of treatment areas. Then, the channels for acquiring this data are determined, such as electronic medical record (EMR) databases, paper medical record scanning and entry systems, hospital radiology equipment, and clinical laboratory instruments.
[0056] After determining the data source, data can be collected from various data sources to ensure its completeness and accuracy. For example, medical record text data, including chief complaint, current medical history, past medical history, family history, personal history (such as lifestyle, occupation, etc.), and physical examination results, can be collected from hospital information systems (HIS); medical imaging data, such as X-rays, CT images, and MRI images, can be obtained from medical imaging equipment (such as X-ray machines, CT scanners, and MRIs).
[0057] For example, the collected data can be de-identified to protect patient privacy.
[0058] In addition, the collected data can be preliminarily screened to remove obviously erroneous or irrelevant data.
[0059] Finally, the same type of data from different data sources is merged to form a unified dataset. Exemplarily, the data can be labeled and classified for subsequent processing and analysis.
[0060] In some embodiments, for the data cleaning and processing steps, first check whether there are missing values in the dataset. For missing values, methods such as mean filling, median filling, mode filling, etc. can be used for processing, or the records containing missing values can be deleted according to the specific situation. Then identify the noise data in the data, such as outliers, duplicate values, etc.; for outliers, they can be judged and processed through statistical methods or domain knowledge, such as deletion, correction or transformation; for duplicate values, perform a de-duplication operation to retain a single unique record. Finally, ensure that the data formats in the dataset are consistent, such as date format, numerical format, text encoding format, etc.; convert or standardize the data that does not meet the format requirements.
[0061] In some embodiments, for the multi-modal feature extraction step, it can be divided into text data processing, image data processing, audio data processing, video data processing, etc.
[0062] Among them, for text data processing, first perform lexical analysis, tokenize the text data, decompose the text into words or vocabulary, and remove stop words, such as common meaningless words like "of", "is", "in", etc. Then perform stemming or lemmatization to restore the words to their original forms. Second, perform syntactic analysis, analyze the grammatical structure of the sentence, determine the dependency relationships and syntactic components between words, and then construct a syntactic tree to represent the syntactic structure of the sentence. Third, perform semantic analysis, extract semantic information such as entities, relationships, events, etc. in the text by understanding the meaning of the text. Additionally, natural language processing techniques, such as named entity recognition, relation extraction, etc., can be used to perform semantic annotation on the text. Finally, perform text feature extraction. For example, methods such as the bag-of-words model, word vector models (such as Word2Vec, GloVe), or pre-trained Transformer models (such as BERT) can be used to convert the text data into text feature vectors.
[0063] Image data processing begins with image enhancement. This involves performing operations like rotation, flipping, scaling, and cropping to increase image diversity and robustness. Color parameters like brightness, contrast, and saturation can also be adjusted to improve image quality. Next, normalization is performed. For example, pixel values can be normalized to a specific range, such as [0, 1] or [-1, 1]. Furthermore, the image dimensions are normalized to a uniform size. Finally, feature extraction is performed. Deep learning algorithms such as convolutional neural networks (CNNs) are used to extract visual feature vectors from the image. For example, pre-trained CNN models such as ResNet and VGG can be used for image feature extraction.
[0064] For audio data processing, first perform audio noise reduction to remove noise signals and improve audio quality. For example, methods such as filters and spectral subtraction can be used for noise reduction. Feature extraction is then performed, such as using Mel-Frequency Cepstral Coefficients (MFCCs) and chrominance vectors to extract audio features. For example, audio analysis toolkits such as Librosa can also be used to extract acoustic features from the audio. Finally, perform audio conversion to convert the audio data into a format suitable for model processing, such as a spectrogram or waveform.
[0065] Video data processing begins with video frame extraction, breaking the video down into a series of image frames. Frame extraction can be performed, for example, at fixed time intervals or using keyframe extraction methods. Next, video enhancement is performed, for example, by performing image enhancement operations on the video frames, such as rotation, flipping, and scaling. Parameters such as the frame rate and resolution can also be adjusted to improve video quality. Finally, feature extraction can be performed using deep learning algorithms such as three-dimensional convolutional neural networks (3D CNNs) or spatiotemporal neural networks (ST-CNNs) to extract the spatiotemporal feature vectors of the video. For example, comprehensive feature extraction can be performed by combining image and audio feature extraction methods.
[0066] In some embodiments, the feature fusion step can first perform feature vector concatenation, which involves simply concatenating feature vectors from different modalities to form a combined feature vector, ensuring that the concatenated feature vector has the appropriate dimensionality and order. For example, feature vectors from the same modality are aligned to ensure temporal and spatial consistency; for example, methods such as dynamic time warping (DTW) can be used for feature alignment. Furthermore, if the dimensionality of the fused feature vector is too high, feature dimensionality reduction can be performed; for example, methods such as principal component analysis (PCA) and linear discriminant analysis (LDA) can be used for feature dimensionality reduction.
[0067] It can be understood that the above multimodal data preprocessing process can convert the original multimodal data into a feature representation suitable for classification and grading model processing, providing a good data foundation for subsequent classification and grading tasks. In practical applications, the preprocessing process can be adjusted and optimized according to the specific data characteristics and needs of the medical industry.
[0068] In some embodiments, in step S2, the method for constructing the vertical domain small model includes: Collect structured and unstructured data in the vertical fields of the target industry, and build a domain expertise database based on structured and unstructured data; Based on the domain expertise database and preset annotated data, several preset initialization models are used for training to obtain several corresponding vertical domain small models; among them, the preset initialization models include the BioBERT model for processing text data, the DenseNet model for processing image data, and the CLIP model for processing multimodal data; The routing mechanism of each vertical field small model is configured to activate the corresponding type of vertical field small model according to the modal type of the input data.
[0069] In some embodiments, when building a small vertical domain model, structured data (such as electronic medical record system databases) and unstructured data (such as domain literature, expert experience, etc.) in the vertical domain of the medical industry are first collected. This data is then cleaned, annotated, and vectorized to construct a domain expertise graph or embedded representation. For example, a graph database (such as Neo4j) or a vector database (such as FAISS) can be used to store this collected domain knowledge, supporting efficient retrieval and matching.
[0070] Then, select a pre-trained small model based on the characteristics of the vertical field. For example, for text data, you can choose a variant of the BERT model (such as BioBERT for text data in the medical industry); for image data, you can choose a variant of the ResNet model (such as DenseNet for medical imaging); for multimodal data, you can choose a variant of the CLIP model (such as CLIP fine-tuned in the medical field).
[0071] The model is then initialized by loading pre-trained weights and freezing most of the parameters to reduce computational overhead. In addition, a lightweight adaptation layer (such as Adapter and LoRA) is added to the top layer of the model to inject medical expertise.
[0072] For example, an adapter module (a small neural network) can be inserted into each layer of the pre-trained model. During lightweight fine-tuning, only the adapter parameters are trained, leaving the original model parameters unchanged. The adapter module's input is the common features extracted by the multimodal large model, and its output is a feature representation adapted to medical knowledge.
[0073] For example, low-rank decomposition (LoRA) can be introduced into the weight matrix of the pre-trained model, and efficient parameter updates can be achieved by training the low-rank matrix. Based on this, the training parameters can be significantly reduced and the computational cost can be reduced.
[0074] For example, during the model fine-tuning process, you can use labeled data in the domain knowledge base (such as medical images + diagnostic reports) for fine-tuning; you can also perform data augmentation to expand the fine-tuning training set through data generated by a large multimodal model.
[0075] It should be noted that the routing mechanism for the vertical domain small model is configured as input for the common features and data modality labels (such as text and images) extracted by the multimodal large model, and the output is configured to activate the branch of the corresponding vertical domain small model. The routing strategy is modality-based routing: the corresponding vertical domain small model is selected based on the modality of the input data (such as case text → BERT, image → ResNet). For example, computing resources can be dynamically allocated by combining modality and medical industry domain information. In addition, reinforcement learning (such as the PPO algorithm) can be used to optimize the routing strategy to maximize the balance between classification accuracy and efficiency.
[0076] For example, a cross-modal attention mechanism can also be introduced to achieve feature alignment (such as text-image alignment) by introducing an attention mechanism between multimodal features and small model features; in addition, the embedding representation (such as medical term vector) in the medical industry domain knowledge base can be used to enhance the feature representation of the small model.
[0077] It is understandable that the output of the vertical domain model is a fine-grained classification and grading result, such as "personal privacy level P3", "commercial confidentiality level C2", etc. For example, when outputting the classification and grading results, the corresponding confidence score can also be output for subsequent decision fusion.
[0078] For example, the classification and grading results can be fed back to the medical industry professional knowledge base to dynamically update the knowledge graph. In addition, small models can be incrementally trained with new data regularly to improve model performance.
[0079] In some embodiments, in step S2, using a pre-trained classification and grading model to process the multimodal features to obtain data classification and grading results output by the classification and grading model may include: The multimodal features are processed using the vertical domain small model to obtain the vertical domain feature data extracted by the vertical domain small model; Fuse the multimodal features and vertical domain feature data to obtain a fused feature vector; The fused feature vector is input into the pre-trained classification and grading converter model to obtain the data classification and grading results.
[0080] It should be noted that the fusion process of a large multimodal model and a small vertical model can include three aspects: the first is feature extraction and alignment (feature fusion), the second is model integration or distillation (model fusion), and the third is decision fusion (weighted fusion or rule-based fusion).
[0081] Among them, feature fusion can adopt early or mid-term feature fusion. For early feature fusion, the multimodal features processed by the multimodal large model (multimodal data preprocessing module) and the features of the vertical domain small model can be spliced (fused). For example, the text feature vector and image feature vector extracted by the large model are spliced with the domain feature vector extracted by the small model to form a new comprehensive feature vector. For mid-term feature fusion, the intermediate representation (such as hidden layer state) of the multimodal large model and the vertical domain small model can be obtained and fused through a specific fusion mechanism (such as weighted summation, attention mechanism, etc.). For example, in the task of processing the fusion of natural language and medical data, the attention mechanism is used to combine the attention weight of the large model on the text with the feature weight of the small model on the medical data to highlight the important information in the data.
[0082] In some embodiments, in step S2, using a pre-trained classification and grading model to process the multimodal features to obtain data classification and grading results output by the classification and grading model may include: The multimodal large model and the vertical domain small model are used to make predictions based on the multimodal features, respectively, to obtain the corresponding first classification and second classification results; The first classification and grading results and the second classification and grading results are integrated based on the voting method to obtain the data classification and grading results.
[0083] Specifically, model integration or distillation (model fusion) can include three methods: voting-based model integration, stacking-based model integration, and model distillation. For voting-based model integration, a voting method can be used. In classification or grading tasks, predictions are first made based on a large multimodal model and a small vertical model. The final data classification and grading results are then determined using a voting strategy (such as majority voting).
[0084] In some embodiments, in step S2, using a pre-trained classification and grading model to process the multimodal features to obtain data classification and grading results output by the classification and grading model may include: Input the multimodal features into the multimodal large model to obtain the preliminary classification and grading results output by the multimodal large model; The preliminary classification and grading results are output to the vertical domain small model to obtain the data classification and grading results output by the vertical domain small model.
[0085] It's important to note that for model integration based on stacking, the output of a large multimodal model can be used as input for smaller vertical domain models. For example, the large multimodal model first processes medical data (including case text and images) to obtain preliminary topic classification results. This result is then input into a smaller vertical domain model (a data classification and grading model specifically for the medical industry) to output the final data classification and grading results.
[0086] In some embodiments, model distillation can also be employed. For example, a large multimodal model serves as the teacher model, and a smaller, specialized model serves as the student model. The soft targets (e.g., probability distributions) output by the teacher model serve as training labels for the student model. For example, in a medical data recognition task, the large model has already learned a rich set of feature representations for specialized medical vocabulary. By transferring this knowledge to the smaller model through model distillation, the smaller model can quickly acquire similar specialized data recognition capabilities even with a small amount of data.
[0087] In some embodiments, in step S2, using a pre-trained classification and grading model to process the multimodal features to obtain data classification and grading results output by the classification and grading model may include: The multimodal large model is used to make predictions based on the multimodal features to obtain a first classification and grading result and a corresponding first confidence score. At the same time, the vertical domain small model is used to make predictions based on the multimodal features to obtain a second classification and grading result and a corresponding second confidence score. A weighting coefficient is determined based on the first confidence score and the second confidence score, and the first classification and grading result and the second classification and grading result are weightedly fused based on the weighting coefficient to obtain a data classification and grading result.
[0088] In an embodiment of the present application, the accuracy and reliability of data classification and grading are improved by utilizing a large multimodal model and a small vertical domain model to output prediction results and their corresponding confidence scores respectively, and fusing the outputs of the two models in combination with the confidence scores.
[0089] It should be noted that there are three main methods for decision fusion: weighted fusion based on confidence, weighted fusion based on performance indicators, and rule-based decision fusion.
[0090] Specifically, for confidence-based weighted fusion, the classification and grading results and corresponding confidence levels of the multimodal large model and the vertical domain small model are first obtained. For example, in a specific data classification and grading task, the confidence level is determined based on the probability value output by the large model and the accuracy of the small model in existing classifications in the medical industry. A weighting coefficient is then determined based on the confidence level, and the prediction results of the two models are weighted and summed to obtain the final classification and grading decision.
[0091] In some embodiments, for weighted fusion based on performance metrics, the weighting coefficients can be determined based on the performance metrics (e.g., accuracy, precision, recall, etc.) of the two models on the validation set or test set. For example, in an initial classification task, if the large model exhibits a higher recall rate in identifying unknown data types, while the small model has a higher precision rate in detecting known data types, the weighting ratio can be adjusted based on these performance metrics so that the final decision can comprehensively consider the advantages of both models.
[0092] In some embodiments, rule-based decision fusion can be implemented by combining the knowledge of medical experts to formulate rules that integrate the decisions of large and small models. For example, in the integrated decision-making of data classification and grading, the multimodal large model's recognition of data types and the vertical domain small model's judgment of specific magnitudes are combined according to expert-defined rules (e.g., when the magnitude of data items of the same category exceeds a threshold, a decision should be made at a higher data level), to produce the final classification and grading decision.
[0093] In some embodiments, the joint data classification and grading decision process may include two steps: joint confirmation of the data classification and grading strategy and strategy implementation and monitoring.
[0094] Specifically, the steps for jointly confirming data classification and grading strategies begin with model prediction and result analysis. For example, the fused model can be used to perform inference predictions on the test set data to determine the data classification and grading strategy. The prediction results are then analyzed in depth to compare the data distribution under different classification and grading conditions, as well as metrics such as model accuracy and recall.
[0095] Expert evaluation and feedback are then conducted. Specifically, medical industry experts can be invited to evaluate the model strategy results and, based on their expertise and experience, determine the rationality and effectiveness of the classification and grading strategy. Experts can provide opinions and suggestions based on business logic, industry standards, and practical application value. The classification and grading strategy can then be adjusted based on expert feedback. For example, if experts believe that a certain level of classification is too detailed or coarse and does not meet actual business needs, the classification and grading strategy can be modified accordingly.
[0096] Finally, the process involves iterative optimization and validation. This involves applying the adjusted classification and grading strategy to the model, retraining, and re-evaluating it. Through multiple iterations, the classification and grading strategy is continuously optimized until the model performance and expert acceptance reach a high level. Finally, the current data classification and grading strategy is finalized, ensuring it meets medical business needs while fully leveraging the advantages of integrating large and small multimodal models.
[0097] Regarding the steps of strategy implementation and monitoring, the first is strategy implementation. The data classification and grading strategy can be applied to the actual medical business system to classify and grade new data; by ensuring that the system can accurately classify and grade medical data according to the strategy, a reliable basis is provided for subsequent decision-making.
[0098] Next comes monitoring and adjustment. Because data classification and grading strategies are not static, a real-time monitoring mechanism must be established to evaluate the effectiveness of the strategy in real time or regularly. By monitoring changes in data distribution and fluctuations in model performance, problems can be identified and optimized promptly. For example, when medical business scenarios change or new data features emerge, the classification and grading strategy and model must be updated promptly.
[0099] It should be noted that the performance of data classification and grading can also be continuously improved through model optimization and evaluation processes. For example, a reasonable loss function can be designed by comprehensively considering multiple indicators such as classification accuracy, recall rate, F1 value, grading accuracy, mean square error, etc. to measure the performance of the model; different weight settings can be used to balance the importance of each indicator for different classification tasks and grading requirements. For example, a suitable optimization algorithm can be selected to optimize the parameters of the model to improve the convergence speed and generalization ability of the model; overfitting should be prevented during the optimization process. Finally, various methods such as cross-validation, confusion matrix, ROC curve, etc. are used to evaluate and verify the model; by conducting experiments on different data sets in the medical industry, the performance indicators of the model are compared, and the structure and parameters of the model are continuously adjusted until satisfactory performance is achieved.
[0100] For example, a professional dataset in the medical field can be used to train and validate the system. During training, model hyperparameters (such as learning rate, batch size, and number of training rounds) are adjusted based on the validation set's loss function and performance metrics. Early stopping is used to terminate training when the validation set loss stops decreasing over multiple epochs. L1 and L2 regularization terms are used to constrain model parameters and prevent excessive model complexity.
[0101] For example, a cross-validation method can be used to evaluate the stability and generalization ability of a model. The dataset is divided into multiple subsets, and one subset is used alternately as the validation set, while the remaining subsets are used as the training set for training and evaluation. Based on the results of cross-validation, the model structure and algorithm are further optimized. For example, k-fold cross-validation is used to divide the dataset into k subsets, and each subset is used as the validation set for evaluation.
[0102] For example, the system is updated and expanded by regularly collecting new medical data to adapt to the continuous development of the medical field and the emergence of new data types. Newly collected case data is added to the training set at regular intervals to retrain the model, allowing it to learn and grasp the latest medical knowledge and classification and grading patterns.
[0103] Exemplarily, the model optimization and evaluation process may include the following aspects: 1. Comprehensive evaluation and bottleneck analysis; 2. Model structure adjustment; 3. Hyperparameter optimization; 4. Data enhancement and cleaning; 5. Functional evaluation; 6. Performance evaluation; 7. Robustness evaluation; 8. Interpretability evaluation.
[0104] Specifically, comprehensive evaluations can be performed regularly on the fused model using validation and test sets. Evaluation metrics include, but are not limited to, accuracy, recall, F1-score, mean squared error (for regression tasks), and area under the receiver operating characteristic (ROC) curve (AUC). Separate evaluations can also be performed for different modalities and classification tasks, such as evaluating image recognition accuracy in image classification tasks and assessing text accuracy and relevance in medical record text generation tasks. Regarding bottleneck analysis, performance bottlenecks can be identified through in-depth analysis of the model's performance across different aspects. For example, this can be done to examine overfitting or underfitting, determining whether these are caused by model structure or data issues. Furthermore, for multimodal fusion, the effectiveness of the fusion of features from different modalities can be analyzed to determine whether certain features are underutilized or poorly fused. For example, visualization tools can be used to visualize the representation and contribution of features from different modalities in the fused model.
[0105] For model structure adjustments, you can consider increasing or decreasing the model's hierarchy based on the performance evaluation results. If the model is underfitting, you can appropriately increase the number of neural network layers to improve the model's complexity and fitting ability. For example, you can increase the number of layers in the Transformer encoder part of a large multimodal model. For overfitting, you can reduce the number of layers or adopt a simpler network structure, such as changing a complex multi-layer perceptron to a single-layer perceptron. In addition, you can also make adjustments to the specific processing modules of each modality. For the imaging modality, you can adjust the number, size, and step size of the convolutional neural network's filters. For the medical record text modality, you can change the number of heads, hidden unit size, and other parameters of the Transformer model. You can also optimize the level and method of inter-modal fusion, such as adjusting the weight coefficient during fusion, or trying different fusion positions (from early and mid-term fusion to late fusion).
[0106] For hyperparameter optimization, you can choose an appropriate learning rate optimization algorithm, such as an adaptive learning rate algorithm (AdaGrad, AdaDelta, Adam, etc.), and use techniques such as grid search or random search to find the optimal learning rate combination. Dynamically adjust the learning rate based on the changes in the loss function during training. For example, when the loss function decreases slowly, appropriately reduce the learning rate; when the loss function oscillates and fails to converge, consider increasing the learning rate. In addition, you can also optimize other hyperparameters such as the regularization term hyperparameters (such as the strength of L1 and L2 regularization), batch size, and the momentum parameter of the optimizer. Optimizing these hyperparameters can further balance the model's fit and generalization capabilities. It is understood that by using methods such as cross-validation to determine the optimal hyperparameter combination, we can ensure that the model performs well on different data subsets.
[0107] For data enhancement and cleaning, data quality can be rechecked to remove outliers, noisy data, and incorrectly labeled data. Missing values can be filled or deleted as appropriate. Data consistency can be checked, such as whether the labels corresponding to different modal data are consistent, to avoid model bias caused by data inconsistency. In addition, for image data, data enhancement can be performed by rotating, flipping, scaling, cropping, adding noise, etc. to increase data diversity and improve the generalization ability of the model. For medical record text data, synonym replacement, sentence reorganization, random insertion and deletion of words can be performed to expand the data set.
[0108] Functional evaluation can include task accuracy evaluation and functional integrity evaluation. In task accuracy evaluation, the output accuracy of the model can be evaluated for specific data classification and grading tasks. For example, in the task of automatic image data labeling, the accuracy of the image category determined by the model can be checked. In the task of medical record text analysis, the accuracy of the model's judgment of the sensitivity of text data can be evaluated. For multimodal generation, classification and grading labeling tasks, the quality of the generated annotations and their relevance to the input modality can be evaluated. This can be done through a combination of manual evaluation and automated indicators, such as using indicators such as BLEU and METEOR to evaluate the quality of the generated annotations.
[0109] Functional integrity assessments examine whether a model can fully implement its intended functionality. For example, a medical image must not only accurately identify its data type but also determine its classification based on its importance and impact. By assessing the model's functional stability across various scenarios and input conditions, we ensure that the model can function properly under a variety of complex circumstances.
[0110] Performance evaluation can include efficiency and resource utilization assessments. Efficiency assessments assess the model's computational efficiency, including training and inference time. By recording the model's runtime on different hardware devices (such as the CPU, GPU, and NPU), we can analyze which parts can be further optimized to improve efficiency. Resource utilization assessments measure the model's utilization of computing resources (such as memory and video memory) and storage resources during runtime. By analyzing the model's resource utilization, we can identify areas for optimization, such as reducing resource usage through techniques like model compression and pruning.
[0111] Robustness assessment can include noise resistance assessment and adversarial attack assessment. In the noise resistance assessment, noisy data (such as disturbed images, medical records containing spelling errors) can be input into the model to evaluate whether the model's output is still stable and accurate, so as to check the model's robustness under different noise types and intensities. The performance of the model can be observed by adjusting the noise parameters. In addition, adversarial attack techniques (such as adversarial sample generation) can be used to evaluate the adversarial robustness of the model. By generating adversarial samples and inputting them into the model, it is observed whether the model will produce incorrect outputs to evaluate the model's ability to resist different types of adversarial attacks (such as malicious attacks and misleading attacks). This is an important evaluation indicator for the implementation of data classification and grading in the medical industry, which has high security and accuracy requirements.
[0112] For interpretability assessment, appropriate model explanation methods can be applied, such as feature importance analysis, attention visualization, and concept explanation. For models that combine a large multimodal model with a small vertical model, explain the contribution of each modality's features to the final decision. For the evaluation of explanation results, the explanation results should be consistent with the actual behavior of the model and help users understand the model's reasoning and decision-making process. For example, the attention weights displayed by visualization tools should be able to highlight important image regions or medical record text fragments, and these regions or fragments should be relevant to the predicted results. Medical industry experts are invited to review the explanation results to ensure that the explanations are consistent with medical business logic and professional knowledge.
[0113] Based on this, through the above comprehensive model optimization and evaluation process, the performance and quality of the fusion of multimodal large models and vertical field small models can be continuously improved to meet the needs of actual classification and grading applications of medical data, and provide a strong basis for subsequent model improvement and application expansion.
[0114] It should be noted that the embodiments of the present application have the following beneficial effects: 1. Effectively improves the accuracy and efficiency of data classification and grading: By introducing a multimodal large model, the system addresses the challenges of processing and analyzing multi-source heterogeneous data in the medical industry, resolving issues that can lead to insufficient information utilization, insufficient domain adaptability, data bias, modality alignment issues, poor interpretability, and a lack of domain knowledge. The multimodal large model integrates heterogeneous data from multiple sources, including text, imaging, pathology, and genetics, and combines it with refined analysis using small vertical models. This significantly improves data recognition accuracy, reduces reliance on manual analysis, and enables precise multi-dimensional classification and grading of medical data.
[0115] 2. Effectively balances the universality and domain specificity of data processing: A collaborative working framework is established between small models in vertical medical fields and large multimodal models. The large model provides a common knowledge base, covering common data and basic medical logic; the small vertical models are deeply optimized for medical scenarios to form a medical professional knowledge map. Through feature interaction and information sharing, the two complement each other's strengths, achieving a collaborative reasoning model of "global generalization + local precision." This collaborative working mechanism fully leverages the global semantic understanding capabilities of the large multimodal model and the specialized targeting of the small vertical models, improving the accuracy and reliability of data classification and grading.
[0116] 3. Localized Deployment and Data Isolation: Small, vertically-specific models are constructed to reduce data transmission links through localized data processing. Physical security measures (such as private server deployment) and logical isolation technologies are combined to mitigate the risk of sensitive data leakage. Furthermore, knowledge distillation technology is used to migrate the capabilities of large models to lightweight, small models, significantly reducing inference resource consumption. This adapts to the low-computing environments of primary healthcare institutions, lowering computing costs and deployment complexity.
[0117] 4. Strengthening Data Security and Compliance Governance: Small vertical models focus on specific areas, reducing redundant data exposure. Combined with automated desensitization technology, they facilitate the processing of sensitive data, reduce the risk of leakage, and prevent data leaks. Furthermore, they ensure that the classification and grading of highly sensitive information, such as genetic data and rare disease data, complies with laws and regulations such as the Data Security Law. A hierarchical permission control module dynamically adjusts access permissions based on data sensitivity to prevent unauthorized use.
[0118] 5. Promote intelligent upgrades in medical scenarios: Support data classification and grading applications in multiple scenarios, including drug development, intelligent diagnosis and treatment, and scientific research analysis. For example, in AI-powered pharmaceutical manufacturing, multimodal data classification and grading accelerates target discovery and shortens the R&D cycle. This technology also assists primary healthcare institutions in rapidly completing standardized processing of unstructured data (such as fuzzy images and handwritten medical records), improving the efficiency of primary healthcare data security management.
[0119] Please refer to Figure 2 , Figure 2 The following is a block diagram showing the composition of the data classification and grading processing device provided by some embodiments of the present application. It should be understood that the data classification and grading processing device is similar to the above Figure 1 Corresponding to the method embodiment, it is able to execute each step involved in the above method embodiment. The specific functions of the data classification and grading processing device can be found in the description above. To avoid repetition, the detailed description is appropriately omitted here.
[0120] Figure 2 The data classification and grading processing device includes at least one software function module that can be stored in a memory in the form of software or firmware or fixed in the data classification and grading processing device, and the data classification and grading processing device includes: The feature extraction module 210 is used to collect multi-source data corresponding to the classification and grading tasks, and extract multimodal features of the multi-source data using a pre-trained multimodal feature extraction model; The fusion output module 220 is used to process the multimodal features using a pre-trained classification and grading model to obtain data classification and grading results output by the classification and grading model; wherein the classification and grading model is formed by fusing a pre-trained multimodal large model and a pre-trained vertical domain small model; The strategy determination module 230 is used to determine the final classification and grading decision strategy based on the classification and grading results.
[0121] It can be understood that the above-mentioned device embodiment corresponds to the method embodiment of the present invention. The data classification and grading processing device provided by the embodiment of the present invention can implement the data classification and grading processing method provided by any method embodiment of the present invention.
[0122] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method, and will not be described in detail here.
[0123] like Figure 3 As shown, some embodiments of the present application provide an electronic device 300, which includes: a memory 310, a processor 320, and a computer program stored in the memory 310 and executable on the processor 320, wherein the processor 320 reads the program from the memory 310 through the bus 330 and executes the program to implement a method of any embodiment included in the above-mentioned data classification and grading processing method.
[0124] Processor 320 can process digital signals and can include various computing architectures, such as a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements a combination of multiple instruction sets. In some examples, processor 320 can be a microprocessor.
[0125] The memory 310 can be used to store instructions executed by the processor 320 or data related to the execution of instructions. These instructions and / or data may include code for implementing some or all functions of one or more modules described in the embodiments of this application. The processor 320 of the embodiment of the present disclosure can be used to execute the instructions in the memory 310 to implement the method shown above. The memory 310 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memory known to those skilled in the art.
[0126] Some embodiments of the present application further provide a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method described in the method embodiment is executed.
[0127] Some embodiments of the present application further provide a computer program product, which, when executed on a computer, enables the computer to execute the method described in the method embodiment.
[0128] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similarities between the various embodiments can be referred to in conjunction with each other. For device embodiments, since they are generally similar to method embodiments, their description is relatively simple, and for relevant details, reference can be made to the description of the method embodiments.
[0129] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a portion of code, and the module, program segment, or a portion of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0130] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0131] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.
[0132] The foregoing is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.
[0133] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0134] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
Claims
1. A data classification and grading method, characterized in that: include: Collecting multi-source data corresponding to the task to be classified and graded, and extracting multimodal features of the multi-source data using a pre-trained multimodal feature extraction model; The multimodal features are processed using a pre-trained classification and grading model to obtain data classification and grading results output by the classification and grading model; wherein the classification and grading model is composed of a fusion of a pre-trained multimodal large model and a pre-trained vertical domain small model; A final classification and grading decision strategy is determined based on the classification and grading results.
2. The data classification and grading method according to claim 1, characterized in that: The method of processing the multimodal features using a pre-trained classification and grading model to obtain data classification and grading results output by the classification and grading model includes: Inputting the multimodal features into the multimodal large model to obtain a preliminary classification and grading result output by the multimodal large model; The preliminary classification and grading results are output to the vertical domain small model to obtain the data classification and grading results output by the vertical domain small model.
3. The data classification and grading method according to claim 1, characterized in that: The method of processing the multimodal features using a pre-trained classification and grading model to obtain data classification and grading results output by the classification and grading model includes: Using the multimodal large model and the vertical domain small model to perform prediction based on the multimodal features, respectively, to obtain corresponding first classification and grading results and second classification and grading results; The first classification and grading result and the second classification and grading result are integrated based on a voting method to obtain the data classification and grading result.
4. The data classification and grading method according to claim 1, characterized in that: The method of processing the multimodal features using a pre-trained classification and grading model to obtain data classification and grading results output by the classification and grading model includes: Using the large multimodal model to make predictions based on the multimodal features to obtain a first classification and grading result and a corresponding first confidence score; and simultaneously using the small vertical domain model to make predictions based on the multimodal features to obtain a second classification and grading result and a corresponding second confidence score; A weighting coefficient is determined based on the first confidence score and the second confidence score, and the first classification and grading result and the second classification and grading result are weightedly fused based on the weighting coefficient to obtain the data classification and grading result.
5. The data classification and grading method according to claim 1, characterized in that: The method of processing the multimodal features using a pre-trained classification and grading model to obtain data classification and grading results output by the classification and grading model includes: Processing the multimodal features using the vertical domain small model to obtain vertical domain feature data extracted by the vertical domain small model; Fusing the multimodal features with the vertical domain feature data to obtain a fused feature vector; The fused feature vector is input into a pre-trained classification and grading converter model to obtain the data classification and grading result.
6. The data classification and grading method according to claim 1, characterized in that: The collecting of multi-source data corresponding to the to-be-classified and graded tasks, and extracting multi-modal features of the multi-source data using a pre-trained multi-modal feature extraction model, includes: For the text data in the multi-source data, extract features from the text data using a preset text feature extraction model to obtain text feature data; wherein the text feature extraction model includes a bag-of-words model, a word vector model, or a pre-trained transformer model; For the image data in the multi-source data, extract features of the image data using a pre-trained convolutional neural network model to obtain image feature data; For the audio data in the multi-source data, using a preset audio analysis toolkit to perform feature extraction on the image data to obtain acoustic feature data; A feature fusion operation is performed based on the text feature data, the image feature data and the acoustic feature data to obtain multimodal features of the multi-source data; wherein the feature fusion operation includes feature vector splicing, feature alignment and feature dimensionality reduction.
7. The data classification and grading method according to claim 1, characterized in that: The method for constructing the vertical domain small model includes: Collect structured data and unstructured data in the vertical field of the target industry, and build a domain expertise database based on the structured data and the unstructured data; Based on the domain expertise database and preset annotated data, several preset initialization models are used for training to obtain corresponding small models in vertical fields; wherein the preset initialization models include the BioBERT model for processing text data, the DenseNet model for processing image data, and the CLIP model for processing multimodal data; The routing mechanism of each vertical field mini-model is configured to activate the corresponding type of vertical field mini-model according to the modality type of the input data.
8. An electronic device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor can implement the data classification and grading processing method described in any one of claims 1 to 7 when executing the program.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the data classification and grading processing method according to any one of claims 1 to 7 is executed.
10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the data classification and grading processing method according to any one of claims 1 to 7.
Citation Information
Cited By
Multi-model anomaly detection method based on multi-modal time series data
CN121093238A
A multi-model anomaly detection method based on multi-modal time series data
CN121093238B
Automatic classification and grading method for unstructured data based on large model
CN121638418A