Multi-modal fusion unstructured data multivariate classification and grading method and system

Through multimodal fusion technology and multi-classification and grading method, the problems of low information extraction efficiency and low accuracy in unstructured data classification and grading are solved, and more efficient and flexible data management and protection are achieved to adapt to complex and changeable unstructured data.

CN120296558AActive Publication Date: 2025-07-11CHINESE PEOPLES LIBERATION ARMY 92493 UNIT INFORMATION TECH CENT
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510357941.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The prior art has problems such as low information extraction efficiency, low accuracy and poor flexibility in unstructured data classification and classification. Especially when dealing with complex and changeable unstructured data, traditional methods are difficult to capture complex semantic relationships and context information in documents, while deep learning methods face the problems of high data annotation cost and poor model interpretability.

Method used

Multimodal fusion technology is adopted to characterize the metadata and itself of unstructured data through semantic understanding models, combine with convolutional neural networks to fusion information, output multivariate data categories and levels, utilize multiple modal information of metadata and content data, and use multivariate classification hierarchical methods to calculate the maximum data level and expected data level.

Benefits of technology

It significantly improves the classification and grading accuracy and flexibility of unstructured data, and can give multiple categories to unstructured data at the same time, enhances the flexibility of data management and protection, and ensures information security while effectively utilizing existing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296558A_ABST
    Figure CN120296558A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data security classification and grading, and discloses a multi-modal fusion unstructured data multivariate classification and grading method and system. The classification method comprises: multi-modal data fusion, including document data multi-modal fusion, picture data multi-modal fusion, audio data multi-modal fusion and video data multi-modal fusion, to obtain a fusion representation vector; the fusion representation vector is input into a neural network module to output an M-dimensional vector, and M is a category number; converting the M-dimensional vector to obtain a probability vector P; obtaining a multivariate data category of the multi-modal data based on the probability vector P; a maximum data level and an expected data level of the unstructured data are calculated. According to the method, the multi-modal fusion technology is introduced, various effective information in the unstructured data is fully mined and utilized, the multi-element classification and grading technology is adopted, multiple categories, the maximum data level and the expected data level are given to the unstructured data, and the classification and grading effect of the unstructured data is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data security classification and grading, and in particular to a multi-modal fused unstructured data multivariate classification and grading method and system. Background Art

[0002] At present, the data classification and grading work of various organizations mainly focuses on structured data, while less attention is paid to the classification and grading of unstructured data. However, in various organizations, unstructured data (such as text, pictures, audio, video, etc.) accounts for more than 80% of the total data. Classifying and grading unstructured data is an important means to improve data management efficiency, reduce security risks, ensure compliance and mine data value. Due to the irregular and non-standard characteristics of unstructured data, its classification and grading is more challenging than structured data. At present, the mainstream classification and grading of unstructured data mainly relies on traditional methods, such as rule matching and simple text analysis. These methods have problems such as low information extraction efficiency and low accuracy when dealing with complex and changeable unstructured data. In recent years, with the development of deep learning technology, deep learning methods have also been applied to the classification and grading of unstructured data, but this method still faces challenges such as high data annotation costs and poor model interpretability.

[0003] Many unstructured data contain a large amount of information that is difficult for humans to describe with rules. For example, a piece of text may contain a variety of information such as emotions, themes, and opinions. The relationship between these information is intricate, and it is difficult for rule-based systems to capture and process this ambiguity. For example, pictures contain various irregular entity objects, and it is difficult to describe their internal laws or characteristics with a rule system. Therefore, it is very difficult to classify and grade unstructured data using rule-based technology. Even if some unstructured data, such as document data, can be barely classified and graded using rules, because document data is usually long and information is sparsely distributed, traditional rule matching methods are difficult to capture the complex semantic relationships and contextual information in the document, resulting in low information extraction efficiency and low accuracy. In addition, the flexibility and scalability of rule matching methods are poor, and it is difficult to adapt to the ever-changing data environment and classification needs. Therefore, for the classification and grading of document data, it is necessary to explore more efficient, accurate and flexible technical means to meet its unique challenges.

[0004] Although deep learning-based methods have shown great potential in processing unstructured data, this approach also has significant limitations. The primary issue is that the effective training of deep learning models highly depends on large-scale and high-quality labeled datasets. However, in many data classification and grading application scenarios involving sensitive or confidential information, obtaining a sufficient number of labeled samples is not only extremely difficult but also costly. This is mainly because data ownership and privacy protection have become major obstacles to data sharing. In addition, most existing deep learning methods focus on directly analyzing the content of unstructured data, such as text, images, and sounds. However, this approach ignores an important characteristic of unstructured data - the sparsity of information distribution. Coupled with the limited amount of sample data available for training, this makes even the most advanced deep learning models unable to achieve ideal performance levels in all cases. It should be noted that in the actual business environment, unstructured data not only contains valuable information in its intrinsic content but also the metadata surrounding this data, such as file names, file descriptions, file sources, file sizes, and the identities of file creators, often contains rich contextual information, which is also crucial for data understanding and utilization. However, most current deep learning-based unstructured data analysis solutions do not fully consider or utilize this additional metadata information, thus limiting their effectiveness and efficiency in practical applications.

[0005] In addition, most current methods for classifying and grading unstructured data tend to simply assign data to a single category or level. This method may be effective when dealing with data with relatively simple information, but for unstructured data with complex and diverse content, it is difficult to meet the diverse needs of practical applications. The complexity and diversity of unstructured data mean that it may belong to multiple categories simultaneously. Therefore, when dealing with such data, it is urgent to adopt more flexible and detailed strategies to accurately capture its multiple attributes. Summary of the Invention

[0006] The object of the present invention is to provide a multi-modal fusion-based method and system for multi-class classification and grading of unstructured data, innovatively introducing multi-modal fusion technology to fully explore and utilize various effective information in unstructured data. At the same time, by adopting multi-class classification and grading technology, multiple categories, a maximum data level, and an expected data level are assigned to unstructured data, which can significantly improve the classification and grading effect of unstructured data.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] In the first aspect, the present invention provides a multi-modal fusion-based method for multi-class classification and grading of unstructured data, including the following steps:

[0009] S1. Perform multimodal data fusion to obtain a fused representation vector;

[0010] S2. Input the fused representation vector into a neural network module to output an M-dimensional vector, where M is the number of classes;

[0011] S3. Transform the M-dimensional vector to obtain a probability vector P;

[0012] S4. Obtain the multivariate data classes of the multimodal data based on the probability vector P, specifically including:

[0013] S40. Set parameters k and t, where k is the maximum value of the selectable classes and t is the probability threshold;

[0014] S41. Screen out the classes from the probability vector P whose probability values are greater than or equal to the probability threshold t;

[0015] S42. Sort the screened classes in descending order of probability values. When the number of the screened classes is greater than or equal to k, take the first k classes as the multivariate data classes of the unstructured data; when the number of the screened classes is less than k, take all the classes as the multivariate data classes of the unstructured data;

[0016] S5. Calculate the maximum data level and the expected data level of the unstructured data; among them, the maximum data level is calculated by the following method:

[0017] S50. Calculate the data level vectors corresponding to each class in the multivariate data classes;

[0018] S51. Select the maximum value of the data levels from the data level vectors as the maximum data level of the unstructured data;

[0019] The expected data level is calculated by the following method:

[0020] S52. According to the index of each class in the multivariate data classes in the probability vector, obtain its corresponding logits from the M-dimensional vector, and further denote it as CLogits;

[0021] S53. Convert CLogits to a probability vector CP, and calculate the expected data level according to the probability vector CP.

[0022] As a possible implementation, S1 includes:

[0023] S10. Apply a semantic understanding model to represent the metadata of the unstructured data to obtain a metadata representation vector;

[0024] S11. Apply a semantic understanding model to represent the unstructured data itself to obtain an unstructured data representation vector;

[0025] S12. Stack the metadata representation vector and the unstructured data representation vector in the channel dimension;

[0026] S13. Apply the convolution kernel of the convolutional neural network to fuse the data in the above-mentioned multiple channels to obtain a fused representation vector.

[0027] As a possible implementation, the metadata includes at least file name, directory name, creator, descriptive information, data source, and creation date; and is represented by the RoBERTa semantic understanding model.

[0028] As a possible implementation, the unstructured data includes document data, which has three modalities: metadata, text data, and pictures. Among them, the metadata is represented by the RoBERTa semantic understanding model to obtain a metadata representation vector; the text data is represented by the Longformer semantic understanding model to obtain a text data representation vector; the pictures are first represented by the ViT semantic understanding model for each picture, and then the representation vectors of all pictures are fused by the Transformer encoder to obtain a picture representation vector; the metadata representation vector, the text data representation vector, and the picture representation vector are stacked in the channel dimension and then input into the convolutional neural network; the convolution kernel of the convolutional neural network fuses the metadata representation vector, the text data representation vector, and the picture representation vector on the three channels to obtain the fused representation of the document data.

[0029] As a possible implementation, the unstructured data includes picture data, which has two modalities: metadata and pictures. The metadata is represented by the RoBERTa semantic understanding model to obtain a metadata representation vector; the pictures are first represented by the ViT semantic understanding model for each picture, and then the representation vectors of all pictures are fused by the Transformer encoder to obtain a picture representation vector; the metadata representation vector and the picture representation vector are stacked in the channel dimension and then input into the convolutional neural network; the convolution kernel of the convolutional neural network fuses the metadata representation vector and the picture representation vector on the two channels to obtain the fused representation of the picture data.

[0030] As a possible implementation, the unstructured data includes audio data, which has two modalities: metadata and audio signals; among them, the metadata is represented by the RoBERTa semantic understanding model to obtain a metadata representation vector; the audio signals are preprocessed to obtain Mel-frequency cepstral coefficients and are represented by the Conformer semantic understanding model to obtain an audio signal representation vector; the metadata representation vector and the audio signal representation vector are stacked in the channel dimension and then input into the convolutional neural network; the convolution kernel of the convolutional neural network fuses the metadata representation vector and the audio signal representation vector on the two channels to obtain the fused representation of the audio data.

[0031] As a possible implementation, the unstructured data includes video data, and the video data has four modalities: metadata, subtitle sequence, audio signal, and frame sequence; among them, the metadata is characterized by a RoBERTa semantic understanding model to obtain a metadata characterization vector; the subtitle sequence is characterized by a RoBERTa semantic understanding model to obtain a subtitle sequence characterization vector; the audio signal is characterized by a Conformer semantic understanding model to obtain an audio signal characterization vector; the frame sequence is characterized by a VTN-ENCODER semantic understanding model to obtain a frame sequence characterization vector; the metadata characterization vector, subtitle sequence characterization vector, audio signal characterization vector, and frame sequence characterization vector are stacked in the channel dimension and then input into a convolutional neural network; the convolutional kernels of the convolutional neural network fuse the metadata characterization vector, subtitle sequence characterization vector, audio signal characterization vector, and frame sequence characterization vector on the four channels to obtain a fused characterization of the video data.

[0032] As a possible implementation, the formula for the expected data level is:

[0033]

[0034] where edl is the expected data level; 1 ≤ l ≤ m, l is the l-th category in the multi-modal data categories of the unstructured data, and m is the number of categories in the multi-modal data categories of the unstructured data; cp l represents the probability of the data category dc l , idx l is the index of the l-th category in the multi-modal data categories of the unstructured data in the probability vector P, is the logit corresponding to the l-th category in the multi-modal data categories of the unstructured data, j is a variable representing 1, 2, 3, …, m, is the logit corresponding to the j-th category in the multi-modal data categories of the unstructured data, the relevant data attribute of the unstructured data is attr, for a data category dc l , its corresponding data level is dl l , dl l = f(dc l , attr).

[0035] As a possible implementation, define the relevant data attribute of the unstructured data as attr, for one of its data categories dc l , its corresponding data level dl l can be expressed as:

[0036] dl l = f(dcl , attr)

[0037] According to dl l Calculate the data level vector DL corresponding to each category, and represent it as:

[0038]

[0039] Select the value with the largest data level from DL as the maximum data level mdl of the unstructured data, and represent it as:

[0040]

[0041] In a second aspect, the present invention provides a multi-modal fusion-based unstructured data multi-classification and grading system, including:

[0042] A multi-modal data fusion module that receives unstructured data, characterizes the metadata and itself of the unstructured data, stacks them in the channel dimension, and applies the convolutional kernels of a convolutional neural network to fuse the data in multiple channels to obtain a fused representation vector;

[0043] A neural network module that receives the fused representation vector and outputs an M-dimensional vector, where M is the number of categories;

[0044] A conversion module that converts the M-dimensional vector to obtain a probability vector P;

[0045] A multi-source data category determination module for performing S40-S42 of the first aspect to obtain multi-source data categories;

[0046] A data level calculation module for performing S50-S51 of the first aspect to obtain the maximum data level; and also for performing S52-S53 of the first aspect to obtain the expected data level.

[0047] Compared with the prior art, the present invention has the following effects:

[0048] 1. The multi-modal fusion-based unstructured data multi-classification and grading method provided by the present invention uses multi-modal fusion technology to combine the unstructured data content itself and its metadata, makes full use of the available information, can reduce the requirement for the number of samples, and improve the accuracy of unstructured data classification and grading.

[0049] 2. The multi-modal fusion-based unstructured data multi-classification and grading method provided by the present invention can simultaneously assign multiple categories to unstructured data, and obtain the maximum data level and the expected data level, and can perform more refined and flexible management and protection on the data.

[0050] 3. The multi-modal fusion-based unstructured data multi-classification and grading method provided by the present invention can flexibly select the maximum data level or the desired data level according to specific application scenarios, greatly enhancing the flexibility of data management and protection, and ensuring that organizations can effectively utilize existing resources while ensuring information security. The set maximum security level is designed to provide comprehensive data protection in extreme situations. For example, in actual operations, any change operation on unstructured data can be protected by mdl to prevent sensitive information from being illegally tampered with. During daily access, corresponding security measures can be implemented according to edl, thus maintaining the convenience of operations and work efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The drawings described herein are used to provide a further understanding of the present invention and form a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0052] Figure 1 is the schematic diagram of the multi-modal fusion-based unstructured data multi-classification and grading method in the embodiment of the present invention;

[0053] Figure 2 is the flow chart of the multi-modal fusion-based unstructured data multi-classification and grading method in the embodiment of the present invention;

[0054] Figure 3 is the flow chart of obtaining the fusion feature vector by multi-modal data fusion in the embodiment of the present invention;

[0055] Figure 4 is the schematic diagram of the multi-modal fusion model architecture of document data in the embodiment of the present invention;

[0056] Figure 5 is the schematic diagram of the multi-modal fusion model architecture of picture data in the embodiment of the present invention;

[0057] Figure 6 is the schematic diagram of the multi-modal fusion model architecture of audio data in the embodiment of the present invention;

[0058] Figure 7 is the schematic diagram of the multi-modal fusion model architecture of video data in the embodiment of the present invention;

[0059] Figure 8 is the flow chart of obtaining the multi-modal data multi-classification by the probability vector P in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0061] It should be noted that when an element is referred to as being "fixedly arranged on" or "arranged on" another element, it can be directly on the other element or indirectly on the other element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element.

[0062] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, the meaning of "a plurality of" is two or more unless otherwise specifically defined.

[0063] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "upper" and "lower" is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention.

[0064] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the term "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meaning of the above terms in the present invention can be understood according to specific circumstances.

[0065] Existing deep learning methods have problems such as insufficient information extraction ability, low accuracy, and poor flexibility when dealing with the classification and grading tasks of unstructured data. To solve these problems, the embodiments of the present invention propose a multi-modal fusion-based multi-classification and grading method and system for unstructured data, innovatively introducing multi-modal fusion technology, aiming to fully explore and utilize various effective information in unstructured data. At the same time, by adopting multi-classification and grading technology, multiple categories, maximum data levels, and expected data levels are assigned to unstructured data, which can significantly improve the classification and grading effect of unstructured data.

[0066] In a first aspect, the present invention provides a multi-modal fusion-based method for multi-classification and grading of unstructured data. Refer to Figures 1 to 2 , which includes the following steps:

[0067] S1. Fuse multi-modal data to obtain a fused feature vector;

[0068] As a possible implementation, refer to Figure 3 , S1 includes:

[0069] S10. Apply a semantic understanding model to represent the metadata of unstructured data to obtain a metadata feature vector; Exemplarily, the metadata at least includes file name, directory name, creator, descriptive information, data source, and creation date; Apply the RoBERTa semantic understanding model for representation.

[0070] S11. Apply a semantic understanding model to represent the unstructured data itself to obtain an unstructured data feature vector;

[0071] S12. Stack the metadata feature vector and the unstructured data feature vector in the channel dimension;

[0072] S13. Apply the convolution kernel of a convolutional neural network to fuse the data in the above multiple channels to obtain a fused feature vector.

[0073] In an actual application scenario, in addition to the data itself, unstructured data also covers a large amount of metadata that helps data classification, such as file name, directory name, creator, descriptive information, data source, and creation date. The semantic content contained in this information is combined with the unstructured data body, which can make full use of more context information, thereby improving the accuracy of data classification. The present invention uses multi-modal deep fusion technology to combine this information and then provides it to the classification model for prediction.

[0074] Metadata is natural language in text form, and unstructured data is data in different modalities. Information fusion cannot be directly performed. In the embodiments of the present invention, a deep neural network is used to perform deep representation on data in different modalities respectively, and then a convolutional neural network is used to fuse the feature vectors of different modalities. Metadata is usually short and has a high information content, and the RoBERTa semantic understanding model can be used for representation. The most crucial thing in representing metadata is how to organize the metadata. Combining the characteristics of the semantic understanding model, the present invention organizes cloud data in the following manner:

[0075] "File Name: '{file_name}' [sep] Directory Name: '{folder_name}' [sep] Creator: '{creator}' [sep] Description: '{description}' [sep] Data Source: '{data_provider}'".

[0076] Among them, [sep] is a special token used to prompt the AI model to distinguish different types of metadata. {file_name} is the file name of unstructured data, {folder_name} is the directory name where the unstructured data is located, {creator} is the creator of the unstructured data, {description} is the description or caption of the unstructured data, and {data_provider} is the data source. If the metadata of a certain part cannot be obtained, the corresponding value is directly set to empty. The advantage of adopting this method is that it is consistent with the RoBERTa pre-training method, which can maximize the utilization of the model's capabilities. In addition, the processing of metadata is relatively flexible. Even if some meta-information of certain unstructured data is missing, the AI model does not need to be changed and can still make classification predictions based on other meta-information. However, the more comprehensive the information provided, the more accurate the category prediction made by the model.

[0077] The metadata of all unstructured data is organized in this way, characterized by RoBERTa or other similar semantic understanding models, and then the characterization is fused with the information characterizations of other modalities through a convolutional neural network.

[0078] As an example, unstructured data includes document data, and document data includes formats such as txt, word, ppt, and pdf. Except for txt, documents in other formats not only contain text content but also many pictures. Therefore, document data has three modalities: metadata, text data, and pictures.

[0079] For the multi-modal fusion model architecture of document data, see Figure 4 , where the metadata is characterized by the RoBERTa semantic understanding model to obtain the metadata characterization vector; the text data is characterized by the Longformer semantic understanding model to obtain the text data characterization vector; for the pictures, each picture is first characterized by the ViT semantic understanding model, and then the characterization vectors of all pictures are fused using the Transformer encoder to obtain the picture characterization vector; the metadata characterization vector, text data characterization vector, and picture characterization vector are stacked in the channel dimension and then input into the convolutional neural network; the convolutional kernels of the convolutional neural network fuse the metadata characterization vector, text data characterization vector, and picture characterization vector on the three channels to obtain the fusion characterization of the document data.

[0080] As an example, unstructured data includes image data, which has two modalities: metadata and images. The multi-modal fusion model architecture of image data is shown in Figure 5 , where the metadata is characterized by a RoBERTa semantic understanding model to obtain a metadata characterization vector; for each image, the ViT semantic understanding model is first used to characterize the image, and then the Transformer encoder is used to fuse the characterization vectors of all images to obtain an image characterization vector; the metadata characterization vector and the image characterization vector are stacked in the channel dimension and then input into a convolutional neural network; the convolutional kernel of the convolutional neural network fuses the metadata characterization vector and the image characterization vector on the two channels to obtain a fused characterization of the image data.

[0081] As an example, unstructured data includes audio data, which has two modalities: metadata and audio signals; the multi-modal fusion model architecture of audio data is shown in Figure 6 , where the metadata is characterized by a RoBERTa semantic understanding model to obtain a metadata characterization vector; the audio signals are preprocessed to obtain mel-frequency cepstral coefficients and are characterized by a Conformer semantic understanding model to obtain an audio signal characterization vector; the metadata characterization vector and the audio signal characterization vector are stacked in the channel dimension and then input into a convolutional neural network; the convolutional kernel of the convolutional neural network fuses the metadata characterization vector and the audio signal characterization vector on the two channels to obtain a fused characterization of the audio data.

[0082] As an example, unstructured data includes video data, which has four modalities: metadata, subtitle sequences, audio signals, and frame sequences; the multi-modal fusion model architecture of video data is shown in Figure 7 , where the metadata is characterized by a RoBERTa semantic understanding model to obtain a metadata characterization vector; the subtitle sequences are characterized by a RoBERTa semantic understanding model to obtain a subtitle sequence characterization vector; the audio signals are characterized by a Conformer semantic understanding model to obtain an audio signal characterization vector; the frame sequences are characterized by a VTN-ENCODER semantic understanding model to obtain a frame sequence characterization vector; the metadata characterization vector, the subtitle sequence characterization vector, the audio signal characterization vector, and the frame sequence characterization vector are stacked in the channel dimension and then input into a convolutional neural network; the convolutional kernel of the convolutional neural network fuses the metadata characterization vector, the subtitle sequence characterization vector, the audio signal characterization vector, and the frame sequence characterization vector on the four channels to obtain a fused characterization of the video data.

[0083] S2. Input the fused feature vector into a neural network module to output an M-dimensional vector, where M is the number of categories; Exemplarily, input the fused feature vector into a neural network module for classification, which can be a multi-layer perceptron or a convolutional neural network, etc., and the embodiments of the present invention do not make specific limitations. What this neural network module outputs is an M-dimensional vector, called Logits, where M is the number of categories in the classification and grading task. Logits is represented by the following formula:

[0084]

[0085] S3. Transform the M-dimensional vector to obtain a probability vector P; Exemplarily, use the softmax function to transform the Logits vector to obtain a probability vector P, which is represented by the following formula:

[0086]

[0087] S4. Obtain the multi-modal data's multi-dimensional data categories based on the probability vector P; Exemplarily, P represents a probability distribution, and each probability value represents the probability that the input data belongs to the corresponding category, and the sum of all probabilities is 1. Because the information contained in unstructured data is complex and diverse, a file may simultaneously have multiple categories. Therefore, the embodiments of the present invention obtain the multi-modal data's multi-dimensional data categories based on the probability vector P. See Figure 8 , specifically including:

[0088] S40. Set parameters k and t, where k is the maximum value of the selectable categories and t is the probability threshold; Exemplarily, k can be set according to the actual business situation, and t can be obtained by statistically analyzing the prediction probabilities of the data categories in the training set and the test set based on the classification model.

[0089] S41. Screen out the categories from the probability vector P whose probability values are greater than or equal to the probability threshold t;

[0090] S42. Sort the screened-out categories in descending order of probability values. When the number of the screened-out categories is greater than or equal to k, take the top k categories as the multi-dimensional data categories of the unstructured data; when the number of the screened-out categories is less than k, then take all the categories as the multi-dimensional data categories of the unstructured data; Exemplarily, define DC as the multi-dimensional data categories of the unstructured data. Suppose m categories are finally assigned to an unstructured data, then DC can be expressed as:

[0091]

[0092] Correspondingly, the indices of the categories in DC in the probability vector P are expressed as:

[0093]

[0094] S5. Calculate the maximum data level and the expected data level of unstructured data;

[0095] Before performing multi - level classification on unstructured data to determine its maximum data level and expected data level, it is necessary to first calculate the corresponding level for each potential category. According to GB / T 43697 - 2024 "Data Security Technology - Data Classification and Grading Rules", the level of data is jointly determined by the data - affected object and data - grading elements (such as domain, group, region, precision, scale, depth, coverage, importance, etc.). These elements can all be obtained based on the data category and data attributes through certain rules. Therefore, it can be considered that there is a mapping relationship between the data level and the data category and data attributes. The present invention does not explore the specific mapping relationship, but defines it as a function.

[0096] As a possible implementation, define the relevant data attributes of unstructured data as attr, and for one of its data categories dc l , its corresponding data level dl l can be expressed as:

[0097] dl l = f(dcl, atte) (6)

[0098] Among them, the maximum data level is calculated by the following method:

[0099] S50. Calculate the data - level vector corresponding to each category in the multi - variable data category;

[0100] Exemplarily, according to dl l calculate the data - level vector DL corresponding to each category, and express it as:

[0101]

[0102] S51. Select the value with the maximum data level from the data - level vector as the maximum data level of the unstructured data;

[0103] Exemplarily, select the value with the maximum data level from DL as the maximum data level mdl of the unstructured data, and express it as:

[0104]

[0105] The expected data level is calculated by the following method:

[0106] S52. According to the index of each category in the multi - variable data category in the probability vector, obtain its corresponding logits from the M - dimensional vector, and further express it as CLogits;

[0107] Exemplarily, according to CIndex, obtain its corresponding logits from the Logits vector, denoted as CLogits:

[0108]

[0109] S53. Convert CLogits to a probability vector CP, and calculate the expected data level according to the probability vector CP.

[0110] Exemplarily, use the softmax function to convert CLogits to a probability vector CP, as shown in the following formula:

[0111]

[0112] As a possible implementation, the calculation formula for the expected data level is:

[0113]

[0114]

[0115] where edl is the expected data level; 1 ≤ l ≤ m, l is the l-th category in the multi-category of unstructured data, and m is the number of categories in the multi-category of unstructured data; cp l represents the probability of the data category dc l , idx l is the index of the l-th category in the multi-category of unstructured data in the probability vector P, is the logit corresponding to the l-th category in the multi-category of unstructured data, j is a variable representing 1, 2, 3, …, m, is the logit corresponding to the j-th category in the multi-category of unstructured data, the relevant data attribute of the unstructured data is attr, and for a data category dc l , its corresponding data level is dl l , dl l = f(dc l , attr).

[0116] So far, a maximum data level mdl and an expected data level edl have been set for each piece of unstructured data, and one of them can be flexibly selected according to specific application scenarios. This strategy not only enhances the flexibility of data management and protection but also ensures that the organization can effectively utilize existing resources while ensuring information security. The set maximum security level is designed to provide comprehensive data protection in extreme situations. For example, in actual operations, any change operation on unstructured data can use mdl to implement protective measures to prevent sensitive information from being illegally tampered with. During daily access, corresponding security measures can be implemented according to edl to maintain the convenience and work efficiency of operations.

[0117] In a second aspect, an embodiment of the present invention provides a multi-modal fusion-based multi-variate classification and grading system for unstructured data, including:

[0118] A multi-modal data fusion module that receives unstructured data, stacks the metadata of the unstructured data and itself along the channel dimension, and fuses the data in multiple channels by applying the convolutional kernels of a convolutional neural network to obtain a fused representation vector;

[0119] A neural network module that receives the fused representation vector and outputs an M-dimensional vector, where M is the number of categories;

[0120] A conversion module that converts the M-dimensional vector to obtain a probability vector P;

[0121] A multi-variate data category determination module for performing S40 - S42 of the first aspect to obtain multi-variate data categories;

[0122] A data level calculation module for performing S50 - S51 of the first aspect to obtain the maximum data level; and also for performing S52 - S53 of the first aspect to obtain the expected data level.

[0123] In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0124] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A multi-modal fusion-based multi-classification and grading method for unstructured data, characterized in that, It includes the following steps: S1. Multimodal data fusion to obtain a fused feature vector; S2. Input the fused feature vector into a neural network module to output an M-dimensional vector, where M is the number of categories; S3. Transform the M-dimensional vector to obtain a probability vector P; S4. Obtain the multi-modal data categories of the multimodal data based on the probability vector P, specifically including: S40. Set parameters k and t, where k is the maximum value of the selectable categories and t is the probability threshold; S41. Select the categories from the probability vector P whose probability values are greater than or equal to the probability threshold t; S42. Sort the selected categories in descending order of probability values. When the number of selected categories is greater than or equal to k, take the first k categories as the multi-modal data categories of the unstructured data; when the number of selected categories is less than k, take all the categories as the multi-modal data categories of the unstructured data; S5. Calculate the maximum data level and the expected data level of the unstructured data; among them, the maximum data level is calculated by the following method: S50. Calculate the data level vector corresponding to each category in the multi-modal data categories; S51. Select the maximum value of the data level from the data level vector as the maximum data level of the unstructured data; The expected data level is calculated by the following method: S52. According to the index of each category in the multi-modal data categories in the probability vector, obtain its corresponding logits from the M-dimensional vector, and further denote it as CLogits; S53. Convert CLogits to a probability vector CP, and calculate the expected data level according to the probability vector CP.

2. The multi-modal fusion-based unstructured data multi-classification and grading method according to claim 1, wherein The S1 includes: S10. Apply a semantic understanding model to represent the metadata of the unstructured data to obtain a metadata representation vector; S11. Apply a semantic understanding model to represent the unstructured data itself to obtain an unstructured data representation vector; S12. Stack the metadata representation vector and the unstructured data representation vector in the channel dimension; S13. Apply the convolutional kernels of a convolutional neural network to fuse the data in the above multiple channels to obtain a fused feature vector.

3. The multi-modal fusion-based unstructured data multi-classification and grading method according to claim 2, wherein, The metadata at least includes file name, directory name, creator, descriptive information, data source, and creation date; it is represented by the RoBERTa semantic understanding model.

4. The multi-modal fusion-based method for multi-classification and grading of unstructured data according to claim 3, wherein The unstructured data includes document data, and the document data has three modalities: metadata, text data, and pictures. Among them, the metadata is represented by the RoBERTa semantic understanding model to obtain a metadata representation vector; the text data is represented by the Longformer semantic understanding model to obtain a text data representation vector; each picture is first represented by the ViT semantic understanding model, and then the representation vectors of all pictures are fused by the Transformer encoder to obtain a picture representation vector; the metadata representation vector, the text data representation vector, and the picture representation vector are stacked in the channel dimension and then input into a convolutional neural network; the convolutional kernels of the convolutional neural network fuse the metadata representation vector, the text data representation vector, and the picture representation vector on the three channels to obtain the fused representation of the document data.

5. The multi-modal fusion-based unstructured data multi-classification and grading method according to claim 3, wherein The unstructured data includes image data, and the image data has two modalities of metadata and images. The metadata is characterized by a RoBERTa semantic understanding model to obtain a metadata characterization vector; each image is first characterized by a ViT semantic understanding model, and then the characterization vectors of all images are fused by a Transformer encoder to obtain an image characterization vector; the metadata characterization vector and the image characterization vector are stacked in the channel dimension and then input into a convolutional neural network; the convolutional kernels of the convolutional neural network fuse the metadata characterization vector and the image characterization vector on the two channels to obtain a fused characterization of the image data.

6. The multi-modal fusion-based multi-classification and multi-grading method for unstructured data according to claim 3, wherein, The unstructured data includes audio data, and the audio data has two modalities of metadata and audio signals; among them, the metadata is characterized by a RoBERTa semantic understanding model to obtain a metadata characterization vector; the audio signals are preprocessed to obtain mel-frequency cepstral coefficients and are characterized by a Conformer semantic understanding model to obtain an audio signal characterization vector; the metadata characterization vector and the audio signal characterization vector are stacked in the channel dimension and then input into a convolutional neural network; the convolutional kernels of the convolutional neural network fuse the metadata characterization vector and the audio signal characterization vector on the two channels to obtain a fused characterization of the audio data.

7. The multi-modal fusion-based method for multi-category and multi-level classification of unstructured data according to claim 3, wherein The unstructured data includes video data, and the video data has four modalities of metadata, subtitle sequences, audio signals, and frame sequences; among them, the metadata is characterized by a RoBERTa semantic understanding model to obtain a metadata characterization vector; the subtitle sequences are characterized by a RoBERTa semantic understanding model to obtain subtitle sequence characterization vectors; the audio signals are characterized by a Conformer semantic understanding model to obtain audio signal characterization vectors; the frame sequences are characterized by a VTN-ENCODER semantic understanding model to obtain frame sequence characterization vectors; the metadata characterization vector, the subtitle sequence characterization vector, the audio signal characterization vector, and the frame sequence characterization vector are stacked in the channel dimension and then input into a convolutional neural network; the convolutional kernels of the convolutional neural network fuse the metadata characterization vector, the subtitle sequence characterization vector, the audio signal characterization vector, and the frame sequence characterization vector on the four channels to obtain a fused characterization of the video data.

8. The multi-modal fusion-based multi-classification and grading method for unstructured data according to claim 1, wherein, The calculation formula for the desired data level is: Among them, edl is the expected data level; 1 ≤ l ≤ m, where l is the l-th category in the multi-dimensional data categories of unstructured data, and m is the number of categories in the multi-dimensional data categories of unstructured data; cp l represents the probability of the data category dc l cp l idx l is the index of the l-th category in the multi-dimensional data categories of unstructured data in the probability vector P, is the logit corresponding to the l-th category in the multi-dimensional data categories of unstructured data, and j is a variable representing 1, 2, 3, …, m, is the logit corresponding to the j-th category in the multi-dimensional data categories of unstructured data. The relevant data attribute of the unstructured data is attr. For a data category dc l , its corresponding data level is dl l dl l = f(dc l , attr).

9. The multi-modal fusion-based multi-classification and grading method for unstructured data according to claim 1, wherein Define the relevant data attribute of unstructured data as attr, and for one of its data categories dc l , its corresponding data level dl l can be expressed as: dl l = f(dc l , attr), According to dc l Calculate the data level vector DL corresponding to each category and represent it as: Select the maximum value of the data level from DL as the maximum data level mdl of the unstructured data, denoted as:

10. A multi-modal fusion unstructured data multi-classification and grading system, characterized in that, including: A multi-modal data fusion module that receives unstructured data, characterizes the metadata and itself of the unstructured data and stacks them in the channel dimension, and applies the convolutional kernels of a convolutional neural network to fuse the data in multiple channels to obtain a fused characterization vector; A neural network module that receives the fused characterization vector and outputs an M-dimensional vector, where M is the number of categories; A conversion module that converts the M-dimensional vector to obtain a probability vector P; A multi-source data category determination module for performing S40 to S42 described in claim 1 to obtain a multi-source data category; The data level calculation module is used to execute S50 - S51 described in claim 1 to obtain the maximum data level; it is also used to execute S52 - S53 described in claim 1 to obtain the expected data level.

Citation Information

Patent Citations

  • Sensitive information discovery and automatic classification and grading method based on multi-modal fusion

    CN116049397A

  • Classification and grading method and device for unstructured documents, equipment and medium

    CN117290758A

  • Cloud edge collaboration and AI fusion data classification and grading method and system and storage medium

    CN119557493A

  • Multi-modal information fused corn disease intelligent grading evaluation and treatment recommendation method

    CN119625530A

  • Methods and apparatuses for training service model and determining text classification category

    US11216620B1