A multi-modal information fusion-based electronic archive image multi-level classification method and device
Patent Information
- Application Number
- CN202410016368.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-04
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2044-01-04
AI Technical Summary
然而,这些传统方法缺乏可伸缩性和通用性,难以适应不断增长和多样化的电子档案需求
[0062]This invention classifies electronic archives by integrating text and image features, overcoming the shortcomings of single-modality classification and improving the accuracy and robustness of electronic archive classification. Furthermore, it utilizes knowledge graph ontology for electronic archive classification, enabling multi-level classification and increasing the refinement of the classification results, thereby solving the problem of insufficient classification refinement in existing related technologies.
Smart Images

Figure CN117951092B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to a method and device for multi-level classification of electronic archive images based on multimodal information fusion. Background Technology
[0002] Society has now fully entered the digital age, and archival management is gradually shifting from traditional paper-based document management to electronic document management. However, in the process of archiving electronic records, the sheer volume of electronic records makes manual classification difficult and time-consuming, severely impacting the efficiency of electronic record archiving.
[0003] Current approaches to classifying electronic records primarily include using rule-based classification or employing a single modality (considering only visual features or relying solely on textual information). However, these traditional methods lack scalability and universality, making it difficult to adapt to the ever-growing and diverse needs of electronic records.
[0004] In other words, existing technologies for classifying electronic records can usually only provide classification results for a single category, and cannot achieve multi-level classification with a low degree of detail. Summary of the Invention
[0005] To address the aforementioned problems in the existing technology, this invention provides a method and device for multi-level classification of electronic archival images based on multimodal information fusion.
[0006] The technical problem to be solved by this invention is achieved through the following technical solution:
[0007] This invention provides a multi-level classification method for electronic archival images based on multimodal information fusion, comprising:
[0008] Acquire images of electronic archives to be classified;
[0009] Calculate the text information entropy and image information entropy of the electronic archive images to be classified, respectively;
[0010] The trained text classification model, image classification model, and multimodal information fusion classification model are used to classify the electronic archive images to be classified, respectively, to obtain text classification result sets, image classification result sets, and fusion classification result sets. The trained text classification model, image classification model, and multimodal information fusion classification model are trained using different training sets, which are constructed based on a pre-built multi-level category knowledge graph ontology of electronic archives within a predetermined domain.
[0011] The predicted category of the electronic archive image to be classified is determined based on the text information entropy, the image information entropy, the text classification result set, the image classification result set, and the fusion classification result set.
[0012] Based on the multi-level category knowledge graph ontology and the predicted category, the multi-level classification result of the electronic archive image to be classified is determined.
[0013] In some embodiments, determining the multi-level classification result of the electronic archive image to be classified based on the multi-level category knowledge graph ontology and the predicted category includes:
[0014] From the multi-level category knowledge graph ontology, find the leaf node where the predicted category is located;
[0015] The sequence of all nodes traversed from the root node to the leaf node of the multi-level category knowledge graph ontology is used as the multi-level classification result of the electronic archive image to be classified.
[0016] In some embodiments, determining the predicted category of the electronic archive image to be classified based on the text information entropy, the image information entropy, the text classification result set, the image classification result set, and the fused classification result set includes:
[0017] Adjust the text information entropy and the image information entropy to the same order of magnitude;
[0018] The first fusion coefficient, the second fusion coefficient, and the third fusion coefficient are calculated based on the text information entropy and image information entropy of the same order of magnitude.
[0019] The text classification result set, the image classification result set, and the fused classification result set are fused using the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient to obtain the prediction result classification set;
[0020] The category with the highest probability in the predicted result classification set is taken as the predicted category of the electronic archive image to be classified.
[0021] In some embodiments, the text classification result set includes: n different categories and a first probability for each of the n different categories; the image classification result set includes: the n different categories and a second probability for each of the n different categories; the fusion classification result set includes: the n different categories and a third probability for each of the n different categories;
[0022] The step of fusing the text classification result set, the image classification result set, and the fused classification result set using the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient to obtain a predicted result classification set includes:
[0023] Calculate the product between the first probability of the i-th category and the second fusion coefficient to obtain the first product value; i = 1, 2, ..., n, where n is an integer greater than or equal to 2;
[0024] Calculate the product between the second probability of the i-th category and the third fusion coefficient to obtain the second product value;
[0025] Calculate the product between the third probability of the i-th category and the first fusion coefficient to obtain the third product value;
[0026] The sum of the first product value, the second product value, and the third product value is taken as the fourth probability of the i-th category;
[0027] The set consisting of the n different categories and the fourth probability of each of the n different categories is used as the prediction result classification set.
[0028] In some embodiments, calculating the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient based on text information entropy and image information entropy of the same order of magnitude includes:
[0029] The reciprocal of the absolute value of the difference between text information entropy and image information entropy of the same order of magnitude is used as the first fusion coefficient;
[0030] The second fusion coefficient is calculated based on the sum of text information entropy and image information entropy of the same order of magnitude, the first fusion coefficient, and the text information entropy.
[0031] The third fusion coefficient is calculated based on the sum of text information entropy and image information entropy of the same order of magnitude, the first fusion coefficient, and the image information entropy.
[0032] In some embodiments, the different training sets include: a text feature training set, an image feature training set, and a multimodal information fusion feature training set;
[0033] Before classifying the electronic archive images to be classified using the trained text classification model, trained image classification model, and trained multimodal information fusion classification model, respectively, and obtaining the text classification result set, image classification result set, and fusion classification result set, the process includes:
[0034] Construct a multi-level category knowledge graph ontology of electronic archives in the preset domain;
[0035] Acquire multiple electronic archive images belonging to the preset field and having real categories;
[0036] Based on the multiple electronic archive images with real categories and the multi-level category knowledge graph ontology, the text feature training set, the image feature training set, and the multimodal information fusion feature training set are constructed respectively.
[0037] The initial text classification model is trained using the text feature training set to obtain the trained text classification model;
[0038] The initial image classification model is trained using the image feature training set to obtain the trained image classification model;
[0039] The initial multimodal information fusion classification model is trained using the multimodal information fusion feature training set to obtain the trained multimodal information fusion classification model.
[0040] In some embodiments, constructing the text feature training set, the image feature training set, and the multimodal information fusion feature training set based on the plurality of electronic archival images with real categories and the multi-level category knowledge graph ontology includes:
[0041] The multiple electronic archival images with real categories are preprocessed to obtain multiple preprocessed electronic archival images;
[0042] Based on the true category of each preprocessed electronic archive image, the label of the preprocessed electronic archive image is marked as a leaf node in the multi-level category knowledge graph ontology;
[0043] A subset of the labeled preprocessed electronic archive images is used as the original training set.
[0044] Text features are extracted from each electronic archive image in the original training set, and the text feature training set is constructed based on the extracted text features and the labels of each electronic archive image.
[0045] Image features are extracted from each electronic archive image in the original training set, and the image feature training set is constructed based on the extracted image features and the labels of each electronic archive image.
[0046] The text features of each electronic archive image in the original training set are concatenated and concatenated with the image features of the electronic archive image column by column to obtain the multimodal information fusion features of the electronic archive image. Based on the multimodal information fusion features of each electronic archive image and the labels of each electronic archive image, the training set of the multimodal information fusion features is constructed.
[0047] In some embodiments, a method for calculating the text information entropy of the electronic archive image to be classified includes:
[0048] Extract the text from the electronic archive image to be classified;
[0049] Calculate the total number of characters in the text, and count the frequency of each character in the text;
[0050] Calculate the probability of each character in the text based on the total number of characters and the frequency of each character's occurrence;
[0051] The text information entropy of the electronic archive image to be classified is calculated based on the probability of all characters in the text.
[0052] In some embodiments, a method for calculating the image information entropy of the electronic archive image to be classified includes:
[0053] Convert the electronic archive image to be classified into a grayscale image;
[0054] Convert the grayscale image into a NumPy array;
[0055] Using the NumPy array, the frequency of each pixel in the grayscale image is counted, and the total number of pixels in the grayscale image is calculated.
[0056] Calculate the probability of each pixel value in the grayscale image based on the total number of pixels and the frequency of each pixel's occurrence;
[0057] The image information entropy of the electronic archive image to be classified is calculated based on the probability of all pixel values in the grayscale image.
[0058] The present invention also provides a multi-level classification device for electronic archive images based on multimodal information fusion, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0059] The memory is used to store computer programs;
[0060] When the processor executes the program stored in the memory, it implements the steps of the above-mentioned multi-level classification method for electronic archive images based on multimodal information fusion.
[0061] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0062] This invention classifies electronic archives by integrating text and image features, overcoming the shortcomings of single-modality classification and improving the accuracy and robustness of electronic archive classification. Furthermore, it utilizes knowledge graph ontology for electronic archive classification, enabling multi-level classification and increasing the refinement of the classification results, thereby solving the problem of insufficient classification refinement in existing related technologies.
[0063] The present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description
[0064] Figure 1 This is a flowchart illustrating the multi-level classification method for electronic archive images based on multimodal information fusion provided in an embodiment of the present invention.
[0065] Figure 2 This is a schematic diagram of the structure of an exemplary knowledge graph ontology provided in an embodiment of the present invention;
[0066] Figure 3 This is an implementation block diagram of an exemplary method for multi-level classification of electronic archive images based on multimodal information fusion provided in this invention.
[0067] Figure 4 This is a model training block diagram of an exemplary multi-level classification method for electronic archive images based on multimodal information fusion provided in an embodiment of the present invention. Detailed Implementation
[0068] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.
[0069] Figure 1 This is a flowchart illustrating a multi-level classification method for electronic archival images based on multimodal information fusion, provided by an embodiment of the present invention. The method includes:
[0070] S101. Obtain the image of the electronic archive to be classified.
[0071] Here, electronic archive images refer to electronic archives in image form.
[0072] S102. Calculate the text information entropy and image information entropy of the electronic archive images to be classified, respectively.
[0073] Entropy is an indicator of information content in information theory. In this invention, entropy value is used as an indicator of text information content. Generally, a higher entropy value indicates that the text is rich in information, possibly containing more different characters or symbols, or that the characters in the text are more evenly distributed. Conversely, a lower entropy value indicates that the text contains relatively less information, possibly containing fewer characters or symbols, or that the characters in the text are more biased towards certain specific characters.
[0074] In some embodiments, the method for calculating the textual information entropy of an electronic archive image to be classified is as follows:
[0075] S1. Extract the text from the electronic archive image to be classified.
[0076] For example, OCR recognition methods can be used to extract the text from the electronic archive image to be classified.
[0077] S2. Calculate the total number of characters m1 in the text, and count the frequency s of each character in the text. j1 .
[0078] S3. Based on the total number of characters m1 and the frequency s of each character. j1 Calculate the probability of each character in the text.
[0079] S4. Based on the probability of all characters in the text, calculate the text information entropy H1 of the electronic archive image to be classified.
[0080] Specifically, the expression for the text information entropy H1 of the electronic archive image to be classified is:
[0081]
[0082] Similar to text information entropy, the higher the entropy value of an image, the more uniform the pixel distribution in the image and the greater the amount of information. Conversely, a lower entropy value indicates that the pixel distribution in the image is uneven and the amount of information is less.
[0083] In some embodiments, the method for calculating the image information entropy of an electronic archive image to be classified is as follows:
[0084] S11. Convert the image of the electronic archive to be classified into a grayscale image.
[0085] S12. Convert the grayscale image to a NumPy array.
[0086] S13. Using a NumPy array, count the frequency s of each pixel in the grayscale image. j2 Calculate the total number of pixels m2 in the grayscale image.
[0087] S14. Based on the total number of pixels m2 and the frequency of each pixel s j2 Calculate the probability of each pixel value in a grayscale image.
[0088] S15. Based on the probability of all pixel values in the grayscale image, calculate the image information entropy H2 of the electronic archive image to be classified.
[0089] Specifically, the expression for the image information entropy H2 of the electronic archive image to be classified is:
[0090]
[0091] S103. Using a pre-trained text classification model, a pre-trained image classification model, and a pre-trained multimodal information fusion classification model, classify the images of the electronic archives to be classified, respectively, and obtain text classification result sets, image classification result sets, and fusion classification result sets; wherein, the pre-trained text classification model, the pre-trained image classification model, and the pre-trained multimodal information fusion classification model are trained using different training sets; the different training sets are constructed based on a pre-built multi-level category knowledge graph ontology of electronic archives in a preset domain.
[0092] Here, the different training sets include: a text feature training set, an image feature training set, and a multimodal information fusion feature training set. The text feature training set is used to train the text classification model, the image feature training set is used to train the image classification model, and the multimodal information fusion feature training set is used to train the multimodal information fusion classification model.
[0093] Here, the text classification model can extract text features from the OCR recognition results and classify the results based on the extracted features. The input to the text classification model is an electronic archival image, and the output is a set consisting of categories and their corresponding probabilities. The model framework of the text classification model can adopt any of the following network frameworks: Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Recursive Neural Networks (RNN), Transformer, Multilayer Perceptron (MLP), BERT (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), XLNet, and ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately).
[0094] Here, the image classification model can extract image features from electronic archival images and classify them based on the extracted features. The input to the image classification model is the electronic archival image, and the output is a set consisting of categories and their corresponding probabilities. The model framework of the image classification model can adopt any of the following network frameworks: Convolutional Neural Networks (CNN), LeNet, AlexNet, VGG (Visual Geometry Group), GoogLeNet, ResNet (Residual Neural Network), DenseNet, MobileNet, and EfficientNet.
[0095] Here, the multimodal information fusion classification model can extract image features from electronic archival images and classify them based on the extracted features. The input to the multimodal information fusion classification model is the electronic archival image, and the output is a set of categories and their corresponding probabilities. The model framework of the multimodal information fusion classification model can adopt any of the following network frameworks: Bidirectional Long Short-Term Memory Network (BiLSTM), Bidirectional Gated Recurrent Unit (BiGRU), Bidirectional Recurrent Neural Network (BiRNN), Transformer Encoder, BERT (Bidirectional Encoder Representations from Transformers), XLNet, and GPT (Generative Pre-trained Transformer).
[0096] Here, the multi-level category knowledge graph ontology includes root nodes, child nodes, and leaf nodes, and each child node or leaf node represents a category. When constructing a multi-level category knowledge graph ontology for electronic archives in a specific domain, multiple entities (i.e., multiple concepts) of the electronic archives in that domain can be defined, as can the hierarchy between concepts and the attributes of the concepts. Once the attributes and concepts are defined, the multi-level category knowledge graph ontology for the electronic archives in that domain can be constructed based on these definitions. For example, Figure 2 This is a schematic diagram of the structure of a multi-level category knowledge graph ontology for electronic archives in the field of computer technology. Figure 2 The knowledge graph ontology shown does not contain attributes, but only multiple concepts, such as "leave request," "news," "academic literature," and "project proposal." Figure 2 In this knowledge graph ontology, "Computer Technology" is the root node; "Leave Request", "News", "Academic Literature", "Project Proposal", and "Transaction Record" are first-level child nodes; "Journal", "Paper", "Patent", and "Research Notes" are second-level child nodes, and "Journal", "Paper", "Patent", and "Research Notes" are subclasses of "Academic Literature"; "Research Paper", "Review Paper", "Investigation Paper", and "Applied Paper" are subclasses of "Paper", and "Research Paper", "Review Paper", "Investigation Paper", and "Applied Paper" are the four leaf nodes of this knowledge graph ontology.
[0097] Here, the text classification result set includes: n distinct categories and the first probability of each of the n distinct categories, that is, the text classification result set is X={(t1,x1),(t2,x2),…,(t n ,x nThe image classification result set includes: n distinct categories and the second probability of each of the n distinct categories, that is, the image classification result set is Y={(t1,y1),(t2,y2),…,(t n ,y n The fusion classification result set includes: n distinct categories and the third probability of each of the n distinct categories, that is, the fusion classification result set is Z={(t1,z1),(t2,z2),…,(t n ,z n )}。 t i Let i represent the i-th category, i = 1, 2, ..., n, where n is an integer greater than or equal to 2.
[0098] S104. Based on the text information entropy, image information entropy, text classification result set, image classification result set, and fusion classification result set, determine the predicted category of the electronic archive image to be classified.
[0099] Specifically, S104 can be achieved through the following steps:
[0100] S1041. Adjust the text information entropy and image information entropy to the same order of magnitude.
[0101] For example, the text information entropy and image information entropy can be normalized first. Then, the normalized image information entropy can be multiplied by an adjustment factor. In this way, text information entropy H'1 and image information entropy H'2 of the same order of magnitude can be obtained. For example, the adjustment factor can be 1000.
[0102] S1042. Calculate the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient based on the text information entropy and image information entropy of the same order of magnitude.
[0103] Specifically, the reciprocal of the absolute value of the difference between text information entropy H'1 and image information entropy H'2, which belong to the same order of magnitude, is used as the first fusion coefficient c; the second fusion coefficient a is calculated based on the sum of text information entropy H'1 and image information entropy H'2, which belong to the same order of magnitude, the first fusion coefficient c, and text information entropy H'1; and the third fusion coefficient b is calculated based on the sum of text information entropy H'1 and image information entropy H'2, which belong to the same order of magnitude, the first fusion coefficient c, and image information entropy H'2.
[0104] Specifically, the expression for the first fusion coefficient c is: The expression for the second fusion coefficient 'a' is: The expression for the third fusion coefficient b is:
[0105]
[0106] S1043. Using the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient, the text classification result set, the image classification result set, and the fused classification result set are fused to obtain the predicted result classification set.
[0107] Specifically, for the i-th category, calculate the first probability x of the i-th category. i The product of the first product and the second fusion coefficient 'a' is used to obtain the first product value; the second probability y of the i-th category is calculated. i The product of the first product and the third fusion coefficient b yields the second product value; the third probability z of the i-th category is then calculated. i The product of the first fusion coefficient c and the third product value is obtained; the first product value a·x i The second product value b·y i and the third product value c·z i The sum of these probabilities serves as the fourth probability l for the i-th category. i That is, l i =a·x i +b·y i +c·z i The set consisting of n distinct categories and the fourth probability of each of the n distinct categories is taken as the prediction result classification set, that is, the prediction result classification set is L={(t1,l1),(t2,l2),…,(t n ,l n )}.
[0108] S1044. The category with the highest probability in the prediction result classification set shall be used as the predicted category of the electronic archive image to be classified.
[0109] Specifically, let L = {(t1,l1),(t2,l2),…,(t n ,l n In the array, the largest element l k Corresponding category t k As the predicted category for the electronic archive image to be classified.
[0110] S105. Based on the multi-level category knowledge graph ontology and the predicted category, determine the multi-level classification result of the electronic archive image to be classified.
[0111] Specifically, from the knowledge graph ontology, a leaf node containing the predicted category is found; the sequence of all nodes traversed from the root node of the knowledge graph ontology to that leaf node is used as the multi-level classification result for the electronic archive image to be classified. For example, for the above... Figure 2 In this case, when the predicted category of an electronic archive image to be classified is "review paper", the multi-level classification result of the electronic archive image to be classified is: {computer technology, academic literature, paper, review paper}.
[0112] For example, Figure 3 This is a block diagram illustrating an implementation of the multi-level classification method for electronic archival images based on multimodal information fusion provided by the present invention. Figure 3 As shown, the electronic archive image to be classified is input into the trained text classification model, the trained image classification model, and the trained multimodal information fusion classification model, respectively, to obtain the classification result sets of the three models, namely, the text classification result set, the image classification result set, and the fusion classification result set. Then, based on the text information entropy and image information entropy of the electronic archive image to be classified, the text classification result set, the image classification result set, and the fusion classification result set are fused and classified to obtain the predicted result classification set. Based on the predicted result classification set, the predicted category of the electronic archive image to be classified is obtained. Then, based on the predicted category and the pre-constructed multi-level category knowledge graph, the multi-level classification result of the electronic archive image to be classified, i.e., the multi-level category, is obtained.
[0113] In some embodiments, prior to S103, the method further includes the following steps:
[0114] S201. Construct a multi-level category knowledge graph ontology of electronic archives in a preset domain.
[0115] Here, the preset field can be arbitrarily selected, and the present invention does not limit it.
[0116] For example, the knowledge graph ontology modeling tool Protege can be used to model the knowledge graph ontology, the knowledge graph ontology description language (OWL) can be used to encode the knowledge graph ontology, and the embedded inference engine of the open-source semantic web application framework Jena can be used to perform logical checks on the OWL encoding results of the knowledge graph ontology, such as hierarchical reasoning and missing class completion. In this way, the hierarchical relationships between concepts in the knowledge graph ontology can be guaranteed to be correct and the relationship chains between concepts can be complete.
[0117] S202. Obtain multiple electronic archive images belonging to the preset field and having real categories.
[0118] Here, electronic archival images of multiple known real categories within the preset domain can be collected.
[0119] S203. Based on multiple electronic archive images with real categories and multi-level category knowledge graph ontology, construct text feature training sets, image feature training sets, and multimodal information fusion feature training sets respectively.
[0120] Specifically, S203 is achieved through the following steps:
[0121] S21. Preprocess multiple electronic archive images with real categories to obtain multiple preprocessed electronic archive images.
[0122] Specifically, each collected electronic archival image with a real category can be preprocessed, such as image enhancement, resizing, and denoising, to improve image quality and processability. After preprocessing, multiple preprocessed electronic archival images are obtained.
[0123] S22. Based on the true category of each preprocessed electronic archive image, label the preprocessed electronic archive image as a leaf node in the knowledge graph ontology.
[0124] For example, for each preprocessed electronic archival image, a label can be assigned to the preprocessed electronic archival image based on its true category through manual annotation. This label can then be set as a leaf node in the knowledge graph ontology constructed above. In this way, the set of categories represented by the leaf nodes in the constructed knowledge graph ontology can be used as the label set for the electronic archival image samples.
[0125] Here, automatic annotation can also be used to set labels for preprocessed electronic archive images to improve annotation efficiency.
[0126] S23. Use a portion of the preprocessed electronic archive images from multiple preprocessed electronic archive images with labels as the original training set.
[0127] For example, 70% of the pre-processed electronic archival images with labels can be used as the original training set, while the remaining 30% can be used as the original test set. Furthermore, the classification and feature distributions of the original training set and the original test set are ensured to be similar during the partitioning process.
[0128] S24. Extract the text features of each electronic archive image in the original training set, and construct a text feature training set based on the extracted text features and the labels of each electronic archive image.
[0129] Specifically, we can first perform OCR recognition on each electronic archive image in the original training set to extract the text from each electronic archive image. Then, for the text extracted from each electronic archive image, we can perform data cleaning, word segmentation, removal of stop words, and part-of-speech tagging. After that, we can use text feature extraction algorithms, including TF-IDF and Word2Vec, to convert the text into vector representation and capture the semantic and contextual information of words to obtain the text features corresponding to the electronic archive image. Then, we can use the label of the electronic archive image as the label of the text features corresponding to the electronic archive image. In this way, we can construct a text feature training set.
[0130] S25. Extract the image features of each electronic archive image in the original training set, and construct an image feature training set based on the extracted image features and the labels of each electronic archive image.
[0131] Specifically, the original training set of electronic archive images can be preprocessed by image enhancement, resizing, and denoising. For example, the resize() function built into the open-source computer vision library OpenCV can be used to uniformly set the size of each electronic archive image to 224*224 and give each electronic archive image RGB color three channels. In this way, the preprocessed electronic archive images can be obtained. Next, for each preprocessed electronic archive image, a pre-trained model (for example, using VGG19 as the pre-trained model, whose structure includes 16 convolutional layers, 5 max pooling layers and 3 fully connected layers, and setting the weight parameter Weights of the VGG19 model to Imagenet) is used to extract the image features of the preprocessed electronic archive image to obtain the image features corresponding to the preprocessed electronic archive image. Then, the label of the preprocessed electronic archive image is used as the label of the image features corresponding to the preprocessed electronic archive image. In this way, the image feature training set is constructed.
[0132] S26. Concatenate the text features of each electronic archive image in the original training set with the image features of the electronic archive image column by column to obtain the multimodal information fusion features of the electronic archive image. Based on the multimodal information fusion features of each electronic archive image and the labels of each electronic archive image, construct a multimodal information fusion feature training set.
[0133] Specifically, for each electronic archive image in the original training set, the text features corresponding to the electronic archive image and the image features corresponding to the electronic archive image can be concatenated and merged into a higher-dimensional feature vector to serve as the multimodal information fusion feature corresponding to the electronic archive image. Then, the label of the electronic archive image is used as the label of the multimodal information fusion feature corresponding to the electronic archive image. In this way, the multimodal information fusion feature training set is constructed.
[0134] S204. The initial text classification model is trained using the text feature training set to obtain a trained text classification model.
[0135] Here, the initial text classification model includes the text classification model framework and a set of weight parameters obtained during initialization. For example, the initial text classification model can be a BERT-based text classification model, and the loss function used when training the initial text classification model can be cross-entropy loss.
[0136] S205. Train the initial image classification model using the image feature training set to obtain a trained image classification model.
[0137] Here, the initial image classification model includes the image classification model framework and a set of weight parameters obtained during initialization. For example, the initial image classification model can be an image classification model based on a VGG19 convolutional neural network, and the loss function used when training the initial image classification model can be cross-entropy loss.
[0138] S206. The initial multimodal information fusion classification model is trained using the multimodal information fusion feature training set to obtain the trained multimodal information fusion classification model.
[0139] Here, the initial multimodal information fusion classification model includes the multimodal information fusion classification model framework and a set of weight parameters obtained during initialization. For example, the initial multimodal information fusion classification model can be a BiLSTM-based multimodal information fusion classification model, and the loss function used when training the initial multimodal information fusion classification model can be cross-entropy loss.
[0140] For example, Figure 4 This is a model training flowchart for the multi-level classification method of electronic archival images based on multimodal information fusion provided by this invention. Figure 4 As shown, the TF-IDF and Word2Vec algorithms are used to process the electronic archive image dataset (i.e., the original training set mentioned above), and a text feature training set is constructed based on the processing results. A pre-trained VGG19 model is used to process the electronic archive image dataset, and an image feature training set is constructed based on the processing results. Next, a multimodal information fusion feature training set is constructed based on the text feature training set and the image feature training set. Then, the initial text classification model is trained using the text feature training set to obtain a trained text classification model; the initial image classification model is trained using the image feature training set to obtain a trained image classification model; and the initial multimodal information fusion classification model is trained using the multimodal information fusion feature training set to obtain a trained multimodal information fusion classification model.
[0141] This invention classifies electronic archives by integrating text and image features, overcoming the shortcomings of single-modality classification and improving the accuracy and robustness of electronic archive classification. Furthermore, it utilizes knowledge graph ontology for electronic archive classification, enabling multi-level classification and increasing the refinement of the classification results, thereby solving the problem of insufficient classification refinement in existing related technologies.
[0142] The present invention also provides a multi-level classification device for electronic archival images based on multimodal information fusion, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to execute the program stored in the memory to implement the steps of the above-mentioned multi-level classification method for electronic archival images based on multimodal information fusion.
[0143] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0144] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0145] In this specification, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. While different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce a good effect.
[0146] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. A multi-level classification method for electronic archival images based on multimodal information fusion, characterized in that, include: Acquire images of electronic archives to be classified; Calculate the text information entropy and image information entropy of the electronic archive images to be classified, respectively; The trained text classification model, image classification model, and multimodal information fusion classification model are used to classify the electronic archive images to be classified, respectively, to obtain text classification result sets, image classification result sets, and fusion classification result sets. The trained text classification model, image classification model, and multimodal information fusion classification model are trained using different training sets, which are constructed based on a pre-built multi-level category knowledge graph ontology of electronic archives within a predetermined domain. The predicted category of the electronic archive image to be classified is determined based on the text information entropy, the image information entropy, the text classification result set, the image classification result set, and the fusion classification result set. Based on the multi-level category knowledge graph ontology and the predicted category, the multi-level classification result of the electronic archive image to be classified is determined; The text classification result set includes: Different categories and the The image classification result set includes: the first probability of each of the different categories; Different categories and the The second probability of each of the different categories; the fusion classification result set includes: the Different categories and the The third probability of each of the three different categories; The step of determining the predicted category of the electronic archive image to be classified based on the text information entropy, the image information entropy, the text classification result set, the image classification result set, and the fused classification result set includes: Adjust the text information entropy and the image information entropy to the same order of magnitude; The first fusion coefficient, the second fusion coefficient, and the third fusion coefficient are calculated based on the text information entropy and image information entropy of the same order of magnitude. Calculate the first The first product value is obtained by multiplying the first probability of each category by the second fusion coefficient. , It is an integer greater than or equal to 2; Calculate the first The product of the second probability of each category and the third fusion coefficient is used to obtain the second product value; Calculate the first The product of the third probability of each category and the first fusion coefficient is used to obtain the third product value; The sum of the first product value, the second product value, and the third product value is taken as the first product value. The fourth probability of each category; The Different categories and the The set of the fourth probability of each of the different categories is used as the classification set of the prediction results; The category with the highest probability in the predicted result classification set is taken as the predicted category of the electronic archive image to be classified.
2. The method for multi-level classification of electronic archival images based on multimodal information fusion according to claim 1, characterized in that, The step of determining the multi-level classification result of the electronic archive image to be classified based on the multi-level category knowledge graph ontology and the predicted category includes: From the multi-level category knowledge graph ontology, find the leaf node where the predicted category is located; The sequence of all nodes traversed from the root node to the leaf node of the multi-level category knowledge graph ontology is used as the multi-level classification result of the electronic archive image to be classified.
3. The multi-level classification method for electronic archive images based on multimodal information fusion according to claim 1, characterized in that, The calculation of the first fusion coefficient, the second fusion coefficient, and the third fusion coefficient based on text information entropy and image information entropy of the same order of magnitude includes: The reciprocal of the absolute value of the difference between text information entropy and image information entropy of the same order of magnitude is used as the first fusion coefficient; The second fusion coefficient is calculated based on the sum of text information entropy and image information entropy of the same order of magnitude, the first fusion coefficient, and the text information entropy. The third fusion coefficient is calculated based on the sum of text information entropy and image information entropy of the same order of magnitude, the first fusion coefficient, and the image information entropy.
4. The method for multi-level classification of electronic archival images based on multimodal information fusion according to claim 1, characterized in that, The different training sets include: text feature training set, image feature training set, and multimodal information fusion feature training set; Before classifying the electronic archive images to be classified using the trained text classification model, trained image classification model, and trained multimodal information fusion classification model, respectively, and obtaining the text classification result set, image classification result set, and fusion classification result set, the process includes: Construct a multi-level category knowledge graph ontology of electronic archives in the preset domain; Acquire multiple electronic archive images belonging to the preset field and having real categories; Based on the multiple electronic archive images with real categories and the multi-level category knowledge graph ontology, the text feature training set, the image feature training set, and the multimodal information fusion feature training set are constructed respectively. The initial text classification model is trained using the text feature training set to obtain the trained text classification model; The initial image classification model is trained using the image feature training set to obtain the trained image classification model; The initial multimodal information fusion classification model is trained using the multimodal information fusion feature training set to obtain the trained multimodal information fusion classification model.
5. The multi-level classification method for electronic archive images based on multimodal information fusion according to claim 4, characterized in that, The step of constructing the text feature training set, the image feature training set, and the multimodal information fusion feature training set based on the multiple electronic archive images with real categories and the multi-level category knowledge graph ontology includes: The multiple electronic archival images with real categories are preprocessed to obtain multiple preprocessed electronic archival images; Based on the true category of each preprocessed electronic archive image, the label of the preprocessed electronic archive image is marked as a leaf node in the multi-level category knowledge graph ontology; A subset of the labeled preprocessed electronic archive images is used as the original training set. Text features are extracted from each electronic archive image in the original training set, and the text feature training set is constructed based on the extracted text features and the labels of each electronic archive image. Image features are extracted from each electronic archive image in the original training set, and the image feature training set is constructed based on the extracted image features and the labels of each electronic archive image. The text features of each electronic archive image in the original training set are concatenated and concatenated with the image features of the electronic archive image column by column to obtain the multimodal information fusion features of the electronic archive image. Based on the multimodal information fusion features of each electronic archive image and the labels of each electronic archive image, the training set of the multimodal information fusion features is constructed.
6. The multi-level classification method for electronic archival images based on multimodal information fusion according to claim 1, characterized in that, A method for calculating the text information entropy of the electronic archive image to be classified includes: Extract the text from the electronic archive image to be classified; Calculate the total number of characters in the text, and count the frequency of each character in the text; Calculate the probability of each character in the text based on the total number of characters and the frequency of each character's occurrence; The text information entropy of the electronic archive image to be classified is calculated based on the probability of all characters in the text.
7. The method for multi-level classification of electronic archival images based on multimodal information fusion according to claim 1, characterized in that, A method for calculating the image information entropy of the electronic archive image to be classified includes: Convert the electronic archive image to be classified into a grayscale image; Convert the grayscale image into a NumPy array; Using the NumPy array, the frequency of each pixel in the grayscale image is counted, and the total number of pixels in the grayscale image is calculated. Based on the total number of pixels and the frequency of each pixel's occurrence, the probability of each pixel value in the grayscale image is calculated; The image information entropy of the electronic archive image to be classified is calculated based on the probability of all pixel values in the grayscale image.
8. A multi-level classification device for electronic archive images based on multimodal information fusion, comprising: It includes a processor, a communication interface, a memory, and a communication bus, characterized in that the processor, the communication interface, and the memory communicate with each other through the communication bus; The memory is used to store computer programs; When the processor executes the program stored in the memory, it implements the steps of the method described in any one of claims 1-7.
Citation Information
Patent Citations
Article classification method and device based on model fusion
CN113934843A
Document picture classification method and device, storage medium and electronic equipment
CN114780773A
Multi-source heterogeneous disaster data fusion understanding method
CN117312548A