Multi-mode intelligent AI classification system and method for archive arrangement
By designing a multimodal intelligent AI classification system for archive sorting, using technical means such as dual-channel generators and PPO algorithms, the problem of traditional methods being difficult to deal with multimodal data is solved, and efficient multimodal data classification and processing is achieved, which significantly improves classification accuracy and system adaptability.
Patent Information
- Application Number
- CN202510652914.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-21
AI Technical Summary
Traditional single-modal processing methods are difficult to effectively utilize multimodal information, and their processing capabilities for incomplete data are limited, which makes it difficult to dynamically select processing paths based on data quality and equipment resources during multimodal data classification processing to ensure classification accuracy and efficiency.
Design a multi-modal intelligent AI classification system for archive sorting, including data acquisition and transmission module, perception layer module, scoring original multi-modal data based on modal quality evaluator, feature extraction and optimization module, classification and verification module. The cross-modal consistency loss function and structured constraint loss are introduced through a dual-channel generator, and the associated image features or pseudo-text descriptions are generated, dynamic routing decisions are made using the PPO algorithm, and closed-loop optimization of quality evaluation is introduced.
Effectively handle common noise, fuzzy, non-standard formats and other problems in archival data, significantly improve classification accuracy and robustness, improve computing efficiency, enhance system adaptability, and reduce the need for manual intervention.
Smart Images

Figure CN120182989A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and more specifically, to a multimodal intelligent AI classification system and method for archive sorting. Background Art
[0002] Archival data usually contains two modes: images (such as scans) and text (such as electronic documents), and the data quality varies (such as blurry images and non-standard text layout). Traditional single-modality processing methods are difficult to effectively utilize multimodal information, and have limited processing capabilities for incomplete data (such as isolated modality data).
[0003] Therefore, when classifying and processing multimodal data, it is necessary to run it on different devices (such as edge devices and high-performance servers). Due to the differences in modalities, the transmission method, resource usage, and processing methods required during operation are all different. Considering the high efficiency of processing, sub-modal routing is a feasible processing method, but how to dynamically select the processing path based on data quality and device resources while ensuring classification accuracy and efficiency is a complex issue. In addition, when classifying modal data, content recognition must be performed first, which involves the conversion of data content or format. Different modalities have different conversion methods, and there may be cross-modal inconsistencies (such as mismatches between images and text descriptions) or low quality (such as unreasonable pseudo-text generation). Therefore, ensuring conversion quality is also a key challenge.
[0004] In view of this, the present invention proposes a multimodal intelligent AI classification system and method for archive organization to solve the above problems. Summary of the invention
[0005] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned purpose, the present invention provides the following technical solutions: A multimodal intelligent AI classification system for archive sorting, comprising a data acquisition and transmission module: obtaining a scan of a paper archive through a scanner or a camera, and combining it with an electronic document to form original multimodal data; Perception layer module: used to receive raw multimodal data, introduce cross-modal consistency loss function and structured constraint loss in the deployed dual-channel generator, generate associated image features or pseudo-text descriptions, and form evolving multimodal data; Based on the modal quality assessor, the original multimodal data and the evolved multimodal data are scored for mutuality; through the conflict identification and resolution mechanism, the enhanced multimodal data is output through cyclic optimization; Feature extraction and optimization module: Extract features based on enhanced multimodal data to obtain quality score triples; use the PPO algorithm to make dynamic routing decisions, introduce closed-loop optimization of quality assessment, optimize and update quality score triples, and output decision actions; execute path rules according to decision actions and output comprehensive features; Classification and verification module: construct a multimodal feature matrix based on comprehensive features, design a multimodal classifier based on Attention-Transformer, introduce a modal weight adaptive adjustment module, set the multimodal feature matrix as input, and output the classification result; add a classification result verification mechanism to form a closed-loop optimization; obtain the final archive classification result.
[0006] Preferably, the enhanced multimodal data is obtained by: Obtain scanned copies of paper archives through scanners or cameras, combine them with electronic documents to form original multimodal data, and then transmit them to the perception layer through the API interface; A dual-channel generator is deployed in the perception layer. The dual-channel generator learns the cross-modal mapping relationship between images and texts through adversarial training. A cross-modal consistency loss function is introduced in the adversarial training, and a structured constraint loss is added. Differentiated quality thresholds are preset for archive types under different modalities, and associated image features are generated based on text, or pseudo-text descriptions are generated based on images to obtain evolving multimodal data. Design a modality quality evaluator to score the co-generacy of original multimodal data and evolved multimodal data; the multi-dimensional evaluation indicators of the modality quality evaluator include image quality or text quality, and cross-modal consistency; Introducing conflict identification and resolution mechanisms to output optimal evolving multimodal data; The output is enhanced multimodal data, including original multimodal data, evolved multimodal data and symbiosis scores.
[0007] Preferably, the dual-channel generator comprises: For image modality data, we deploy Beter-DeblurGAN-v3 to automatically repair images and extract local key areas. Beter-DeblurGAN-v3 includes: adding multi-scale convolution kernels to the encoder of DeblurGAN-v3 to capture image interference of different scales, and adding RPN region proposal network to automatically locate and enhance key areas. For text data, the dynamic blocking strategy of the VS-ADSP document structure parser is used to identify non-standard typesetting, and the SB-GNN graph neural network is used to reconstruct the document logical topology to convert non-standard typesetting documents into standard versions; among them, the VS-ADSP document structure parser includes the introduction of document segmentation evolution based on visual-semantic union in ADSP, and the SB-GNN graph neural network includes adding visual features and semantic features to the node features of GNN.
[0008] For isolated modal data, a conditional variational autoencoder (CVAE) is introduced into the cross-modal generative adversarial network Cross-GAN, and a variational constraint is added to the generator of Cross-GAN. Among them, for extremely isolated modal data, visual question answering enhancement (VQA-E) is introduced, and a semantic description of the image is generated through a pre-trained VQA-E model.
[0009] Preferably, the method for introducing a conflict recognition and resolution mechanism includes: Pre-define a threshold for cross-modal consistency. If the cross-modal consistency value output by the modal quality evaluator is less than the threshold, it is marked as a conflict. For the original multi-modal data and evolved multi-modal data with conflicts, a dual-channel generator is used to generate multiple possible evolved multi-modal data, and then the modal quality evaluator is used to re-evaluate the cross-modal consistency until the conflict is resolved.
[0010] Preferably, the method for obtaining the comprehensive features includes: Perform preliminary extraction of modal features on the enhanced multi-modal data to obtain a quality score triple (s_img, s_txt, s_consist). Use the Proximal Policy Optimization (PPO) algorithm to make dynamic routing decisions. Set the input as the quality score triple (s_img, s_txt, s_consist), and introduce a closed-loop optimization of quality evaluation, including: if the quality score of the modal data is lower than the preset threshold, feedback it to the dual-channel generator to trigger the re-generation of the data, perform modal feature extraction again based on the re-generated modal data, and update the quality score triple until the closed-loop optimization of quality evaluation is completed, and then obtain the final quality score triple (s_img, s_txt, s_consist). Import the final quality score triple into the PPO algorithm to output the final decision action. According to the decision action execution path rule, output the comprehensive features.
[0011] Preferably, the method for performing preliminary extraction of modal features on the enhanced multi-modal data to obtain a quality score triple (s_img, s_txt, s_consist) includes: Step Q1: Detect the resolution of the image data in the evolved multi-modal data, select an appropriate feature extraction model according to the resolution. The features include layout, seal position, signature area confidence, and preliminary OCR recognition results, and extract the visual feature vector f_img. Extract rules for the text data in the evolved multi-modal data, extract key fields, and use them as auxiliary inputs to DistilBERT. The features include keyword frequency, named entities, and paragraph topic distribution, and extract the semantic feature vector f_txt. Step Q2: Calculate the OCR recognition accuracy based on the edit distance and semantic rationality. Introduce multi-scale analysis in the Laplacian variance and SSIM structural integrity to calculate the blur degree Blur and SSIM value. Use YOLOv8 to detect the seal area, and record the highest confidence score as the seal detection confidence. Output the image quality score s_img through the calculation formula; the calculation formula is , is the regularization coefficient, is the variance of each index or other regularization terms; Calculate the semantic coherence score based on NextSentencePrediction of DistilBERT. Use SB-GNN to analyze the document layout and evaluate the structural integrity; output the text quality score s_txt through the calculation formula; the calculation formula is ; Step Q3: Calculate the similarity between the OCR result of the image and the title of the electronic text. Use the CLIP model to calculate the image-text matching degree of the image and text, and output the cross-modal consistency score s_consist. The calculation formula is: s_consist = · Similarity + · CLIP image-text matching degree; Among them, , , , , , , and are weight coefficients, and the preset weights are matched based on the archive metadata; Step Q4: Output the quality score triple (s_img, s_txt, s_consist).
[0012] Preferably, the method of using the PPO algorithm for dynamic routing decision includes: Define the state space S = [s_img, s_txt, s_consist, CPU load, remaining power, network bandwidth, storage capacity], and define the action space A = {pure image path, pure text path, multi-modal fusion path}; Define the reward function , and use the initial quality score as the initial reward for reinforcement learning; Dynamically adjust the hyperparameters for edge devices and high-performance servers, introduce a long-term reward mechanism, and use discounted return to calculate the long-term performance; The designed routing rules are soft constraints, and paths are selected through probability distribution, including: If s_img > 0.9 and s_consist > 0.8, then select the pure image path with a 90% probability and the multimodal fusion path with a 10% probability; If s_txt > 0.9 and s_consist > 0.8, then select the pure text path with a 90% probability and the multimodal fusion path with a 10% probability; If s_img < 0.4, then switch to the text path with an 80% probability and select the multimodal fusion path with a 20% probability; If s_txt < 0.4, then switch to the image path with an 80% probability and select the multimodal fusion path with a 20% probability; If 0.4 ≤ s_img, s_txt ≤ 0.9 or s_consist < 0.6, then select the multimodal fusion path with a 70% probability and the unimodal path with a 30% probability; The output is a decision-making action a ∈ A.
[0013] Preferably, the execution method of the path rules includes: For the pure image path, use a deep CNN to extract visual features and generate text features by combining the preliminary OCR recognition results; For the text modality, use a deep NLP model to extract semantic features and generate structured features by combining the output results of DistilBERT; For the multimodal fusion path, use a cross-modal SP-Transformer model to fuse the structured and text features of the original multimodal data and the evolved multimodal data to generate a modal embedding vector; The cross-modal SP-Transformer model is based on a two-stream Transformer framework, introduces MT-ARNet to process fuzzy features, and adds NSR to intervene in the conflicting data in the original multimodal data and the evolved multimodal data; Record the text features, structured features, and embedding vectors as comprehensive features.
[0014] Preferably, the method for obtaining the final file classification result includes: Step C1: Construct a multimodal feature matrix based on the comprehensive features, where the multimodal feature matrix forms a unified feature representation by normalizing and aligning the dimensions of different modal features; Step C2: Design a multimodal classifier based on a multi-layer perceptron fusion network of Attention-Transformer, and introduce a modal weight adaptive adjustment module in the attention mechanism of Transformer to dynamically adjust the weights of image, text, and structured features according to the modal importance of different file types; Step C3: For the archive classification task, a set of archive type labels is predefined, and a multimodal classifier is pre-trained using a supervised learning method. The training goal is to minimize the classification cross entropy loss, and a regularization term is introduced at the same time; Step C4: Input the multimodal feature matrix into the trained multimodal classifier and output the classification probability distribution of the archive; select the archive type label with the highest confidence as the final classification result through the probability distribution; Step C5: Design a classification result verification mechanism. For archival data whose classification confidence is lower than the preset threshold, trigger the manual review process and feed the manual review results back to the dual-channel generator to form a closed-loop optimization. Step C6: Output the final archive classification results, including the archive type label, classification confidence and corresponding multimodal feature matrix, for subsequent archive management and retrieval.
[0015] The present invention also discloses a multi-modal intelligent AI classification method for archive sorting, comprising: Step S1: Obtain a scanned copy of a paper file through a scanner or camera, combine it with the electronic document to form original multimodal data, and transmit it to the perception layer through an API interface; Step S2: In the perception layer, a dual-channel generator is used to generate multimodal data through adversarial training, a cross-modal consistency loss function and a structured constraint loss are introduced, and associated image features or pseudo-text descriptions are generated based on a preset differential quality threshold; Step S3: Design a modality quality evaluator to score the original multimodal data and the generated multimodal data based on image quality, text quality and cross-modal consistency indicators; through the conflict identification and resolution mechanism, mark the data below the cross-modal consistency threshold, generate multiple possible multimodal data and re-evaluate, and output the enhanced multimodal data; Step S4: extract modal features from the enhanced multimodal data to obtain a quality score triplet (s_img, s_txt, s_consist); use the PPO algorithm to make dynamic routing decisions, set the state space, action space and reward function, introduce closed-loop optimization, and trigger data regeneration if the quality score is lower than the threshold, update the quality score triplet and output the decision action; Step S5: Execute path rules according to the decision action, including pure image path, pure text path and multimodal fusion path, use deep CNN, NLP model or cross-modal SP-Transformer model to extract features and output comprehensive features; Step S6: Construct a multi-modal feature matrix based on the comprehensive features, and form a unified feature representation through normalization and dimension alignment; design a multi-modal classifier based on Attention-Transformer, introduce a modal weight adaptive adjustment module, and use supervised learning for pre-training; Step S7: Input the multi-modal feature matrix into the classifier, output the classification probability distribution of the archives, and select the label with the highest confidence as the classification result; for the archive data with a confidence lower than the threshold, trigger manual review and feedback for optimization; Step S8: Output the final archive classification result, including the type label, classification confidence, and multi-modal feature matrix.
[0016] Technical effects and advantages of a multi-modal intelligent AI classification system for archive sorting according to the present invention: 1. Through the enhancement, quality evaluation, dynamic routing, and closed-loop optimization of multi-modal data, it can effectively handle common problems such as noise, blur, and non-standard formats in archive data, significantly improving the classification accuracy and robustness.
[0017] 2. The dynamic routing decision of the PPO algorithm and the modal weight adaptive adjustment module enable the system to dynamically adjust the processing strategy according to data quality, device resources, and task requirements, which not only improves the computing efficiency but also enhances the self-adaptability of the system.
[0018] 3. Through multi-level closed-loop optimization (data enhancement, feature extraction, classification verification), high levels of intelligence and automation are achieved, reducing the need for manual intervention. The introduction of the classification result verification mechanism ensures the accuracy of low-confidence data, and at the same time, the model is continuously optimized through manual feedback, taking into account both efficiency and accuracy.
[0019] 4. Through technologies such as cross-modal consistency loss, cross-modal generative adversarial network (Cross-GAN), and CLIP model, deep fusion and complementarity of multi-modal data such as images and texts are achieved. This cross-modal ability significantly improves the comprehensiveness and accuracy of archive classification.
[0020] In summary, it can significantly improve the efficiency of archive management, support the digitization of old archives, and provide strong support for intelligent retrieval and cross-domain applications. Brief Description of the Drawings
[0021] Figure 1 It is a schematic structural diagram of a multi-modal intelligent AI classification system for archive sorting according to the present invention; Figure 2 It is a schematic diagram of the steps of a multi-modal intelligent AI classification method for archive sorting according to the present invention. Detailed Embodiments
[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] Embodiment 1
[0024] Please refer to Figure 1 and Figure 2 As shown, a multi-modal intelligent AI classification system for file sorting in this embodiment includes: File data usually includes various modalities such as paper scans and electronic documents, with large differences in data formats and qualities. How to uniformly process and enhance these heterogeneous data is a difficult problem. It is difficult to accurately model the semantic association between images and texts, especially in file data, where images may contain problems such as blurriness and damage, and texts may have non-standard layouts. How to generate high-quality cross-modal data is a challenge. How to design a mechanism that can comprehensively evaluate the quality of multi-modal data to ensure that the generated data meets high-quality standards in terms of images, texts, and cross-modal consistency. During the generation of multi-modal data, situations where the image and text descriptions are inconsistent may occur. How to identify and resolve these conflicts is a key issue.
[0025] To address the above problems, this design learns the cross-modal mapping relationship between images and texts through adversarial training, introduces a cross-modal consistency loss function and a structured constraint loss, sets different quality thresholds for different modalities, and solves the complexity problem of cross-modal mapping. Design multi-dimensional evaluation indicators (image quality, text quality, cross-modal consistency) to provide an objective quality scoring mechanism. Through conflict identification and resolution, ensure that the output enhanced multi-modal data is optimal and solve the problem of processing conflicting data. The specific content includes: Obtain paper file scans (image data, such as JPG, PNG) through a scanner or camera, combine with electronic documents (text data, such as PDF, Word) to form original multi-modal data, and then transmit it to the perception layer through an API interface; In the perception layer, a dual-channel generator is deployed. The dual-channel generator learns the cross-modal mapping relationship between images and texts through adversarial training to ensure the authenticity and consistency of the generated data. A cross-modal consistency loss function is introduced in the adversarial training to enforce the alignment of the generated text with the original image in the semantic space (e.g., calculating the similarity between images and texts through the CLIP model); for special archives such as engineering drawings, a structured constraint loss is added (e.g., the generated CAD description text needs to conform to the SVG vector graphic element syntax). Different quality thresholds are preset for different types of archives in different modalities (e.g., for medical archives, the NER accuracy rate ≥ 90%, and for engineering drawings, the SSIM ≥ 0.85). Associated image features are generated based on the text (e.g., generating abstract symbols according to "circuit diagram"), or pseudo-text descriptions are generated based on the image (e.g., using the VQA model to explain the content of the drawing); for example, for an archive with only scanned drawings, the CVAE can generate multiple possible pseudo-text descriptions (such as "circuit diagram, designed in 2020" or "architectural drawing, floor plan"), and the discriminator is used to select the optimal description. Generate multi-modal data; (for example, for a scanned copy of a blurred contract signature page, the generator not only completes the text but also infers the position of the signer (such as "Technical Director") based on the context and generates the corresponding virtual signature image). The dual-channel generator specifically includes: There may be problems such as blurring and damage in the scanned copies of archives. The automatic repair and extraction of key areas (such as seals and signatures) are difficult points. The text in the archives may be in non-standard layout (such as handwritten and disordered). Parsing and reconstructing the logical structure of the document is a challenge. In some cases, there may be only single-modal data (such as only images or only texts), and high-quality pseudo-modal data (such as generating text descriptions based on images) needs to be generated.
[0026] To address the above problems, this design uses the improved Beter-DeblurGAN-v3 to capture image interference at different scales through multi-scale convolutional kernels and combines the RPN region proposal network to automatically locate key areas, solving the problems of image blurring and key information extraction. Through the VS-ADSP document structure parser (introducing visual-semantic joint segmentation) and the SB-GNN graph neural network (fusing visual and semantic features), the logical reconstruction of non-standard layout documents is achieved, solving the problem of text non-standardness. The conditional variational autoencoder (CVAE) is introduced in Cross-GAN to add variational constraints, and combined with visual question answering enhancement (VQA-E) to generate semantic descriptions, solving the problem of generating pseudo-modal data from single-modal data. The specific content includes: For image modality data, the Beter-DeblurGAN-v3 (an enhanced network obtained by further optimizing the lightweight scanning quality enhancement network based on U-Net++) is deployed to automatically repair the image (e.g., repair blurring, tilting, and shadow interference in the image) and extract local key regions (such as signatures and seals); among them, Beter-DeblurGAN-v3 includes: adding multi-scale convolutional kernels (such as 3x3, 5x5, 7x7) to the encoder of DeblurGAN-v3 (a lightweight scanning quality enhancement network based on U-Net++) to capture image interferences at different scales (such as large-area shadows and small-scale creases), and adding an RPN region proposal network, combined with the structure of Faster R-CNN, to automatically locate and enhance key regions; When adding the RPN region proposal network, end-to-end training of the RPN is required, taking into account the joint optimization of object existence judgment and position regression. The specific process includes: The loss function of the RPN can be defined as: ; where, is the classification loss (judging whether the region is a key region), is the probability that the th anchor box predicted by the model is a key region (foreground); is the true label (1 represents a positive sample, 0 represents a negative sample), and the binary cross-entropy loss is adopted, is the number of anchor boxes in the mini-batch (usually 256); is the regression loss (locating the region boundary), is the boundary offset of the th positive sample anchor box predicted by the model, is the true boundary offset (calculated from the GT box and the anchor box), the smooth L1 loss is adopted, and ; is the number of positive sample anchor boxes (usually 128); is the balance coefficient, used to adjust the weights of the classification and regression tasks, usually set to = 10 (the amplitude of the regression loss is small and needs to be amplified).
[0027] In image modality processing, by improving DeblurGAN-v3 and adding multi-scale convolutional kernels and an RPN region proposal network, the system can more accurately repair blurred images and extract key region features, significantly enhancing the usability of image data.
[0028] For text data, the dynamic chunking strategy of the VS-ADSP document structure parser is used to identify non-standard layouts (such as multi-column mixed layouts and nested tables), and the SB-GNN graph neural network is combined to reconstruct the document logical topology, converting non-standard layout documents into standard versions. Among them, the VS-ADSP document structure parser includes introducing Vision-Semantic Joint Segmentation (VSJS) into ADSP, that is, ADSP combines image layout information (such as the visual segmentation lines of scanned documents) and text semantic information (such as the semantic boundaries of titles and texts) to improve the parsing accuracy of complex layouts (such as multi-column mixed layouts and nested tables). For example, for a multi-column mixed medical report, VS-ADSP can accurately separate the two columns of "diagnosis results" and "test data". The SB-GNN graph neural network includes adding visual features (such as the geometric information of layout bounding boxes) and semantic features (such as word embeddings extracted by BERT) to the node features of GNN.
[0029] The node update formula of GNN can be defined as: ; by assigning independent weights to self-loops and neighbors and , or sharing weights, while retaining the node's own features and neighbor features, enhancing the model's expressive ability. Among them, represents the feature vector of node at the th layer, , represents the feature vector of node at the th layer, represents the normalization coefficient, usually related to the node degree or edge weight, represents the set of neighbor nodes of node , represents the non-linear activation function (such as ReLU, Sigmoid).
[0030] In text modality processing, by introducing vision-semantic joint document segmentation and graph neural networks, the system can effectively parse non-standard layout documents and reconstruct the logical topology, greatly improving the structured processing ability of text data.
[0031] For isolated modal data (such as only scanned drawings), the conditional variational autoencoder CVAE (Conditional Variational Autoencoder) is introduced in the cross-modal generative adversarial network Cross-GAN, and variational constraints are added to the generator of Cross-GAN to ensure the diversity and authenticity of the generated data. For extremely isolated modal data (such as historical photos without any text labels), visual question answering enhancement VQA-E (Visual Question Answering Enhancement) is introduced to generate semantic descriptions of images through pre-trained VQA-E models (such as CLIP-ViT).
[0032] The loss function of Cross-GAN after introducing conditional variational autoencoder CVAE can be optimized as: ; represents the KL divergence term of the variational constraint, is the balance coefficient, represents a generator, represents the discriminator.
[0033] Through the dual-channel generator of the perception layer module, the system introduces cross-modal consistency loss function and structured constraint loss in adversarial training, which can effectively generate associated image features or pseudo-text descriptions. This method significantly improves the quality and consistency of multimodal data and solves the problem of information loss or incompleteness in traditional single-modal processing.
[0034] For isolated modal data, by introducing conditional variational autoencoder (CVAE) and visual question answering-enhanced (VQA-E), the system can generate high-quality cross-modal descriptions to make up for the shortcomings of single modal data.
[0035] Design a modality quality evaluator (based on deep learning classifier) to score the original multimodal data and the generated multimodal data; for example, evaluate the accuracy of the text repaired by Beter-DeblurGAN-v3 or the authenticity of the generated image, and the scoring results are used for the weight allocation of subsequent modality fusion. Among them, the multi-dimensional evaluation indicators of the modality quality evaluator include image quality (such as PSNR (peak signal-to-noise ratio), SSIM (structural similarity), seal detection confidence) or text quality (such as BERT semantic coherence score, named entity recognition (NER) accuracy), cross-modal consistency (such as calculating the image-text matching degree through the CLIP model); In archival data, there may be semantic deviations between images and texts, making it difficult to accurately determine whether there is consistency between image and text descriptions. When conflicts are found, it is difficult to ensure the quality of the final output data.
[0036] To address the above issues, this design accurately identifies conflicting data by pre-defining cross-modal consistency thresholds and combining them with the scores of the modal quality assessor. A dual-channel generator is used to generate a variety of possible evolution data, and the modal quality assessor is used to iteratively evaluate until the conflict is resolved, thereby solving the problem of efficient conflict handling. The specific contents include: Introducing conflict identification and resolution mechanisms to output optimal generated multimodal data (i.e., closest to the original multimodal data); Methods for introducing conflict identification and resolution mechanisms include: A threshold for cross-modal consistency is predefined (e.g., dynamically adjusted based on archival metadata or historical data). If the cross-modal consistency value output by the modality quality assessor is less than the threshold, it is marked as a conflict; For conflicting original multimodal data and generated multimodal data, a dual-channel generator is used to generate multiple possible generated multimodal data, and then a modality quality evaluator is used to re-evaluate the cross-modal consistency until the conflict is resolved.
[0037] The output is enhanced multimodal data (image + text), including original multimodal data, generated multimodal data and co-genesis score.
[0038] The introduction of the modality quality assessor uses multi-dimensional indicators (such as image quality, text quality, and cross-modal consistency) to score the original and evolved data, and combines conflict identification and resolution mechanisms to form a closed-loop optimization. This mechanism ensures the reliability and robustness of the output data, and is particularly suitable for processing ambiguous, damaged, or non-standard format data commonly found in archives.
[0039] Archival images may have problems such as low resolution, blur, and unclear seal areas. Archival texts may have problems such as semantic incoherence and chaotic layout. It is difficult to accurately measure the semantic consistency between images and texts for multimodal data. In view of this, this design combines resolution detection, multi-scale fuzziness analysis, SSIM structural integrity, seal detection confidence and other indicators to calculate image quality scores through weighted formulas to solve the diversity problem of image quality assessment. The semantic coherence analysis of DistilBERT and the layout structure evaluation of SB-GNN are combined with weighted formulas to calculate text quality scores to solve the complexity problem of text quality assessment. The similarity between OCR and text titles and the image-text matching of the CLIP model are used to calculate cross-modal consistency scores through weighted formulas to solve the accuracy problem of consistency assessment.
[0040] Perform preliminary modal feature extraction on the enhanced multimodal data to obtain the quality score triplet (s_img, s_txt, s_consist); Although lightweight models such as MobileNetV3 and DistilBERT are used for feature extraction, the computational overhead of feature extraction can still be relatively high for large-scale archival data. Especially in scenarios with high real-time requirements, the efficiency of feature extraction directly affects the latency of subsequent quality assessment and routing decisions, which may lead to a decline in the overall system performance. Therefore, adaptive improvements are made to the lightweight models. The specific content includes: Step Q1: Detect the resolution of the image data (the image repaired by Beter-DeblurGAN-v3) in the generated multimodal data, and select an appropriate feature extraction model according to the resolution. For example, for low-resolution images, use a lighter model (such as MobileNetV3-Small, dynamically select the Small or Large version), while for high-resolution or complex images, use a more powerful model (such as EfficientNet-B0). The features include layout, seal position, signature area confidence, and preliminary OCR recognition results, and extract the visual feature vector f_img; to improve the accuracy of feature extraction.
[0041] Extract rules from the text data (the standard version text parsed by VS-ADSP and SB-GNN) in the generated multimodal data, extract key fields (such as dates, names), and use them as auxiliary inputs to DistilBERT to reduce the dependence on DistilBERT, thereby improving efficiency. Combine the rule extraction results and DistilBERT. The features include keyword frequencies (such as "confidential", "contract"), named entities (such as names, dates), and paragraph topic distributions, and extract the semantic feature vector f_txt; to improve the structural degree of feature extraction.
[0042] Inaccurate quality assessment may lead to incorrect routing decisions. For example, misjudging low-quality modal data as high-quality, thus selecting the wrong processing path.
[0043] Step Q2: Calculate the OCR recognition accuracy based on the edit distance and semantic rationality. Introduce multi-scale analysis in the Laplacian variance and SSIM structural integrity to calculate the blur Blur and SSIM values, and detect the blur / skew degree through the calculation formula to improve the robustness to images with complex backgrounds. Use YOLOv8 to detect the seal area, and take the highest confidence score as the seal detection confidence. The larger the value, the clearer the seal and the more compliant the position is with the document specification. The output range is [0,1]; output the image quality score s_img through the calculation formula. The calculation formula is , is the regularization coefficient, is the variance or other regularization term of each index; Among them, in the process of calculating the OCR recognition accuracy based on the edit distance and semantic rationality, if there is a truly correct text, it can be directly used for calculation. In the case of no true label, in addition to using BERT to calculate the semantic perplexity (PPL), a pre-trained semantic similarity model (such as SimCSE) can also be considered to further improve the evaluation accuracy of semantic rationality. The OCR recognition accuracy is mapped to the [0, 1] interval.
[0044] The Laplacian variance introduces multi-scale analysis to more comprehensively evaluate the clarity of the image. The closer the Blur value of the blur is to 1, the clearer the image (high variance), and it automatically falls into the [0, 1] interval; the SSIM structure is complete for measuring the similarity of the layout structure between the image and the standard template, and multi-scale analysis is introduced to better capture the details and structural information of the image. The SSIM value is linearly mapped to [0, 1].
[0045] Calculate the semantic coherence score based on NextSentencePrediction of DistilBERT, use SB-GNN to analyze the document layout, and evaluate the structural integrity (such as the missing rate of required fields); output the text quality score s_txt through the calculation formula; the calculation formula is ; For example, Scenario 1: Blurry medical report scan Input: Blurry X-ray report (partial error in OCR recognition, clear seal).
[0046] Index calculation: OCR = 0.65 ("pneumonia" misrecognized as "lung award"); Blur = 0.3, SSIM = 0.4 (layout distortion); = 0.9; Weight selection: Preset weight for medical records (0.4, 0.2, 0.4); Comprehensive score: s_img = 0.686; Scenario 2: High-definition engineering drawing Input: Clear scanned CAD drawing (no seal, accurate OCR).
[0047] Index calculation: OCR = 0.95; Blur = 0.9, SSIM = 0.85 (conforms to the template); = 0.0; Weight selection: Preset weight for engineering drawings (0.1, 0.6, 0.3); Comprehensive score: s_img == 0.626.
[0048] Step Q3: Calculate the similarity (cosine similarity) between the OCR result of the image and the title of the electronic text, and use the CLIP model to calculate the image-text matching degree of the image and the text as a supplement to the consistency score. Output the cross-modal consistency score s_consist, and the calculation formula is: s_consist = · Similarity + · CLIP image-text matching degree; Among them, , , , , , , and are weight coefficients, which are matched with preset weights based on archive metadata (such as file name suffix, upload tags); they can also be dynamically adjusted based on gradient descent or reinforcement learning.
[0049] Step Q4: Output the quality score triple (s_img, s_txt, s_consist).
[0050] Use the PPO algorithm for dynamic routing decision-making to improve training stability. Set the input as the quality score triple (s_img, s_txt, s_consist), and introduce the closed-loop optimization of quality assessment, including: if the quality score of the modal data is lower than the preset threshold, it is fed back to the dual-channel generator to trigger the re-generation of the data, and the modal feature extraction is performed again based on the re-generated modal data to update the quality score triple until the closed-loop optimization of quality assessment is completed, and the final quality score triple (s_img, s_txt, s_consist) is obtained; Through the closed-loop optimization of quality assessment, the system can re-generate and extract features from low-quality data, continuously update the quality score triple, and ensure that the output comprehensive features have high confidence and consistency.
[0051] If the dynamic modal routing scheme detects that the quality of a certain modal data is low (such as s_img < 0.4 or s_txt < 0.4), it is fed back to the original scheme to trigger further image repair (such as calling Beter-DeblurGAN-v3) or text generation (such as calling Cross-GAN and CVAE).
[0052] Use the final quality score triple to import into the PPO algorithm and output the final decision action; According to the decision action execution path rule, output the comprehensive feature.
[0053] In multi-modal data processing, there are significant differences in the quality of different modalities. Given this, how to dynamically select the optimal processing path (such as pure image, pure text, or multi-modal fusion) based on data quality? Therefore, in this design, a scoring triple of image quality, text quality, and cross-modal consistency is obtained through modal feature extraction, providing a basis for subsequent decision-making. The PPO algorithm in reinforcement learning is used to dynamically select the processing path according to the quality scoring triple to solve the optimization problem of path selection. Through a closed-loop feedback mechanism for quality assessment, when the data quality is lower than the threshold, regeneration is triggered, and the feature extraction process is iteratively optimized to ensure the high quality of the final features.
[0054] The method of using the PPO algorithm for dynamic routing decision-making includes: Define the state space S = [s_img, s_txt, s_consist, CPU load, remaining power, network bandwidth, storage capacity]. The system resource status (such as CPU load, network bandwidth) in the next period of time is predicted using the LSTM model, and define the action space A = {pure image path, pure text path, multi-modal fusion path}; Take the initially output quality scores (such as SSIM, NER accuracy, CLIP image-text matching degree) as part of the reinforcement learning reward function, and define the reward function Use the initial quality score as the initial reward for reinforcement learning; accelerate the training convergence.
[0055] Dynamically adjust the hyperparameters for edge devices and high-performance servers, introduce a long-term reward mechanism, and use discounted return to calculate the long-term performance; for example, for edge devices: = 0.7, = 0.5, = 0.5, = 0.5; for high-performance servers: = 1, = 0.3, = 0.2, = 0.2.
[0056] Design the routing rule as a soft constraint, and select the path through probability distribution, including: If s_img > 0.9 and s_consist > 0.8, then select the pure image path with a 90% probability and the multi-modal fusion path with a 10% probability; If s_txt > 0.9 and s_consist > 0.8, then select the pure text path with a 90% probability and the multi-modal fusion path with a 10% probability; If s_img < 0.4, switch to the text path with an 80% probability and select the multimodal fusion path with a 20% probability; If s_txt < 0.4, switch to the image path with an 80% probability and select the multimodal fusion path with a 20% probability.
[0057] If 0.4 ≤ s_img, s_txt ≤ 0.9 or s_consist < 0.6, select the multimodal fusion path with a 70% probability and select the unimodal path with a 30% probability; A pre-trained rule-based routing strategy (such as the threshold judgment method) can be used as the initial strategy to accelerate convergence. Record the basis for each routing decision (such as quality scores, system resource status, action probability distribution) to generate a decision log for subsequent analysis and optimization.
[0058] The output is the decision action a ∈ A.
[0059] The feature extraction and optimization module uses the PPO (Proximal Policy Optimization) algorithm to dynamically select the processing path (such as the pure image path, pure text path, or multimodal fusion path) according to the quality score triple. This dynamic routing mechanism not only improves computational efficiency but also can adaptively adjust the strategy according to data quality and device resources (such as CPU load, network bandwidth, etc.), which is especially suitable for deployment on edge devices and in complex environments.
[0060] The execution methods of the path rules include: For the pure image path, use a deep CNN (such as ResNet) to extract visual features and generate text features by combining the initial OCR recognition results; For the text modality, use a deep NLP model (such as BERT) to extract semantic features and generate structured features by combining the output results of DistilBERT; For the multimodal fusion path, use a cross-modal SP-Transformer model (such as ViLT) to fuse the original multimodal data and generate structured and text features of the multimodal data to generate a generated modal embedding vector; The cross-modal SP-Transformer model is based on the two-stream Transformer framework, introduces MT-ARNet to process fuzzy features, and adds NSR to intervene in the conflicting data in the original multimodal data and the generated multimodal data; Record the text features, structured features, and embedding vectors as comprehensive features; The method for obtaining the final file classification result includes: Step C1: Construct a multimodal feature matrix based on the comprehensive features, where the multimodal feature matrix forms a unified feature representation by normalizing and aligning the dimensions of different modal features; Step C2: Design a multi-modal classifier using a multi-layer perceptron fusion network based on Attention-Transformer. Introduce a modal weight adaptive adjustment module into the attention mechanism of Transformer to dynamically adjust the weights of image, text, and structured features according to the modal importance of different file types; Step C3: For the file classification task, pre-define a file type label set. For example, the label set includes administrative files, technical files, including but not limited to financial files and historical files; Use the supervised learning method to pre-train the multi-modal classifier. Among them, the training data includes the enhanced multi-modal data and its corresponding file type labels, and the training objective is to minimize the classification cross-entropy loss. At the same time, introduce a regularization term to prevent overfitting; Step C4: Input the multi-modal feature matrix into the trained multi-modal classifier to output the classification probability distribution of the file; Select the file type label with the highest confidence through the probability distribution as the final classification result; Step C5: Design a classification result verification mechanism. For file data with a classification confidence lower than the preset threshold, trigger the manual review process, and feedback the manual review result to the training data set of the dual-channel generator to form a closed-loop optimization; Define the quality control and model update strategy of the feedback data through a weighted loss function or incremental learning.
[0061] Step C6: Output the final file classification result, including the file type label, classification confidence, and the corresponding multi-modal feature matrix, for subsequent file management and retrieval.
[0062] The classification result is accompanied by an interpretable report, such as "The priority is higher than the text keywords because the title contains 'technical cooperation' and the official seal is detected". Provide a real-time adjustment function through interactive rule debugging. Users can drag the highlighted area in the heat map or input natural language to modify the rules and immediately view the adjusted classification result. The rule adjustment and classification result are recorded through the blockchain log to support historical traceability. For example, users can query the classification history of a certain file to see if it has changed due to rule adjustment.
[0063] The classification and verification module designs a multi-modal classifier based on Attention-Transformer and introduces a modal weight adaptive adjustment module, which can dynamically adjust the weights of different modal features according to the characteristics of file types. This design significantly improves the classification accuracy and generalization ability, and is particularly suitable for processing scenarios with large differences in modal importance in file data.
[0064] The classification result verification mechanism forms a closed-loop optimization through manual review feedback and model re-training, which can continuously improve the performance of the classifier and reduce the misclassification rate, especially in dealing with complex or low-confidence files.
[0065] Example 2 See also Figure 2 As shown, the part not described in detail in this embodiment is described in Example 1, which provides a multimodal intelligent AI classification method for file organization, including: Step S1: Obtain a scanned copy of a paper file through a scanner or camera, combine it with the electronic document to form original multimodal data, and transmit it to the perception layer through an API interface; Step S2: In the perception layer, a dual-channel generator is used to generate multimodal data through adversarial training, a cross-modal consistency loss function and a structured constraint loss are introduced, and associated image features or pseudo-text descriptions are generated based on a preset differential quality threshold; Step S3: Design a modality quality evaluator to score the original multimodal data and the generated multimodal data based on image quality, text quality and cross-modal consistency indicators; through the conflict identification and resolution mechanism, mark the data below the cross-modal consistency threshold, generate multiple possible multimodal data and re-evaluate, and output the enhanced multimodal data; Step S4: extract modal features from the enhanced multimodal data to obtain a quality score triplet (s_img, s_txt, s_consist); use the PPO algorithm to make dynamic routing decisions, set the state space, action space and reward function, introduce closed-loop optimization, and trigger data regeneration if the quality score is lower than the threshold, update the quality score triplet and output the decision action; Step S5: Execute path rules according to the decision action, including pure image path, pure text path and multimodal fusion path, use deep CNN, NLP model or cross-modal SP-Transformer model to extract features and output comprehensive features; Step S6: construct a multimodal feature matrix based on the comprehensive features, and form a unified feature representation through normalization and dimension alignment; design a multimodal classifier based on Attention-Transformer, introduce a modal weight adaptive adjustment module, and use supervised learning pre-training; Step S7: Input the multimodal feature matrix into the classifier, output the archive classification probability distribution, and select the highest confidence label as the classification result; for archive data with confidence lower than the threshold, trigger manual review and feedback optimization; Step S8: Output the final archive classification results, including type labels, classification confidence, and multimodal feature matrix.
[0066] Example 3 This embodiment discloses an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the operation mode of the above-provided multi-modal intelligent AI classification system for file sorting.
[0067] Since the electronic device introduced in this embodiment is the electronic device adopted for implementing a multi-modal intelligent AI classification system for file sorting in the embodiments of the present application, based on the multi-modal intelligent AI classification system introduced in the embodiments of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of the present application will not be described in detail here. As long as those skilled in the art implement the electronic device adopted for the multi-modal intelligent AI classification system in the embodiments of the present application, it falls within the scope of protection of the present application.
[0068] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain a formula that is closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.
[0069] The above are only the preferred implementation manners of the present invention. The protection scope of the present invention is not limited to the above embodiments. Any technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for ordinary technical users in the technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.
Claims
1. A multimodal intelligent AI classification system for archive sorting, characterized by: include: Data collection and transmission module: obtain scanned copies of paper archives through scanners or cameras, and combine them with electronic documents to form original multimodal data; Perception layer module: used to receive raw multimodal data, introduce cross-modal consistency loss function and structured constraint loss in the deployed dual-channel generator, generate associated image features or pseudo-text descriptions, and form evolving multimodal data; Based on the modal quality assessor, the original multimodal data and the evolved multimodal data are scored for mutuality; through the conflict identification and resolution mechanism, the enhanced multimodal data is output through cyclic optimization; Feature extraction and optimization module: Extract features based on enhanced multimodal data to obtain quality score triples; use the PPO algorithm to make dynamic routing decisions, introduce closed-loop optimization of quality assessment, optimize and update quality score triples, and output decision actions; execute path rules according to decision actions and output comprehensive features; Classification and verification module: construct a multimodal feature matrix based on comprehensive features, design a multimodal classifier based on Attention-Transformer, introduce a modal weight adaptive adjustment module, set the multimodal feature matrix as input, and output the classification result; add a classification result verification mechanism to form a closed-loop optimization; obtain the final archive classification result.
2. The multimodal intelligent AI classification system for file organization according to claim 1 is characterized in that: The enhanced multimodal data is obtained by: Obtain scanned copies of paper archives through scanners or cameras, combine them with electronic documents to form original multimodal data, and then transmit them to the perception layer through the API interface; A dual-channel generator is deployed in the perception layer. The dual-channel generator learns the cross-modal mapping relationship between images and texts through adversarial training. A cross-modal consistency loss function is introduced in the adversarial training, and a structured constraint loss is added. Differentiated quality thresholds are preset for archive types under different modalities, and associated image features are generated based on text, or pseudo-text descriptions are generated based on images to obtain evolving multimodal data. Design a modality quality evaluator to score the co-generacy of original multimodal data and evolved multimodal data; the multi-dimensional evaluation indicators of the modality quality evaluator include image quality or text quality, and cross-modal consistency; Introducing conflict identification and resolution mechanisms to output optimal evolving multimodal data; The output is enhanced multimodal data, including original multimodal data, evolved multimodal data and symbiosis scores.
3. The multimodal intelligent AI classification system for file organization according to claim 2 is characterized in that: The dual-channel generator comprises: For image modality data, we deploy Beter-DeblurGAN-v3 to automatically repair images and extract local key areas. Beter-DeblurGAN-v3 includes: adding multi-scale convolution kernels to the encoder of DeblurGAN-v3 to capture image interference of different scales, and adding RPN region proposal network to automatically locate and enhance key areas. For text data, the dynamic block segmentation strategy of the VS-ADSP document structure parser is used to identify non-standard typesetting, and the document logical topology is reconstructed in combination with the SB-GNN graph neural network to convert non-standard typesetting documents into standard versions; the VS-ADSP document structure parser includes the introduction of document segmentation evolution based on visual-semantic union in ADSP, and the SB-GNN graph neural network includes adding visual features and semantic features to the node features of GNN; For isolated modal data, conditional variational autoencoder CVAE is introduced in the cross-modal generative adversarial network Cross-GAN, and variational constraints are added to the generator of Cross-GAN. For extreme isolated modal data, visual question answering enhancement VQA-E is introduced, and the semantic description of the image is generated through the pre-trained VQA-E model.
4. The multimodal intelligent AI classification system for file organization according to claim 3 is characterized in that: The method for introducing the conflict identification and resolution mechanism includes: A threshold of cross-modal consistency is predefined. If the cross-modal consistency value output by the modality quality evaluator is less than the threshold, it is marked as a conflict. For conflicting original multimodal data and evolved multimodal data, a dual-channel generator is used to generate multiple possible evolved multimodal data, and then a modal quality assessor is used to re-evaluate the cross-modal consistency until the conflict is resolved.
5. The multimodal intelligent AI classification system for file organization according to claim 4 is characterized in that: The method for obtaining the comprehensive features includes: Perform preliminary modal feature extraction on the enhanced multimodal data to obtain the quality score triplet (s_img, s_txt, s_consist); The PPO algorithm is used for dynamic routing decision making. The input is set as a quality score triplet (s_img, s_txt, s_consist). A closed-loop optimization of quality assessment is introduced, including: if the quality score of the modal data is lower than the preset threshold, it is fed back to the dual-channel generator to trigger the regeneration of the data. The modal feature extraction is performed again based on the regenerated modal data, and the quality score triplet is updated until the closed-loop optimization of quality assessment is completed, and the final quality score triplet (s_img, s_txt, s_consist) is obtained. Use the final quality score triplet to import into the PPO algorithm and output the final decision action; According to the decision action execution path rules, comprehensive features are output.
6. The multimodal intelligent AI classification system for file organization according to claim 5, characterized in that: The method of performing preliminary extraction of modal features on the enhanced multimodal data to obtain a quality score triplet (s_img, s_txt, s_consist) includes: Step Q1: Perform resolution detection on the image data in the evolving multimodal data, select a suitable feature extraction model based on the resolution, the features include layout, seal position, signature area confidence and initial OCR recognition results, and extract the visual feature vector f_img; For the text data in the evolving multimodal data, rules are extracted to extract key fields and use them as auxiliary inputs for DistilBERT. The features include keyword frequency, named entities, and paragraph topic distribution, and the semantic feature vector f_txt is extracted. Step Q2: Calculate the OCR recognition accuracy based on the edit distance and semantic rationality, introduce multi-scale analysis to calculate the blur and SSIM values in the Laplacian variance and SSIM structure integrity, use YOLOv8 to detect the seal area, take the highest confidence score as the seal detection confidence, and output the image quality score s_img through the calculation formula; the calculation formula is , is the regularization coefficient, is the variance of each indicator or other regularization term; Calculate the semantic coherence score based on DistilBERT's NextSentencePrediction, and use SB-GNN to analyze the document layout and evaluate the structural completeness; Output the text quality score s_txt through the calculation formula; the calculation formula is ; Step Q3: Calculate the similarity between the image OCR result and the electronic text title, use the CLIP model to calculate the image-text matching degree between the image and the text, and output the cross-modal consistency score s_consist. The calculation formula is: s_consist= ·Similarity+· CLIP image and text matching degree; in, , , , , , , and is the weight coefficient, which is based on the preset weights of archival metadata matching; Step Q4: Output quality score triplet (s_img, s_txt, s_consist).
7. The multimodal intelligent AI classification system for file organization according to claim 6, characterized in that: The method for making dynamic routing decisions using the PPO algorithm includes: Define the state space S = [s_img, s_txt, s_consist, CPU load, remaining power, network bandwidth, storage capacity], define the action space A = {pure image path, pure text path, multimodal fusion path}; Defining the reward function , using the initial quality score as the initial reward for reinforcement learning; Dynamically adjust hyperparameters for edge devices and high-performance servers, introduce a long-term reward mechanism, and use discounted returns to calculate long-term performance; Design routing rules as soft constraints and select paths through probability distribution, including: If s_img>0.9 and s_consist>0.8, the pure image path is selected with a probability of 90% and the multimodal fusion path is selected with a probability of 10%; If s_txt>0.9 and s_consist>0.8, the pure text path is selected with a probability of 90% and the multimodal fusion path is selected with a probability of 10%; If s_img<0.4, the probability of switching to the text path is 80%, and the probability of selecting the multimodal fusion path is 20%; If s_txt<0.4, the probability of switching to the image path is 80%, and the probability of selecting the multimodal fusion path is 20%; If 0.4≤s_img, s_txt≤0.9 or s_consist<0.6, the multimodal fusion path is selected with a probability of 70% and the single modal path is selected with a probability of 30%; The output is the decision action a∈A; in, , , and is the weight coefficient.
8. The multi-modal intelligent AI classification system for file organization according to claim 7, characterized in that: The execution method of the path rule includes: For the pure image path, deep CNN is used to extract visual features and combined with the initial OCR recognition results to generate text features; For text modality, we use a deep NLP model to extract semantic features and combine them with the output of DistilBERT to generate structured features. For the multimodal fusion path, the cross-modal SP-Transformer model is used to fuse the structural and textual features of the original multimodal data and the evolved multimodal data to generate a generative modality embedding vector; The cross-modal SP-Transformer model is based on the two-stream Transformer framework, introduces MT-ARNet to process fuzzy features, and adds NSR to intervene in the conflicting data in the original multimodal data and the evolved multimodal data; The text features, structural features and embedding vectors are recorded as comprehensive features.
9. The multimodal intelligent AI classification system for file organization according to claim 8, characterized in that: The method for obtaining the final archive classification result includes: Step C1: construct a multimodal feature matrix based on the comprehensive features, wherein the multimodal feature matrix forms a unified feature representation by normalizing and aligning the dimensions of different modal features; Step C2: Design a multimodal classifier based on the Attention-Transformer multi-layer perception fusion network, and introduce a modality weight adaptive adjustment module into the Transformer attention mechanism to dynamically adjust the weights of image, text, and structured features according to the modality importance of different archive types; Step C3: For the archive classification task, a set of archive type labels is predefined, and a multimodal classifier is pre-trained using a supervised learning method. The training goal is to minimize the classification cross entropy loss, and a regularization term is introduced at the same time; Step C4: Input the multimodal feature matrix into the trained multimodal classifier and output the classification probability distribution of the archive; select the archive type label with the highest confidence as the final classification result through the probability distribution; Step C5: Design a classification result verification mechanism. For archival data whose classification confidence is lower than the preset threshold, trigger the manual review process and feed back the manual review results to the dual-channel generator to form a closed-loop optimization; Step C6: Output the final archive classification results, including the archive type label, classification confidence and corresponding multimodal feature matrix, for subsequent archive management and retrieval.
10. A multimodal intelligent AI classification method for file organization, applied to the multimodal intelligent AI classification system for file organization according to any one of claims 1 to 9, characterized in that: The multimodal intelligent AI classification method for archive sorting includes: Step S1: Obtain a scanned copy of a paper file through a scanner or camera, combine it with the electronic document to form original multimodal data, and transmit it to the perception layer through an API interface; Step S2: In the perception layer, a dual-channel generator is used to generate multimodal data through adversarial training, a cross-modal consistency loss function and a structured constraint loss are introduced, and associated image features or pseudo-text descriptions are generated based on a preset differential quality threshold; Step S3: Design a modality quality evaluator to score the original multimodal data and the generated multimodal data based on image quality, text quality and cross-modal consistency indicators; through the conflict identification and resolution mechanism, mark the data below the cross-modal consistency threshold, generate multiple possible multimodal data and re-evaluate, and output the enhanced multimodal data; Step S4: extract modal features from the enhanced multimodal data to obtain a quality score triplet (s_img, s_txt, s_consist); use the PPO algorithm to make dynamic routing decisions, set the state space, action space and reward function, introduce closed-loop optimization, and trigger data regeneration if the quality score is lower than the threshold, update the quality score triplet and output the decision action; Step S5: Execute path rules according to the decision action, including pure image path, pure text path and multimodal fusion path, use deep CNN, NLP model or cross-modal SP-Transformer model to extract features and output comprehensive features; Step S6: construct a multimodal feature matrix based on the comprehensive features, and form a unified feature representation through normalization and dimension alignment; design a multimodal classifier based on Attention-Transformer, introduce a modal weight adaptive adjustment module, and use supervised learning pre-training; Step S7: Input the multimodal feature matrix into the classifier, output the archive classification probability distribution, and select the highest confidence label as the classification result; for archive data with confidence lower than the threshold, trigger manual review and feedback optimization; Step S8: Output the final archive classification results, including type labels, classification confidence, and multimodal feature matrix.
Citation Information
Patent Citations
Construction method and system for multi-modal file auditing
CN118760769A
Layered contrast anti-fact learning method and system oriented to visual question and answer model
CN119166795A
Knowledge data classification multi-modal reasoning evolution system based on depth model
CN119862961A
Multimodal fine-grained mixing method and system, device, and storage medium
US20220237420A1
Cited By
Digital system work order auditing method based on artificial intelligence
CN120598372A
Document structured analysis method based on multi-modal large language model and OCR (optical character recognition) enhancement
CN121074919A
Document structuring analysis method based on multi-modal large language model and OCR enhancement
CN121074919B
Archive data automatic classification method and system based on machine learning
CN121211129A
Archive management system and method based on artificial intelligence
CN121278159A