An archive intelligent opening auditing method and system based on a large model

By building a large-scale model-based intelligent open access review platform and engine for archives in a cloud data center, and using artificial intelligence algorithms for archive text recognition, open access permission classification, and sensitive information detection, the problems of low efficiency and insufficient accuracy in archive review have been solved, and efficient and accurate sensitive information detection and archive data sharing have been achieved.

CN120216743BActive Publication Date: 2025-12-16WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510263565.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-12-16
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

Existing archival review technologies are inefficient, inaccurate, and have a low level of informatization, making it difficult to meet the needs of rapidly processing massive amounts of archival data. Furthermore, they lack effective means of detecting sensitive information, posing a risk of information leakage.

Method used

The method for intelligent open access review of archives based on a large model involves building an intelligent open access review platform and engine for archives in a cloud data center. It utilizes artificial intelligence algorithms to perform text recognition, open access permission classification, sensitive information detection, and automatic review of archives, generating review reports and enabling visualization.

Benefits of technology

It has automated and systematized the process of document review, improved efficiency and accuracy, reduced misjudgments and omissions, ensured the accurate identification and masking of sensitive information, avoided the risk of information leakage, and enabled the rapid review and sharing of document data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216743B_ABST
    Figure CN120216743B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of file review, and discloses a file intelligent opening review method and system based on a large model.The method comprises the following steps: a cloud data center, building a file intelligent opening review platform, and constructing a file intelligent opening review engine; the cloud data center, using the file intelligent opening review engine, generating second file text data of each second file, corresponding first sensitive information and second sensitive information mask; the cloud data center, using the file intelligent opening review engine, generating opening review second file text data and corresponding second automatic review data; the cloud data center, according to the second artificial review data and the second automatic review data, using the file intelligent opening review engine, generating a second review report, and visualizing on the file intelligent opening review platform.The application solves the problems of low efficiency, insufficient accuracy and low informatization degree in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of file review, and particularly relates to a file intelligent opening review method and system based on a large model. BACKGROUND

[0002] With the rapid development of information technology, file digitization has become an irreversible trend. Digitization not only improves the storage, retrieval and utilization efficiency of files, but also makes file information more easily shared and disseminated. File opening review serves the opening and utilization of file data, prevents sensitive information from being opened, and avoids threats to national security and invasion of personal privacy through file content review. At the same time, with the continuous progress of technology, the methods and means of file review will continue to innovate and improve, providing more efficient and convenient services for file data opening and utilization work.

[0003] The existing file review technology has the following defects:

[0004] 1) Low efficiency: Traditional file review relies on manual operation, and in the face of massive file data, manual review is slow and difficult to meet the demand for rapid processing;

[0005] 2) Lack of accuracy: Manual review is greatly affected by subjective factors, and is prone to misjudgment or omission, affecting the accuracy of file review, and the existing technology lacks effective sensitive information detection means, making it difficult to accurately identify and mask sensitive information, and there is a risk of information leakage;

[0006] 3) Low informationization: The existing technology has low informationization, making it difficult to realize rapid review and sharing of file data, and the file data format is diverse, lacking unified standards and specifications, making data integration and exchange difficult. SUMMARY

[0007] In order to solve the problems of low efficiency, lack of accuracy and low informationization in the existing technology, the application aims to provide a file intelligent opening review method and system based on a large model.

[0008] The technical solution adopted by the application is:

[0009] A file intelligent opening review method based on a large model, comprising the following steps:

[0010] A cloud data center is built to build a file intelligent opening review platform, an artificial intelligence algorithm is used to build a file intelligent opening review engine in the cloud data center, and the file intelligent opening review engine is connected to the file intelligent opening review platform;

[0011] The cloud data center receives a plurality of second archive files sent by the data server, uses an archive intelligent open audit engine to generate second archive text data of each second archive file, corresponding first sensitive information, and second sensitive information masks;

[0012] The cloud data center uses the archive intelligent open audit engine to generate open audit second archive text data and corresponding second automatic audit data according to all second archive text data, corresponding second sensitive information, and second sensitive information masks, and visualizes the open audit second archive text data and the corresponding second automatic audit data on the archive intelligent open audit platform.

[0013] The cloud data center uses the archive intelligent open audit engine to generate a second audit report according to the second manual audit data returned by the user terminal for the open audit second archive text data and the corresponding second automatic audit data, and visualizes the second audit report on the archive intelligent open audit platform.

[0014] Further, the archive intelligent open audit platform includes a user login module, a data upload module, an archive visualization module, an archive audit module, and a report visualization module.

[0015] The archive intelligent open audit engine includes an archive text recognition model, an archive open permission classification model, a sensitive information detection model, an archive automatic audit model, and an audit report generation model.

[0016] Further, the cloud data center builds an archive intelligent open audit platform, uses artificial intelligence algorithms to build an archive intelligent open audit engine in the cloud data center, and connects the archive intelligent open audit engine to the archive intelligent open audit platform, including the following steps:

[0017] The cloud data center builds an archive intelligent open audit framework, sets a user login module, a data upload module, an archive visualization module, an archive audit module, and a report visualization module, and obtains an archive intelligent open audit platform.

[0018] A plurality of first archive files and a plurality of first manual audit data are collected and preprocessed to obtain a plurality of preprocessed first archive files and a plurality of preprocessed first manual audit data.

[0019] According to the plurality of preprocessed first archive files, a text recognition algorithm is used to build an archive text recognition model to generate a plurality of first archive text data.

[0020] According to the plurality of first archive text data, a deep learning algorithm is used to build an archive open permission classification model to generate a first archive open permission classification result for each first archive text data.

[0021] According to the external text big data, and a plurality of first archive text data and first archive open permission classification results, using a large model algorithm, a sensitive information detection model is constructed, and a plurality of first sensitive information is generated;

[0022] According to a plurality of first archive text data, corresponding first archive open permission classification results and first sensitive information, using a deep learning algorithm, an archive automatic review model is constructed, and a plurality of first automatic review data is generated;

[0023] According to a plurality of pre-processed first artificial review data and corresponding first automatic review data, using a deep learning algorithm, an audit report generation model is constructed;

[0024] The archive text recognition model, the archive open permission classification model, the sensitive information detection model, the archive automatic review model and the audit report generation model are integrated, an archive intelligent open review engine is constructed in a cloud data center, and the archive intelligent open review engine is connected to an archive intelligent open review platform.

[0025] Further, the text recognition model is constructed based on an FPN-LSTM-CRF algorithm, and the text recognition model comprises an image feature extraction module constructed based on an FPN algorithm, a sequence feature extraction module constructed based on an LSTM algorithm and a recognized text label generation module constructed based on a CRF algorithm connected in sequence;

[0026] The archive open permission classification model is constructed based on an LSTM-DBN algorithm, and the archive open permission classification model comprises a semantic feature extraction module constructed based on an LSTM algorithm and an archive open permission classification module constructed based on a DBN algorithm connected in sequence;

[0027] The sensitive information detection model is constructed based on a RoBERTa-Transfomer-CRF algorithm, and the sensitive information detection model comprises a word embedding module based on RoBERTa, a deep feature extraction module constructed based on a Transfomer algorithm and a sensitive information label generation module constructed based on a CRF algorithm connected in sequence;

[0028] The archive automatic review model is constructed based on an RF-MLP algorithm, and the archive automatic review model comprises a key feature extraction module constructed based on an RF algorithm and an archive automatic review module constructed based on an MLP algorithm connected in sequence;

[0029] The audit report generation model is constructed based on a cGAN-MLP algorithm, and the audit report generation model comprises a generator and a discriminator both constructed based on an RNN algorithm, and a condition embedding module and a condition processing module constructed based on an MLP algorithm, the generator is connected with the discriminator and the condition embedding module respectively, and the condition processing module is connected with the discriminator.

[0030] Further, the cloud data center receives a plurality of second archive files sent by the data server, uses an archive intelligent open audit engine to generate second archive text data, corresponding second sensitive information, and second sensitive information masks of each second archive file, including the following steps:

[0031] The cloud data center receives a plurality of second archive files sent by the data server, and uses an archive text recognition model to generate second archive text data of each second archive file;

[0032] Using an archive open permission classification model, the archive open permission classification of each second archive text data is performed to obtain the corresponding second archive open permission classification result;

[0033] According to the second archive open permission classification result, a sensitive information detection model is used to detect sensitive information in each second archive text data to obtain a plurality of second sensitive information and corresponding second sensitive information masks.

[0034] Further, the first archive open permission classification result includes the first archive field and the first open permission of the first archive text data;

[0035] The second archive open permission classification result includes the second archive field and the second open permission of the second archive text data.

[0036] Further, the cloud data center uses an archive intelligent open audit engine to generate open audit second archive text data and corresponding second automatic audit data according to all second archive text data, corresponding second sensitive information, and second sensitive information masks, and visualizes it on an archive intelligent open audit platform, including the following steps:

[0037] The cloud data center uses an archive automatic audit model to perform archive automatic audit according to a plurality of second archive text data, corresponding second archive open permission classification results, and second sensitive information to obtain corresponding second automatic audit data;

[0038] According to each second archive text data and corresponding second sensitive information masks, open audit second archive text data is generated;

[0039] The archive intelligent open audit platform is used to visualize the open audit second archive text data.

[0040] Further, the cloud data center uses an archive intelligent open audit engine to generate a second audit report according to the second artificial audit data returned by the user terminal for the open audit second archive text data and the corresponding second automatic audit data, and visualizes it on an archive intelligent open audit platform, including the following steps:

[0041] The cloud data center receives second artificial auditing data returned by the user terminal for the second artificial auditing of the second archival text data using the archival intelligent open auditing platform;

[0042] According to the second artificial auditing data and the corresponding second automatic auditing data, a second auditing report is generated using the archival intelligent open auditing engine;

[0043] The second auditing report is visualized using the archival intelligent open auditing platform.

[0044] An archival intelligent open auditing system based on a large model is used to implement an archival intelligent open auditing method. The system includes a cloud data center and a plurality of user terminals, the plurality of user terminals are in communication connection with the cloud data center, the cloud data center is provided with an archival intelligent open auditing platform and an archival intelligent open auditing engine, and the cloud data center includes a platform initialization unit, an archival text generation unit, an automatic auditing processing unit and an auditing report generation unit connected in sequence.

[0045] The archival intelligent open auditing system based on a large model has the following beneficial effects:

[0046] The archival intelligent open auditing method and system based on a large model provided by the present application use an archival intelligent open auditing engine constructed by an artificial intelligence algorithm to realize a systematic auditing process of automatic archival text recognition, archival open permission classification, sensitive information detection, archival automatic auditing and auditing report generation, significantly improving the efficiency of archival auditing, shortening the auditing period and meeting the demand for rapid processing of a large number of archives. The use of advanced archival text recognition models, archival open permission classification models and sensitive information detection models reduces misjudgment and omission and ensures the accuracy of archival auditing. The archival data is automatically audited by the archival automatic auditing model to improve the auditing efficiency and assist artificial auditing. The sensitive information detection model based on a large model has strong semantic understanding ability, which can improve the accuracy of sensitive information recognition and judgment, realize effective sensitive information detection means, accurately identify and generate sensitive information masks and avoid information leakage risks. The archival intelligent open auditing platform realizes online intelligent open auditing of archives, realizes rapid auditing and sharing of archival data, and uses the cloud data center to adopt unified standards and specifications to uniformly manage and analyze archival files of different data sources, simplifies the difficulty of data integration and exchange, avoids information silos and improves the informatization level.

[0047] Other beneficial effects of the present application will be further described in the specific embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 is a flowchart of the archival intelligent open auditing method based on a large model in the present application.

[0049] Figure 2 is a structural diagram of the large model-based archive intelligent open audit system in the present application. DETAILED DESCRIPTION

[0050] The present application will be further explained in conjunction with the accompanying drawings and specific embodiments.

[0051] Embodiment 1:

[0052] As shown in Figure 1 , the present embodiment provides a large model-based archive intelligent open audit method, comprising the following steps:

[0053] S1: cloud data center, build an archive intelligent open audit platform, use artificial intelligence algorithm, build an archive intelligent open audit engine in the cloud data center, and connect the archive intelligent open audit engine to the archive intelligent open audit platform, comprising the following steps:

[0054] S1-1: cloud data center, build an archive intelligent open audit framework, and set user login module, data upload module, archive visualization module, archive audit module and report visualization module, to obtain an archive intelligent open audit platform;

[0055] The archive intelligent open audit platform includes a user login module, a data upload module, an archive visualization module, an archive audit module and a report visualization module;

[0056] The user login module is used to receive the login data input by the user; the data upload module is used to upload the login data to the cloud data center for user login; the archive visualization module is used to visualize the open audit second archive text data; the archive audit module is used to receive the second artificial audit data of the user for the open audit second archive text data; and the report visualization module is used to visualize the second audit report;

[0057] S1-2: collect a plurality of first archive files and a plurality of first artificial audit data, and pre-process to obtain a plurality of pre-processed first archive files and a plurality of pre-processed first artificial audit data;

[0058] The preprocessing includes data cleaning, format conversion and normalization processing, etc., to improve the data quality and provide data support for subsequent model training;

[0059] S1-3: according to the plurality of pre-processed first archive files, use a character recognition algorithm to build an archive text recognition model, and generate a plurality of first archive text data;

[0060] The text recognition model is built on the Feature Pyramid Networks (FPN) - Long Short-Term Memory (LSTM) - Conditional Random Field (CRF) algorithm. The text recognition model includes an image feature extraction module built on the FPN algorithm, a sequence feature extraction module built on the LSTM algorithm, and a recognition text label generation module built on the CRF algorithm, which are connected in sequence.

[0061] The image feature extraction module utilizes the top-down and bottom-up paths of FPN to effectively extract image features at different scales, ensuring accurate detection of various details in archival images and achieving the fusion of features at different levels, thus enhancing the expressive power of features. The sequence feature extraction module's LSTM is adept at processing sequence data and can capture long-short-term dependencies in text sequences, providing rich sequence features for subsequent text label generation. The text label generation module uses CRF constraints to optimize the generated labels, ensuring the rationality and accuracy of the labels and making the generated text label sequences more syntactically and semantically sound.

[0062] The entire text recognition model achieves a complete process from image feature extraction to text label generation through the organic combination of FPN, LSTM, and CRF. FPN ensures the comprehensiveness and accuracy of image features; LSTM captures the temporal and contextual information of the text sequence; and CRF optimizes and constrains the generated labels. This structure not only improves the accuracy of text recognition but also enhances the model's adaptability to various complex scenarios.

[0063] S1-4: Based on several first-archive text data, use deep learning algorithms to construct an archive access permission classification model and generate the first archive access permission classification result for each first-archive text data;

[0064] The archive access permission classification model is based on LSTM-Deep Belief Network. The archive access permission classification model is constructed using the Network (DBN) algorithm, and includes a semantic feature extraction module based on the LSTM algorithm and an archive access permission classification module based on the DBN algorithm, which are connected sequentially.

[0065] The LSTM of the semantic feature extraction module can capture long-distance dependencies in the text, making the model better understand the semantics and context of the archival text data, thereby improving the accuracy of classification; the deep structure of the DBN of the archival open access permission classification module enables the model to learn more complex and abstract feature representations, thereby improving the accuracy and robustness of archival open access permission classification;

[0066] By organically combining LSTM and DBN, the archival open access permission classification model realizes a complete process from semantic feature extraction to classification decision. The LSTM module is responsible for capturing deep semantic features and temporal information in the archival text, providing rich feature representations for subsequent classification. The DBN module further refines and abstracts these features using its deep structure and pre-training mechanism, and makes accurate classification decisions. This structure not only improves the accuracy of archival open access permission classification, but also enhances the model's adaptability and generalization ability for complex archival text.

[0067] S1-5: According to external text big data and a plurality of first archival text data and first archival open access permission classification results thereof, a sensitive information detection model is constructed using a large model algorithm, and a plurality of first sensitive information is generated;

[0068] The sensitive information detection model is constructed based on a Robustly optimized BERT approach (RoBERTa)-Transfomer-CRF algorithm, and the sensitive information detection model comprises a RoBERTa-based word embedding module, a deep feature extraction module constructed based on a Transfomer algorithm, and a sensitive information label generation module constructed based on a CRF algorithm connected in sequence.

[0069] The word embedding module is an improved and adjusted model based on the Transformer-based bidirectional encoder representation (BERT) to improve performance. The module converts the input text into word vector representation, each word vector containing semantic information, contextual information and positional information of the word, improving the understanding of polysemous words and complex sentences, preserving the position information of the word in the sentence, which helps the model understand the order and structure of the word; the Transformer of the deep feature extraction module uses self-attention mechanism to capture long-distance dependency in the text, and generates a feature representation containing global information for each word. The self-attention mechanism effectively captures long-distance dependencies in the text, improves the detection ability of complex sensitive information, and the multi-layer stacking structure enables the model to learn multi-level and multi-angle text features, enhancing the expression ability of the features; the sensitive information label generation module introduces constraints such as the legality constraint between labels, further optimizes the generation of label sequences, and has certain robustness to noise data and abnormal situations;

[0070] The sensitive information detection model realizes the complete process from word embedding to sensitive information label generation through the organic combination of RoBERTa, Transformer and CRF. The RoBERTa word embedding module provides rich semantic representation and contextual information; the Transformer deep feature extraction module further captures deep features and long-distance dependencies in the text; and the CRF sensitive information label generation module generates a globally optimal sensitive information label sequence using sequence labeling technology. This structure not only improves the accuracy of sensitive information detection, but also enhances the model's adaptability and robustness to complex text and noise data;

[0071] According to external text big data and a plurality of first archive text data and first archive open permission classification results, a sensitive information detection model is constructed using a large model algorithm, and a plurality of second sensitive information is generated, including the following steps:

[0072] S1-5-1: Collect external text big data from sources such as the Internet, databases, file systems, etc. These data cover various topics, fields and language styles, and preprocess the text big data to obtain a plurality of preprocessed text data; clean, remove duplicates, segment, remove stop words, etc. operations are performed on the collected text big data to eliminate noise and redundant information and improve data quality;

[0073] S1-5-2: Use the RoBERTa-Transfomer-CRF algorithm to construct an initial sensitive information detection model;

[0074] S1-5-3: Pre-training the initial sensitive information detection model using a plurality of pre-processed text data to obtain a pre-trained sensitive information detection model, so that the model has strong semantic understanding and text processing capability; pre-training the model using a large amount of pre-processed text data, so that the model learns general language representation and features, and pre-training enables the model to have strong generalization ability and better adapt to different text data and tasks;

[0075] S1-5-4: Fine-tuning the pre-trained sensitive information detection model according to a plurality of first archive text data and first archive open permission classification results to obtain a final sensitive information detection model; fine-tuning makes the model more suitable for specific sensitive information detection tasks, improves the accuracy and effectiveness of detection, and through the use of first archive text data and classification results, the model can inherit and utilize existing knowledge to further optimize detection performance;

[0076] S1-6: Constructing an archive automatic review model using a deep learning algorithm according to a plurality of first archive text data, corresponding first archive open permission classification results, and first sensitive information, and generating a plurality of first automatic review data;

[0077] The archive automatic review model is constructed based on a Random Forest (RF)-Multilayer Perceptron (MLP) algorithm, and the archive automatic review model includes a key feature extraction module constructed based on the RF algorithm and an archive automatic review module constructed based on the MLP algorithm connected in sequence;

[0078] The key feature extraction module improves prediction accuracy by constructing multiple decision trees and aggregating their predictions, each decision tree is trained on a random subset of the training set, which can evaluate the importance of each feature to the prediction result, thereby identifying key features. By extracting key features, the dimensionality of the data is reduced, simplifying the complexity of the subsequent model, and the extraction of key features helps to remove noise and irrelevant information, improving the prediction accuracy of the model. The archive automatic review module generates automatic review labels by learning the weights and biases between input features and output labels, and can adjust the number of layers and neurons of the MLP as needed to adapt to different complexity requirements;

[0079] By organically combining RF and MLP, the archive automatic review model realizes a complete process from key feature extraction to archive automatic review, the RF module is responsible for screening the most important features for archive automatic review from a large number of features, which reduces the complexity of data and improves the interpretability of the model; the MLP module uses its powerful nonlinear modeling capability to generate accurate open permissions based on the screened key features. This structure not only improves the accuracy and efficiency of archive automatic review, but also enhances the adaptability and interpretability of the model to complex scenarios;

[0080] S1-7: According to a plurality of pre-processed first artificial review data and corresponding first automatic review data, a deep learning algorithm is used to construct an audit report generation model;

[0081] The audit report generation model is constructed based on a conditional generative adversarial network (cGAN)-MLP algorithm, and the audit report generation model includes a generator and a discriminator both constructed based on a recurrent neural network (RNN) algorithm, and a condition embedding module and a condition processing module constructed based on an MLP algorithm. The generator is connected with the discriminator and the condition embedding module, and the condition processing module is connected with the discriminator;

[0082] The conditional embedding module utilizes a Multilayer Perceptron (MLP) to convert the features of the archive access permission classification results, the features of the manual review data, and the text data of the open review archives into embedding vectors that the generator can understand. These embedding vectors serve as conditional information, guiding the generator's generation process. The MLP algorithm effectively handles non-linear relationships, transforming features into useful embedding representations. The conditional embedding module ensures that the generator considers the specific decision context when generating output, improving the relevance and accuracy of the report generation process. The generator receives the embedding vectors from the conditional embedding module and other relevant information, and uses an RNN algorithm to generate an output that matches the input conditions—the review report. The generator attempts to generate sufficiently realistic data to deceive the discriminator. The generator can generate customized outputs based on decision information, improving the model's flexibility and adaptability. The discriminator receives the output from the generator and the output from the conditional processing module. It uses an RNN algorithm to determine whether the audit report generated by the generator is sufficiently realistic, i.e., whether it matches the actual situation of the audit report. The discriminator guides the generator's training process through feedback signals. Adversarial training improves the quality and realism of the generator's generated data, and the adversarial process helps improve the overall performance of the model and the consistency of the output. The conditional processing module receives a portion of the generator's output and the output from the conditional embedding module. It uses an MLP algorithm to process this information, providing the discriminator with additional conditional information to help it better understand the context of the generator's output. The conditional processing module enhances the execution model's ability to evaluate the generator's output, improving the overall reliability of the model. By combining conditional information, it improves the matching degree between the generator's output and the actual application scenario.

[0083] Based on several preprocessed first manual review data and corresponding first automatic review data, a review report generation model is constructed using a deep learning algorithm, including the following steps:

[0084] S1-7-1: Set a corresponding real report for each preprocessed first manually reviewed data. The real report serves as the standard or target when training the model.

[0085] S1-7-2: Use the cGAN-MLP algorithm to build the initial audit report generation model;

[0086] S1-7-3: Combine the first loss function of the initial generator and the second loss function of the initial discriminator in the initial audit report generation model to obtain the comprehensive loss function; ensure that the generator and discriminator promote each other during training to improve the overall model performance;

[0087] S1-7-4: Extract the first manual review data features of the first manual review data and the first automatic review data features of the first automatic review data after preprocessing;

[0088] S1-7-5: using the initial audit report generation model, the conditional information embedder is used to conditionally embed the first artificial audit data features and the first automatic audit data features to obtain a plurality of first conditional embedding features; the ability of the generator to generate reports under certain conditions is improved;

[0089] S1-7-6: according to the plurality of first conditional embedding features, the initial generator of the initial audit report generation model is trained to obtain a plurality of generated reports;

[0090] S1-7-7: using the conditional information processor, the real report and the corresponding generated report are subjected to conditional information processing to obtain first conditional information; it is ensured that the discriminator can be trained based on accurate conditional information;

[0091] S1-7-8: according to the real report, the corresponding generated report and the first conditional information, the initial discriminator is trained to obtain a plurality of first data discrimination results; the ability of the discriminator to identify real and generated data is improved, so as to push the generator to generate higher quality reports;

[0092] S1-7-9: according to each generated report and the corresponding first data discrimination result, a first comprehensive loss value in the training process is obtained using a comprehensive loss function; the training progress of the model is evaluated, which provides a basis for whether to continue training;

[0093] S1-7-10: if the first loss value is lower than the first loss value threshold, the optimized generator and the optimized discriminator are output, otherwise, the optimization training is continued;

[0094] S1-7-11: integrating the optimized generator, the optimized discriminator, the conditional information embedder and the conditional information processor, a final audit report generation model is obtained;

[0095] S1-8: integrating the archive text recognition model, the archive open permission classification model, the sensitive information detection model, the archive automatic audit model and the audit report generation model, an archive intelligent open audit engine is constructed in the cloud data center, and the archive intelligent open audit engine is connected to the archive intelligent open audit platform;

[0096] The archive intelligent open audit engine includes an archive text recognition model, an archive open permission classification model, a sensitive information detection model, an archive automatic audit model and an audit report generation model;

[0097] S2: the cloud data center receives a plurality of second archive files sent by the data server, and uses the archive intelligent open audit engine to generate second archive text data of each second archive file, corresponding second sensitive information and second sensitive information mask, including the following steps:

[0098] S2-1: The cloud data center receives a plurality of second archive files sent by the data server, and uses an archive text recognition model to generate second archive text data of each second archive file, including the following steps:

[0099] S2-1-1: The cloud data center receives a plurality of second archive files sent by the data server, converts the second archive files into second archive images, and inputs the plurality of second archive images into the archive text recognition model;

[0100] S2-1-2: The image feature extraction module of the archive text recognition model extracts second image features of the second archive images;

[0101] S2-1-3: The sequence feature extraction module of the archive text recognition model generates corresponding second sequence features according to the second image features;

[0102] S2-1-4: The recognized text label generation module of the archive text recognition model generates corresponding second archive text data according to the second sequence features;

[0103] S2-1-5: All second archive files are traversed to obtain second archive text data of each second archive file;

[0104] S2-2: Using an archive open permission classification model, each second archive text data is classified according to the archive open permission classification, to obtain a corresponding second archive open permission classification result, including the following steps:

[0105] S2-2-1: The second archive text data output by the archive text recognition model is input into the archive open permission classification model;

[0106] S2-2-2: The semantic feature extraction module of the archive open permission classification model extracts second semantic features of the second archive text data;

[0107] S2-2-3: The archive open permission classification module of the archive open permission classification model classifies the archive open permission according to the second semantic features, to obtain a corresponding second archive open permission classification result;

[0108] S2-2-4: All second archive text data is traversed to obtain a second archive open permission classification result of each second archive text data;

[0109] The second archive open permission classification result includes a second archive field and a second open permission of the second archive text data;

[0110] S2-3: According to the second archive open permission classification result, using the sensitive information detection model, sensitive information detection is performed on each second archive text data to obtain a plurality of second sensitive information and corresponding second sensitive information masks, including the following steps:

[0111] S2-3-1: Using the word embedding module of the sensitive information detection model, word embedding is performed on the second archive text data to obtain the corresponding second word vector;

[0112] S2-3-2: Using the deep feature extraction module of the sensitive information detection model, the second archive open permission classification result feature of the second archive open permission classification result and the second word vector feature of the second word vector are extracted;

[0113] S2-3-3: Using the sensitive information label generation module of the sensitive information detection model, according to the second archive open permission classification result feature and the second word vector feature, sensitive information label generation is performed to obtain a plurality of corresponding second sensitive information;

[0114] S2-3-4: According to the different open permissions and the second position data of the second sensitive information in the second archive text data, a plurality of corresponding second sensitive information masks are generated;

[0115] S2-3-5: All second archive text data is traversed to obtain a plurality of second sensitive information and corresponding second sensitive information masks for each second archive text data;

[0116] S3: The cloud data center uses the archive intelligent open audit engine to generate open audit second archive text data and corresponding second automatic audit data according to all second archive text data, corresponding second sensitive information and second sensitive information masks, and visualizes it on the archive intelligent open audit platform, including the following steps:

[0117] S3-1: The cloud data center uses the archive automatic audit model to perform archive automatic audit according to a plurality of second archive text data, corresponding second archive open permission classification results and second sensitive information to obtain corresponding second automatic audit data, including the following steps:

[0118] S3-1-1: The cloud data center inputs each second archive text data, corresponding second archive open permission classification result and second sensitive information into the archive automatic audit model;

[0119] S3-1-2: Using the key feature extraction module of the archive automatic audit model, a plurality of second key features of each second archive text data, corresponding second archive open permission classification result and second sensitive information are extracted;

[0120] S3-1-3: using the file automatic review module of the file automatic review model to perform file automatic review to obtain corresponding second automatic review data;

[0121] S3-2: generating open review second file text data according to each second file text data and corresponding second sensitive information mask;

[0122] S3-4: using the file intelligent open review platform to visualize the open review second file text data;

[0123] S4: the cloud data center uses the file intelligent open review engine to generate a second review report according to the second artificial review data returned by the user terminal for the open review second file text data and corresponding second automatic review data, and visualizes the file intelligent open review platform, including the following steps:

[0124] S4-1: the cloud data center uses the file intelligent open review platform to receive the second artificial review data returned by the user terminal for the open review second file text data;

[0125] S4-2: using the file intelligent open review engine according to the second artificial review data and the corresponding second automatic review data to generate a second review report, including the following steps:

[0126] S4-2-1: extracting the second artificial review data features of the second artificial review data and the second automatic review data features of the corresponding second automatic review data;

[0127] S4-2-2: using the condition embedding module of the review report generation model to conditionally embed the second artificial review data features and the second automatic review data features to obtain second condition embedding features;

[0128] S4-2-3: using the generator of the review report generation model to generate a review report according to the second condition embedding features to obtain a corresponding second review report;

[0129] S4-3: using the file intelligent open review platform to visualize the second review report.

[0130] Embodiment 2:

[0131] As Figure 2As shown, the embodiment provides an archive intelligent opening audit system based on a large model, which is used to realize an archive intelligent opening audit method. The system includes a cloud data center and a plurality of user terminals. The plurality of user terminals are in communication connection with the cloud data center. The cloud data center is provided with an archive intelligent opening audit platform and an archive intelligent opening audit engine. The cloud data center includes a platform initialization unit, an archive text generation unit, an automatic audit processing unit, and an audit report generation unit connected in sequence.

[0132] The user terminal is used for a user to access the archive intelligent opening audit platform and collect second artificial audit data of the user for the opening audit second archive text data.

[0133] The archive intelligent opening audit platform is used to provide user login function, data upload function, archive visualization function, archive audit function, and report visualization function.

[0134] The archive intelligent opening audit engine is used to realize archive text recognition operation, archive opening permission classification operation, sensitive information detection operation, archive automatic audit operation, and audit report generation operation.

[0135] The platform initialization unit is used to build the archive intelligent opening audit platform, use artificial intelligence algorithm, construct the archive intelligent opening audit engine in the cloud data center, and connect the archive intelligent opening audit engine to the archive intelligent opening audit platform.

[0136] The archive text generation unit is used to receive a plurality of second archive files sent by the data server, use the archive intelligent opening audit engine, generate second archive text data, corresponding second sensitive information, and second sensitive information mask of each second archive file.

[0137] The automatic audit processing unit is used to use the archive intelligent opening audit engine to generate opening audit second archive text data and corresponding second automatic audit data according to all second archive text data, corresponding second sensitive information, and second sensitive information mask, and visualize in the archive intelligent opening audit platform.

[0138] The audit report generation unit is used to use the archive intelligent opening audit engine to generate a second audit report according to the second artificial audit data of the user terminal returned for the opening audit second archive text data and the corresponding second automatic audit data, and visualize in the archive intelligent opening audit platform.

[0139] The application provides a kind of based on big model's file intelligent opening auditing method and system, and the file intelligent opening auditing engine of artificial intelligence algorithm is constructed, realizes the systematized auditing process of automatic file text recognition, file opening permission classification, sensitive information detection, file automatic auditing and the generation of auditing report, significantly improve the efficiency of file auditing, shorten the auditing cycle, meet the demand of rapid processing of a large number of files;Utilize advanced file text recognition model, file opening permission classification model, sensitive information detection model, reduce misjudgment and omission, ensure the accuracy of file auditing;Through file automatic auditing model, the file data is automatically audited, improves the auditing efficiency, and assists artificial auditing;Based on big model's sensitive information detection model, has powerful semantic understanding ability, can improve the accuracy of sensitive information identification determination, realizes effective sensitive information detection means, can accurately identify and generate sensitive information mask, avoids information leakage risk;Through file intelligent opening auditing platform, realizes the online intelligent opening auditing of file, realizes the rapid auditing and sharing of file data, and uses cloud data center to adopt unified standard and specification, carries out unified management and analysis to the file of different data sources, simplifies the difficulty of data integration and exchange, avoids information island, improves the informationization degree.

[0140] The application is not limited to the above optional embodiments, and anyone can derive other various forms of products under the inspiration of the application. The above specific embodiments should not be understood as limiting the protection scope of the application, and the protection scope of the application should be defined by the claims, and the specification can be used to explain the claims.

Claims

1. A method for intelligent open access review of archives based on a large model, characterized in that: Includes the following steps: In a cloud data center, an intelligent open access review platform for archives is built. This platform utilizes artificial intelligence algorithms to construct an intelligent open access review engine for archives within the cloud data center, and then connects this engine to the intelligent open access review platform for archives. The process includes the following steps: A cloud data center is used to build an intelligent open audit framework for archives, and to set up user login modules, data upload modules, archive visualization modules, archive audit modules, and report visualization modules to obtain an intelligent open audit platform for archives; Collect several first-level archive files and several first-level manual review data, and preprocess them to obtain several preprocessed first-level archive files and several preprocessed first-level manual review data; Based on several preprocessed first archive files, an archive text recognition model is constructed using a character recognition algorithm to generate several first archive text data. The text recognition model is constructed based on the FPN-LSTM-CRF algorithm, and the text recognition model includes an image feature extraction module constructed based on the FPN algorithm, a sequence feature extraction module constructed based on the LSTM algorithm, and a recognition text label generation module constructed based on the CRF algorithm, which are connected in sequence. Based on several primary archive text data, a deep learning algorithm is used to construct an archive access permission classification model and generate the primary archive access permission classification result for each primary archive text data. The archive access permission classification model is constructed based on the LSTM-DBN algorithm, and the archive access permission classification model includes a semantic feature extraction module constructed based on the LSTM algorithm and an archive access permission classification module constructed based on the DBN algorithm, which are connected in sequence. Based on external text big data, as well as several first-file text data and their first-file open access permission classification results, a sensitive information detection model is constructed using a large model algorithm, and several first-file sensitive information is generated; The sensitive information detection model is constructed based on the RoBERTa-Transformer-CRF algorithm, and includes a word embedding module based on RoBERTa, a deep feature extraction module based on the Transformer algorithm, and a sensitive information label generation module based on the CRF algorithm, which are connected in sequence. Based on several first-level file text data, the corresponding first-level file access permission classification results, and first-level sensitive information, a file automatic review model is constructed using deep learning algorithms, and several first-level automatic review data are generated. The automatic document review model is built based on the RF-MLP algorithm, and the automatic document review model includes a key feature extraction module built based on the RF algorithm and an automatic document review module built based on the MLP algorithm, which are connected in sequence. Based on several pre-processed first manual review data and corresponding first automatic review data, a review report generation model is constructed using a deep learning algorithm; The audit report generation model is constructed based on the cGAN-MLP algorithm, and includes a generator and a discriminator both constructed based on the RNN algorithm, as well as a condition embedding module and a condition processing module constructed based on the MLP algorithm. The generator is connected to the discriminator and the condition embedding module respectively, and the condition processing module is connected to the discriminator. By integrating archival text recognition models, archival access permission classification models, sensitive information detection models, archival automatic review models, and review report generation models, an intelligent archival access review engine is built in a cloud data center, and the intelligent archival access review engine is connected to the intelligent archival access review platform. The cloud data center receives several second-file files sent by the data server, and uses the intelligent open audit engine to generate the second-file text data, the corresponding second-sensitive information, and the second-sensitive information mask for each second-file file. The cloud data center uses the intelligent open audit engine for archives to generate open audit second archive text data and corresponding second automatic audit data based on all second archive text data, corresponding second sensitive information and second sensitive information mask, and visualizes them on the intelligent open audit platform for archives. The cloud data center uses the intelligent open audit engine for archives to generate a second audit report based on the second manual audit data and the corresponding second automatic audit data of the second archive text data returned by the user terminal, and visualizes it on the intelligent open audit platform for archives.

2. The method for intelligent open access review of archives based on a large model according to claim 1, characterized in that: The aforementioned intelligent open archive review platform includes a user login module, a data upload module, an archive visualization module, an archive review module, and a report visualization module; The aforementioned intelligent archive access review engine includes an archive text recognition model, an archive access permission classification model, a sensitive information detection model, an automatic archive review model, and a review report generation model.

3. The method for intelligent open access review of archives based on a large model according to claim 2, characterized in that: The cloud data center receives several second-file files sent by the data server, and uses the intelligent open auditing engine to generate the second-file text data, corresponding second-sensitive information, and a second-sensitive information mask for each second-file file, including the following steps: The cloud data center receives several second archive files sent by the data server and uses an archive text recognition model to generate second archive text data for each second archive file. Using the archive access permission classification model, the archive access permission is classified for each second archive text data to obtain the corresponding second archive access permission classification result; Based on the classification results of the access permissions of the second archives, a sensitive information detection model is used to detect sensitive information in each second archive text data, resulting in several second sensitive information and corresponding second sensitive information masks.

4. The method for intelligent open access review of archives based on a large model according to claim 3, characterized in that: The classification result of the first archive access permission includes the first archive domain and the first access permission of the first archive text data; The classification results of the second archive open access permissions include the second archive domain and the second open access permissions of the second archive text data.

5. The method for intelligent open access review of archives based on a large model according to claim 4, characterized in that: The cloud data center, based on all secondary archive text data, corresponding secondary sensitive information, and secondary sensitive information masks, uses the intelligent open archive review engine to generate open review secondary archive text data and corresponding secondary automatic review data, and visualizes them on the intelligent open archive review platform, including the following steps: The cloud data center uses an automatic file review model to automatically review files based on several second file text data, the corresponding second file open permission classification results, and second sensitive information, and obtains the corresponding second automatic review data. Based on each second file text data and the corresponding second sensitive information mask, generate open review second file text data; The intelligent open archive review platform is used to visualize the text data of the second archive for open review.

6. The method for intelligent open access review of archives based on a large model according to claim 5, characterized in that: The cloud data center, based on the second manual review data and the corresponding second automatic review data of the open review of the second archival text data returned by the user terminal, uses the intelligent open review engine for archives to generate a second review report, which is then visualized on the intelligent open review platform for archives. This process includes the following steps: The cloud data center uses the intelligent open archive review platform to receive second manual review data for the open review of second archive text data returned by user terminals; Based on the second manual review data and the corresponding second automatic review data, the archive intelligent open review engine is used to generate a second review report; The second review report is visualized using the intelligent open review platform for archives.

7. A large-scale model-based intelligent open access review system for archives, used to implement the intelligent open access review method for archives as described in any one of claims 1-6, characterized in that: The system includes a cloud data center and several user terminals, all of which are communicatively connected to the cloud data center. The cloud data center is equipped with an intelligent open archive review platform and an intelligent open archive review engine. The cloud data center also includes a platform initialization unit, an archive text generation unit, an automatic review processing unit, and an review report generation unit connected in sequence.

Citation Information

Patent Citations

  • Intelligent identification system and method

    CN116578703A

  • Archive data management method and system and electronic equipment

    CN118277511A