Intelligent open auditing method and system for archives based on large model

Through the intelligent open archive review method and system based on large models, the problems of low efficiency, insufficient accuracy and low level of informatization in the existing technology are solved, efficient and accurate archive review and sensitive information detection are achieved, and the risk of information leakage is avoided.

CN120216743AActive Publication Date: 2025-06-27WUHAN UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510263565.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-27
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The existing archive review technology has problems such as inefficient efficiency, insufficient accuracy and low level of informatization, which is difficult to meet the need to quickly process large amounts of archives and poses a risk of information leakage.

Method used

Using intelligent open review methods and systems for archives based on large models, an intelligent open review engine for archives is built through artificial intelligence algorithms to realize automated archive text recognition, open permission classification, sensitive information detection, automatic review and report generation.

Benefits of technology

It significantly improves the efficiency of archive review, shortens the audit cycle, reduces misjudgment and misjudgment, ensures the accuracy of archive review, and avoids the risk of information leakage through sensitive information detection models, realizing rapid review and sharing of archive data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216743A_ABST
    Figure CN120216743A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of archive auditing, and discloses an archive intelligent open auditing method and system based on a large model. The method comprises the following steps: a cloud data center builds a file intelligent open auditing platform and builds a file intelligent open auditing engine; the cloud data center is used for generating second archive text data, corresponding first sensitive information and second sensitive information masks of each second archive file by using an archive intelligent open auditing engine; the cloud data center is used for generating open auditing second archive text data and corresponding second automatic auditing data by using an archive intelligent open auditing engine; and the cloud data center generates a second auditing report by using an intelligent archive open auditing engine according to the second manual auditing data and the second automatic auditing data, and visualizes the second auditing report on an intelligent archive open auditing platform. According to the invention, the problems of low efficiency, insufficient accuracy and low informatization degree in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of file auditing, and particularly relates to an intelligent open auditing method and system for files based on a large model. Background Art

[0002] With the rapid development of information technology, digitalization of files has become an irreversible trend. Digitalization not only improves the storage, retrieval, and utilization efficiency of files, but also makes file information easier to share and disseminate. File open auditing serves the open utilization of file data. Through file content auditing, it prevents the opening of sensitive information and avoids threatening national security and infringing on personal privacy. At the same time, with the continuous progress of technology, the methods and means of file auditing will also be continuously innovated and improved, providing more efficient and convenient services for the work of open utilization of file data.

[0003] The existing file auditing technologies have the following defects:

[0004] 1) Low efficiency: Traditional file auditing relies on manual operations. Facing a large amount of file data, the manual auditing speed is slow and it is difficult to meet the requirements of rapid processing.

[0005] 2) Insufficient accuracy: Manual auditing is greatly affected by subjective factors, prone to misjudgment or missed judgment, affecting the accuracy of file auditing. Moreover, the existing technologies lack effective means for detecting sensitive information, making it difficult to accurately identify and mask sensitive information, and there is a risk of information leakage.

[0006] 3) Low level of informatization: The existing technologies have a low level of informatization, making it difficult to achieve rapid auditing and sharing of file data. The file data formats are diverse and lack unified standards and specifications, resulting in difficulties in data integration and exchange. Summary of the Invention

[0007] In order to solve the problems of low efficiency, insufficient accuracy, and low level of informatization existing in the prior art, the purpose of the present invention is to provide an intelligent open auditing method and system for files based on a large model.

[0008] The technical solution adopted by the present invention is as follows:

[0009] An intelligent open auditing method for files based on a large model, comprising the following steps:

[0010] Build a file intelligent open auditing platform in a cloud data center, use artificial intelligence algorithms to build a file intelligent open auditing engine in the cloud data center, and connect the file intelligent open auditing engine to the file intelligent open auditing platform;

[0011] The cloud data center receives a number of second archive files sent by the data server, and uses the intelligent open audit engine for archives to generate the second archive text data, corresponding first sensitive information, and second sensitive information mask for each second archive file;

[0012] The cloud data center, based on all the second archive text data, corresponding first sensitive information, and second sensitive information mask, uses the intelligent open audit engine for archives to generate the open audit second archive text data and corresponding second automatic audit data, and visualizes them on the intelligent open audit platform for archives;

[0013] The cloud data center, based on the second manual audit data and corresponding second automatic audit data for the open audit second archive text data returned by the user terminal, uses the intelligent open audit engine for archives to generate a second audit report and visualizes it on the intelligent open audit platform for archives.

[0014] Furthermore, the intelligent open audit platform for archives includes a user login module, a data upload module, an archive visualization module, an archive audit module, and a report visualization module;

[0015] The intelligent open audit engine for archives includes an archive text recognition model, an archive open permission classification model, a sensitive information detection model, an archive automatic audit model, and an audit report generation model.

[0016] Furthermore, the cloud data center builds the intelligent open audit platform for archives, uses artificial intelligence algorithms to build the intelligent open audit engine for archives in the cloud data center, and connects the intelligent open audit engine for archives to the intelligent open audit platform for archives, including the following steps:

[0017] The cloud data center builds the intelligent open audit framework for archives, and sets up a user login module, a data upload module, an archive visualization module, an archive audit module, and a report visualization module to obtain the intelligent open audit platform for archives;

[0018] Collect a number of first archive files and a number of first manual audit data, and perform preprocessing to obtain a number of preprocessed first archive files and a number of preprocessed first manual audit data;

[0019] Based on a number of preprocessed first archive files, use an optical character recognition algorithm to build an archive text recognition model and generate a number of first archive text data;

[0020] Based on a number of first archive text data, use a deep learning algorithm to build an archive open permission classification model and generate a first archive open permission classification result for each first archive text data;

[0021] According to external text big data, as well as a number of first archive text data and their first archive open permission classification results, using large model algorithms, construct a sensitive information detection model and generate a number of first sensitive information;

[0022] According to a number of first archive text data, the corresponding first archive open permission classification results, and the first sensitive information, using deep learning algorithms, construct an archive automatic review model and generate a number of first automatic review data;

[0023] According to a number of preprocessed first manual review data and the corresponding first automatic review data, using deep learning algorithms, construct a review report generation model;

[0024] Integrate the archive text recognition model, the archive open permission classification model, the sensitive information detection model, the archive automatic review model, and the review report generation model, construct an archive intelligent open review engine in the cloud data center, and connect the archive intelligent open review engine to the archive intelligent open review platform.

[0025] Furthermore, the text recognition model is constructed based on the FPN-LSTM-CRF algorithm, and the text recognition model includes an image feature extraction module constructed based on the FPN algorithm, a sequence feature extraction module constructed based on the LSTM algorithm, and a recognition text label generation module constructed based on the CRF algorithm, which are connected in sequence;

[0026] The archive open permission classification model is constructed based on the LSTM-DBN algorithm, and the archive open permission classification model includes a semantic feature extraction module constructed based on the LSTM algorithm and an archive open permission classification module constructed based on the DBN algorithm, which are connected in sequence;

[0027] The sensitive information detection model is constructed based on the RoBERTa-Transfomer-CRF algorithm, and the sensitive information detection model includes a word embedding module based on RoBERTa, a deep feature extraction module constructed based on the Transfomer algorithm, and a sensitive information label generation module constructed based on the CRF algorithm, which are connected in sequence;

[0028] The archive automatic review model is constructed based on the RF-MLP algorithm, and the archive automatic review model includes a key feature extraction module constructed based on the RF algorithm and an archive automatic review module constructed based on the MLP algorithm, which are connected in sequence;

[0029] The review report generation model is constructed based on the cGAN-MLP algorithm, and the review report generation model includes a generator and a discriminator both constructed based on the RNN algorithm, as well as a conditional embedding module and a conditional processing module constructed based on the MLP algorithm. The generator is connected to the discriminator and the conditional embedding module respectively, and the conditional processing module is connected to the discriminator.

[0030] Further, the cloud data center receives a number of second archive files sent by the data server, and uses the archive intelligent open audit engine to generate the second archive text data, the corresponding first sensitive information, and the second sensitive information mask for each second archive file, including the following steps:

[0031] The cloud data center receives a number of second archive files sent by the data server and uses the archive text recognition model to generate the second archive text data for each second archive file;

[0032] Use the archive open permission classification model to classify the archive open permissions of each second archive text data to obtain the corresponding second archive open permission classification result;

[0033] According to the second archive open permission classification result, use the sensitive information detection model to detect the sensitive information of each second archive text data to obtain a number of second sensitive information and the corresponding second sensitive information mask.

[0034] Further, the first archive open permission classification result includes the first archive field and the first open permission of the first archive text data;

[0035] The second archive open permission classification result includes the second archive field and the second open permission of the second archive text data.

[0036] Further, the cloud data center uses the archive intelligent open audit engine to generate the open audit second archive text data and the corresponding second automatic audit data according to all the second archive text data, the corresponding first sensitive information, and the second sensitive information mask, and visualizes them on the archive intelligent open audit platform, including the following steps:

[0037] The cloud data center uses the archive automatic audit model to perform archive automatic audit according to a number of second archive text data, the corresponding second archive open permission classification result, and the second sensitive information to obtain the corresponding second automatic audit data;

[0038] Generate the open audit second archive text data according to each second archive text data and the corresponding second sensitive information mask;

[0039] Use the archive intelligent open audit platform to visualize the open audit second archive text data.

[0040] Further, the cloud data center uses the archive intelligent open audit engine to generate a second audit report according to the second manual audit data and the corresponding second automatic audit data of the open audit second archive text data returned by the user terminal, and visualizes it on the archive intelligent open audit platform, including the following steps:

[0041] A cloud data center uses an intelligent open audit platform for archives to receive the second manual audit data of the second archive text data returned by a user terminal.

[0042] According to the second manual audit data and the corresponding second automatic audit data, use the intelligent open audit engine for archives to generate a second audit report.

[0043] Use the intelligent open audit platform for archives to visualize the second audit report.

[0044] An intelligent open audit system for archives based on a large model is used to implement the intelligent open audit method for archives. The system includes a cloud data center and a number of user terminals. The number of user terminals are all communicatively connected to the cloud data center. The cloud data center is provided with an intelligent open audit platform for archives and an intelligent open audit engine for archives. And the cloud data center includes a platform initialization unit, an archive text generation unit, an automatic audit processing unit, and an audit report generation unit that are connected in sequence.

[0045] The beneficial effects of the present invention are:

[0046] An intelligent open audit method and system for archives based on a large model provided by the present invention realizes a systematic audit process of automatic archive text recognition, archive open permission classification, sensitive information detection, archive automatic audit, and audit report generation through an intelligent open audit engine constructed by artificial intelligence algorithms, significantly improving the efficiency of archive audit, shortening the audit cycle, and meeting the need for quickly processing a large number of archives; using advanced archive text recognition models, archive open permission classification models, and sensitive information detection models to reduce misjudgments and missed judgments and ensure the accuracy of archive audit; through an archive automatic audit model, automatically auditing archive data, improving the audit efficiency, and assisting manual audit; a sensitive information detection model based on a large model has strong semantic understanding ability, can improve the accuracy of sensitive information identification and determination, realizes an effective sensitive information detection means, can accurately identify and generate sensitive information masks, and avoids the risk of information leakage; through the intelligent open audit platform for archives, realizes online intelligent open audit of archives, realizes rapid audit and sharing of archive data, and uses the cloud data center to adopt unified standards and specifications to uniformly manage and analyze archive files from different data sources, simplifies the difficulty of data integration and exchange, avoids information islands, and improves the degree of informatization.

[0047] Other beneficial effects of the present invention will be further described in the specific implementation manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a flowchart of the intelligent open audit method for archives based on a large model in the present invention.

[0049] Figure 2 It is the structural block diagram of the intelligent open audit system for archives based on large models in the present invention. Specific implementation manners

[0050] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments.

[0051] Embodiment 1:

[0052] As Figure 1 shown, this embodiment provides an intelligent open audit method for archives based on large models, including the following steps:

[0053] S1: In the cloud data center, build an intelligent open audit platform for archives, use artificial intelligence algorithms, build an intelligent open audit engine for archives in the cloud data center, and connect the intelligent open audit engine for archives to the intelligent open audit platform for archives, including the following steps:

[0054] S1-1: In the cloud data center, build an intelligent open audit framework for archives, and set up a user login module, a data upload module, an archive visualization module, an archive audit module, and a report visualization module to obtain an intelligent open audit platform for archives;

[0055] The intelligent open audit platform for archives includes a user login module, a data upload module, an archive visualization module, an archive audit module, and a report visualization module;

[0056] The user login module is used to receive the login data input by the user; the data upload module is used to upload the login data to the cloud data center for user login; the archive visualization module is used to visualize the second archive text data for open audit; the archive audit module is used to receive the second manual audit data of the user for the second archive text data for open audit; the report visualization module is used to visualize the second audit report;

[0057] S1-2: Collect a number of first archive files and a number of first manual audit data, and perform preprocessing to obtain a number of preprocessed first archive files and a number of preprocessed first manual audit data;

[0058] The preprocessing includes data cleaning, format conversion, normalization processing, etc., to improve the data quality and provide data support for subsequent model training;

[0059] S1-3: According to a number of preprocessed first archive files, use an optical character recognition algorithm to build an archive text recognition model and generate a number of first archive text data;

[0060] The text recognition model is constructed based on the Feature Pyramid Networks (FPN)-Long Short-Term Memory (LSTM)-Conditional Random Field (CRF) algorithm. The text recognition model includes an image feature extraction module constructed based on the FPN algorithm, a sequence feature extraction module constructed based on the LSTM algorithm, and a recognized text label generation module constructed based on the CRF algorithm, which are connected in sequence;

[0061] The image feature extraction module uses the top-down and bottom-up paths of FPN to effectively extract image features at different scales, ensuring that various details in the archival image can be accurately detected, realizing the fusion of features at different levels, and enhancing the feature expression ability; the LSTM in the sequence feature extraction module is good at processing sequence data and can capture the long-term and short-term dependencies in the text sequence, providing rich sequence features for subsequent text label generation; the recognized text label generation module uses the constraint conditions of CRF to optimize the generated labels, ensuring the rationality and accuracy of the labels, and ensuring that the generated text label sequence is more reasonable in grammar and semantics;

[0062] The entire text recognition model realizes the complete process from image feature extraction to text label generation through the organic combination of FPN, LSTM, and CRF. FPN ensures the comprehensiveness and accuracy of image features; LSTM captures the temporal and context information of the text sequence; CRF optimizes and constrains the generated labels. This structure not only improves the accuracy of text recognition but also enhances the adaptability of the model to various complex scenarios;

[0063] S1-4: According to a number of first archival text data, use deep learning algorithms to construct an archival access permission classification model and generate the first archival access permission classification results for each first archival text data;

[0064] The archival access permission classification model is constructed based on the LSTM-Deep Belief Network (DBN) algorithm. The archival access permission classification model includes a semantic feature extraction module constructed based on the LSTM algorithm and an archival access permission classification module constructed based on the DBN algorithm, which are connected in sequence;

[0065] The LSTM in the semantic feature extraction module can capture the long-distance dependencies in the text, enabling the model to better understand the semantics and context of the archival text data, thereby improving the accuracy of classification; the deep structure of the DBN in the archival access permission classification module enables the model to learn more complex and abstract feature representations, thereby improving the accuracy and robustness of archival access permission classification;

[0066] By organically combining LSTM and DBN, the classification model for file opening permissions realizes the complete process from semantic feature extraction to classification decision-making. The LSTM module is responsible for capturing the deep semantic features and temporal information in the file text, providing rich feature representations for subsequent classification; the DBN module uses its deep structure and pre-training mechanism to further refine and abstract these features and make accurate classification decisions. This structure not only improves the accuracy of file opening permission classification but also enhances the adaptability and generalization ability of the model to complex file texts;

[0067] S1-5: According to external text big data, as well as a number of first file text data and their first file opening permission classification results, use the large model algorithm to construct a sensitive information detection model and generate a number of first sensitive information;

[0068] The sensitive information detection model is constructed based on the Robustly optimized BERT approach (RoBERTa)-Transformer-CRF algorithm, and the sensitive information detection model includes a word embedding module based on RoBERTa, a deep feature extraction module constructed based on the Transformer algorithm, and a sensitive information label generation module constructed based on the CRF algorithm connected in sequence;

[0069] The word embedding module improves and adjusts the Bidirectional Encoder Representations from Transformers (BERT) model based on Transformer to enhance performance. This module converts the input text into word vector representations, and each word vector contains the semantic information, context information, and position information of the word, improving the understanding ability of polysemous words and complex sentences, and retaining the position information of the words in the sentence, which helps the model understand the order and structure of the words; the Transformer in the deep feature extraction module uses the self-attention mechanism to capture the long-distance dependence relationships in the text and generate a feature representation containing global information for each word. The self-attention mechanism effectively captures the long-distance dependence relationships in the text and improves the detection ability for complex sensitive information. The multi-layer stacked structure enables the model to learn multi-level and multi-angle text features and enhances the expression ability of the features; the sensitive information label generation module introduces constraint conditions, such as the legality constraints between labels, to further optimize the generation of the label sequence and has a certain robustness to noise data and abnormal situations;

[0070] The sensitive information detection model realizes the complete process from word embedding to the generation of sensitive information labels through the organic combination of RoBERTa, Transformer, and CRF. The RoBERTa word embedding module provides rich semantic representations and context information; the Transformer deep feature extraction module further captures deep features and long-distance dependencies in the text; the CRF sensitive information label generation module uses sequence labeling technology to generate the globally optimal sensitive information label sequence. This structure not only improves the accuracy of sensitive information detection but also enhances the model's adaptability and robustness to complex texts and noisy data;

[0071] According to external text big data, as well as several first archive text data and their first archive open permission classification results, use large model algorithms to construct a sensitive information detection model and generate several first sensitive information, including the following steps:

[0072] S1-5-1: Collect external text big data from sources such as the Internet, databases, and file systems. These data cover various topics, fields, and language styles, and preprocess the text big data to obtain several preprocessed text data; perform operations such as cleaning, deduplication, word segmentation, and stop word removal on the collected text big data to eliminate noise and redundant information and improve data quality;

[0073] S1-5-2: Use the RoBERTa-Transfomer-CRF algorithm to construct an initial sensitive information detection model;

[0074] S1-5-3: Use several preprocessed text data to pre-train the initial sensitive information detection model to obtain a pre-trained sensitive information detection model, enabling the model to have strong semantic understanding and text processing capabilities; use a large amount of preprocessed text data to pre-train the model so that the model learns general language representations and features. Pre-training enables the model to have strong generalization capabilities and be able to better adapt to different text data and tasks;

[0075] S1-5-4: According to several first archive text data and their first archive open permission classification results, fine-tune the pre-trained sensitive information detection model to obtain the final sensitive information detection model; fine-tuning makes the model more adaptable to specific sensitive information detection tasks, improving the accuracy and effectiveness of detection. By using the first archive text data and their classification results, the model can inherit and utilize existing knowledge to further optimize the detection performance;

[0076] S1-6: According to several first archive text data, the corresponding first archive open permission classification results, and the first sensitive information, use deep learning algorithms to construct an archive automatic review model and generate several first automatic review data;

[0077] The file automatic review model is constructed based on the Random Forest (RF)-Multilayer Perceptron (MLP) algorithm, and the file automatic review model includes a key feature extraction module constructed based on the RF algorithm and a file automatic review module constructed based on the MLP algorithm that are connected in sequence;

[0078] The key feature extraction module improves the prediction accuracy by constructing multiple decision trees and aggregating their predictions. Each decision tree is trained on a random subset of the training set and can evaluate the importance of each feature for the prediction result, thereby identifying the key features. By extracting the key features, the dimension of the data is reduced, the complexity of the subsequent model is simplified, and the extraction of key features helps to remove noise and irrelevant information, improving the prediction accuracy of the model; The file automatic review module generates automatic review labels by learning the weights and biases between the input features and the output labels, and can adjust the number of layers and neurons of the MLP according to needs to adapt to different complexity requirements;

[0079] By organically combining RF and MLP, the file automatic review model realizes the complete process from key feature extraction to file automatic review. The RF module is responsible for screening out the most important features for file automatic review from a large number of features, reducing the data complexity and improving the interpretability of the model; The MLP module then uses its powerful non-linear modeling ability to generate accurate access permissions based on the selected key features. This structure not only improves the accuracy and efficiency of file automatic review, but also enhances the adaptability and interpretability of the model to complex scenarios;

[0080] S1-7: According to a number of preprocessed first manual review data and corresponding first automatic review data, use a deep learning algorithm to construct a review report generation model;

[0081] The review report generation model is constructed based on the Conditional Generative Adversarial Network (cGAN)-MLP algorithm, and the review report generation model includes a generator and a discriminator both constructed based on the Recurrent Neural Network (RNN) algorithm, and a conditional embedding module and a conditional processing module constructed based on the MLP algorithm. The generator is respectively connected to the discriminator and the conditional embedding module, and the conditional processing module is connected to the discriminator;

[0082] The conditional embedding module uses a multi-layer perceptron (MLP) to convert the characteristics of the classified results of the file opening permissions, the characteristics of the manual review data, and the text data of the files under open review into embedding vectors that can be understood by the generator. This embedding vector serves as conditional information to guide the generation process of the generator. The MLP algorithm can effectively handle non-linear relationships and convert the characteristics into useful embedding representations. The conditional embedding module ensures that the generator can take into account the specific decision context when generating the output, improving the relevance and accuracy of the report generation process. The generator receives the embedding vector from the conditional embedding module and other relevant information, and uses the RNN algorithm to generate an output that matches the input conditions, namely the review report. The generator attempts to generate data that is realistic enough to deceive the discriminator. The generator can generate customized outputs based on the decision information, improving the flexibility and adaptability of the model. The discriminator receives the output from the generator and the output from the conditional processing module, and uses the RNN algorithm to determine whether the review report generated by the generator is realistic enough, that is, whether it conforms to the actual situation of the review report. The discriminator guides the training process of the generator through feedback signals. The discriminator improves the quality and authenticity of the data generated by the generator through adversarial training. The adversarial process helps to improve the overall performance of the model and the consistency of the output. The conditional processing module receives part of the output from the generator and the output from the conditional embedding module, and uses the MLP algorithm to process this information to provide additional conditional information for the discriminator, helping the discriminator better understand the context of the generator's output. The conditional processing module enhances the evaluation ability of the execution model for the generator's output, improving the reliability of the entire model. By combining conditional information, the matching degree between the generator's output and the actual application scenario is improved.

[0083] According to a number of preprocessed first manual review data and corresponding first automatic review data, use deep learning algorithms to construct a review report generation model, including the following steps:

[0084] S1-7-1: Set a corresponding true report for each preprocessed first manual review data. The true report serves as the standard or target when training the model.

[0085] S1-7-2: Use the cGAN-MLP algorithm to construct an initial review report generation model.

[0086] S1-7-3: Combine the first loss function of the initial generator and the second loss function of the initial discriminator of the initial review report generation model to obtain a comprehensive loss function; ensure that the generator and the discriminator promote each other during the training process and improve the performance of the overall model.

[0087] S1-7-4: Extract the first manual review data characteristics of the preprocessed first manual review data and the first automatic review data characteristics of the first automatic review data.

[0088] S1-7-5: Use the conditional information embedder of the initial audit report generation model to perform conditional embedding on the first manual audit data features and the first automatic audit data features to obtain a number of first conditional embedding features; improve the ability of the generator to generate reports under specific conditions;

[0089] S1-7-6: Train the initial generator of the initial audit report generation model according to a number of first conditional embedding features to obtain a number of generated reports;

[0090] S1-7-7: Use the conditional information processor to perform conditional information processing on the real report and the corresponding generated report to obtain the first conditional information; ensure that the discriminator can be trained based on accurate conditional information;

[0091] S1-7-8: Train the initial discriminator according to the real report, the corresponding generated report, and the first conditional information to obtain a number of first data discrimination results; improve the ability of the discriminator to identify real and generated data, thereby promoting the generator to generate higher-quality reports;

[0092] S1-7-9: According to each generated report and the corresponding first data discrimination result, use the comprehensive loss function to obtain the first comprehensive loss value during the training process; evaluate the training progress of the model and provide a basis for whether to continue training;

[0093] S1-7-10: If the first loss value is lower than the first loss value threshold, output the optimized generator and the optimized discriminator; otherwise, continue with the optimization training;

[0094] S1-7-11: Integrate the optimized generator, the optimized discriminator, the conditional information embedder, and the conditional information processor to obtain the final audit report generation model;

[0095] S1-8: Integrate the archive text recognition model, the archive open permission classification model, the sensitive information detection model, the archive automatic audit model, and the audit report generation model to build an archive intelligent open audit engine in the cloud data center, and connect the archive intelligent open audit engine to the archive intelligent open audit platform;

[0096] The archive intelligent open audit engine includes an archive text recognition model, an archive open permission classification model, a sensitive information detection model, an archive automatic audit model, and an audit report generation model;

[0097] S2: The cloud data center receives a number of second archive files sent by the data server and uses the archive intelligent open audit engine to generate the second archive text data, the corresponding first sensitive information, and the second sensitive information mask of each second archive file, including the following steps:

[0098] S2-1: The cloud data center receives a number of second archive files sent by the data server and uses an archive text recognition model to generate second archive text data for each second archive file, including the following steps:

[0099] S2-1-1: The cloud data center receives a number of second archive files sent by the data server, converts the second archive files into second archive images, and inputs the number of second archive images into the archive text recognition model;

[0100] S2-1-2: Use the image feature extraction module of the archive text recognition model to extract the second image features of the second archive images;

[0101] S2-1-3: Use the sequence feature extraction module of the archive text recognition model to generate corresponding second sequence features according to the second image features;

[0102] S2-1-4: Use the recognition text label generation module of the archive text recognition model to generate corresponding second archive text data according to the second sequence features;

[0103] S2-1-5: Traverse all the second archive files to obtain the second archive text data of each second archive file;

[0104] S2-2: Use the archive open permission classification model to classify the archive open permissions of each second archive text data to obtain the corresponding second archive open permission classification results, including the following steps:

[0105] S2-2-1: Input the second archive text data output by the archive text recognition model into the archive open permission classification model;

[0106] S2-2-2: Use the semantic feature extraction module of the archive open permission classification model to extract the second semantic features of the second archive text data;

[0107] S2-2-3: Use the archive open permission classification module of the archive open permission classification model to classify the archive open permissions according to the second semantic features to obtain the corresponding second archive open permission classification results;

[0108] S2-2-4: Traverse all the second archive text data to obtain the second archive open permission classification results of each second archive text data;

[0109] The second archive open permission classification results include the second archive field and the second open permission of the second archive text data;

[0110] S2-3: According to the classification results of the second file access permissions, use the sensitive information detection model to detect sensitive information in each second file text data, obtaining a number of second sensitive information and corresponding second sensitive information masks, including the following steps:

[0111] S2-3-1: Use the word embedding module of the sensitive information detection model to perform word embedding on the second file text data to obtain corresponding second word vectors;

[0112] S2-3-2: Use the deep feature extraction module of the sensitive information detection model to extract the second file access permission classification result features of the second file access permission classification results and the second word vector features of the second word vectors;

[0113] S2-3-3: Use the sensitive information label generation module of the sensitive information detection model to generate sensitive information labels based on the second file access permission classification result features and the second word vector features, obtaining corresponding a number of second sensitive information;

[0114] S2-3-4: Generate corresponding a number of second sensitive information masks according to different access permissions and the second position data of the second sensitive information in the second file text data;

[0115] S2-3-5: Traverse all second file text data to obtain a number of second sensitive information and corresponding second sensitive information masks for each second file text data;

[0116] S3: The cloud data center, based on all second file text data, corresponding first sensitive information, and second sensitive information masks, uses the file intelligent open audit engine to generate open audit second file text data and corresponding second automatic audit data, and visualizes them on the file intelligent open audit platform, including the following steps:

[0117] S3-1: The cloud data center, based on a number of second file text data, corresponding second file access permission classification results, and second sensitive information, uses the file automatic audit model to perform file automatic audit to obtain corresponding second automatic audit data, including the following steps:

[0118] S3-1-1: The cloud data center inputs each second file text data, corresponding second file access permission classification result, and second sensitive information into the file automatic audit model;

[0119] S3-1-2: Use the key feature extraction module of the file automatic audit model to extract a number of second key features of each second file text data, corresponding second file access permission classification result, and second sensitive information;

[0120] S3-1-3: Use the file automatic review module of the file automatic review model to perform automatic review of files, and obtain the corresponding second automatic review data;

[0121] S3-2: Generate the open review second file text data according to each second file text data and the corresponding second sensitive information mask;

[0122] S3-4: Use the file intelligent open review platform to visualize the open review second file text data;

[0123] S4: The cloud data center, according to the second manual review data and the corresponding second automatic review data of the open review second file text data returned by the user terminal, uses the file intelligent open review engine to generate a second review report and visualize it on the file intelligent open review platform, including the following steps:

[0124] S4-1: The cloud data center uses the file intelligent open review platform to receive the second manual review data of the open review second file text data returned by the user terminal;

[0125] S4-2: According to the second manual review data and the corresponding second automatic review data, use the file intelligent open review engine to generate a second review report, including the following steps:

[0126] S4-2-1: Extract the second manual review data features of the second manual review data and the second automatic review data features of the corresponding second automatic review data;

[0127] S4-2-2: Use the conditional embedding module of the review report generation model to perform conditional embedding on the second manual review data features and the second automatic review data features to obtain the second conditional embedding features;

[0128] S4-2-3: Use the generator of the review report generation model to generate a review report according to the second conditional embedding features to obtain the corresponding second review report;

[0129] S4-3: Use the file intelligent open review platform to visualize the second review report.

[0130] Embodiment 2:

[0131] As Figure 2As shown in the figure, this embodiment provides an intelligent open audit system for archives based on a large model, which is used to implement an intelligent open audit method for archives. The system includes a cloud data center and a number of user terminals. All the user terminals are communicatively connected to the cloud data center. The cloud data center is provided with an intelligent open audit platform for archives and an intelligent open audit engine for archives. The cloud data center includes a platform initialization unit, an archive text generation unit, an automatic audit processing unit, and an audit report generation unit that are connected in sequence;

[0132] User terminals are used for users to access the intelligent open audit platform for archives and collect second manual audit data of users for the second archive text data of open audit;

[0133] The intelligent open audit platform for archives is used to provide user login functions, data upload functions, archive visualization functions, archive audit functions, and report visualization functions;

[0134] The intelligent open audit engine for archives is used to implement archive text recognition operations, archive open permission classification operations, sensitive information detection operations, archive automatic audit operations, and audit report generation operations;

[0135] The platform initialization unit is used to build the intelligent open audit platform for archives, use artificial intelligence algorithms to build an intelligent open audit engine for archives in the cloud data center, and connect the intelligent open audit engine for archives to the intelligent open audit platform for archives;

[0136] The archive text generation unit is used to receive a number of second archive files sent by the data server, use the intelligent open audit engine for archives to generate second archive text data, corresponding first sensitive information, and second sensitive information masks for each second archive file;

[0137] The automatic audit processing unit is used to generate second archive text data for open audit and corresponding second automatic audit data according to all the second archive text data, corresponding first sensitive information, and second sensitive information masks, using the intelligent open audit engine for archives, and visualize them on the intelligent open audit platform for archives;

[0138] The audit report generation unit is used to generate a second audit report according to the second manual audit data of the second archive text data for open audit returned by the user terminal and the corresponding second automatic audit data, using the intelligent open audit engine for archives, and visualize it on the intelligent open audit platform for archives.

[0139] An intelligent open audit method and system for archives based on large models provided by the present invention realizes a systematic audit process of automated archive text recognition, archive open permission classification, sensitive information detection, automatic archive audit, and audit report generation through an archive intelligent open audit engine constructed by artificial intelligence algorithms, significantly improving the efficiency of archive audits, shortening the audit cycle, and meeting the need to quickly process a large number of archives; by using advanced archive text recognition models, archive open permission classification models, and sensitive information detection models, it reduces misjudgments and missed judgments to ensure the accuracy of archive audits; through an automatic archive audit model, it automatically audits archive data, improves audit efficiency, and assists manual audits; the sensitive information detection model based on large models has strong semantic understanding capabilities, can improve the accuracy of sensitive information recognition and determination, realizes effective sensitive information detection means, can accurately identify and generate sensitive information masks, and avoids the risk of information leakage; through the archive intelligent open audit platform, it realizes the online intelligent open audit of archives, realizes the rapid audit and sharing of archive data, and uses the cloud data center to uniformly manage and analyze archive files from different data sources according to unified standards and specifications, simplifies the difficulty of data integration and exchange, avoids information silos, and improves the degree of informatization.

[0140] The present invention is not limited to the above optional implementation manners, and anyone can obtain other various forms of products under the inspiration of the present invention. The above specific implementation manners should not be construed as limiting the protection scope of the present invention, and the protection scope of the present invention should be defined by the claims, and the specification can be used to interpret the claims.

Claims

1. A method for intelligent file opening and review based on a large model, characterized by: The steps include: Cloud data center, build an intelligent open archive review platform, use artificial intelligence algorithms to build an intelligent open archive review engine in the cloud data center, and connect the intelligent open archive review engine to the intelligent open archive review platform; The cloud data center receives a plurality of second archive files sent by the data server, and uses the archive intelligent open audit engine to generate second archive text data, corresponding first sensitive information, and second sensitive information mask for each second archive file; The cloud data center generates the open review second archive text data and the corresponding second automatic review data using the archive intelligent open review engine according to all the second archive text data, the corresponding first sensitive information and the second sensitive information mask, and visualizes them on the archive intelligent open review platform; The cloud data center uses the archive intelligent open review engine to generate a second review report based on the second manual review data and the corresponding second automatic review data of the second archive text data for open review returned by the user terminal, and visualizes it on the archive intelligent open review platform.

2. According to claim 1, a method for intelligent file opening and review based on a large model is characterized by: The archive intelligent open review platform includes a user login module, a data upload module, an archive visualization module, an archive review module and a report visualization module; The archive intelligent open audit engine includes an archive text recognition model, an archive open permission classification model, a sensitive information detection model, an archive automatic audit model and an audit report generation model.

3. According to claim 2, a method for intelligent file opening and review based on a large model is characterized by: The cloud data center builds an intelligent open archive review platform. Using artificial intelligence algorithms, an intelligent open archive review engine is built in the cloud data center, and the intelligent open archive review engine is connected to the intelligent open archive review platform, including the following steps: The cloud data center builds an intelligent open archive review framework, and sets up a user login module, a data upload module, an archive visualization module, an archive review module, and a report visualization module to obtain an intelligent open archive review platform; Collecting a number of first archive files and a number of first manual review data, and preprocessing them to obtain a number of preprocessed first archive files and a number of preprocessed first manual review data; According to the plurality of pre-processed first archive files, using a text recognition algorithm, an archive text recognition model is constructed to generate a plurality of first archive text data; Based on a number of first archive text data, using a deep learning algorithm, a file open authority classification model is constructed, and a first archive open authority classification result for each first archive text data is generated; Based on external text big data, as well as a number of first archive text data and first archive open permission classification results, a large model algorithm is used to build a sensitive information detection model and generate a number of first sensitive information; According to a number of first archive text data, the corresponding first archive open permission classification results and the first sensitive information, using a deep learning algorithm, construct an archive automatic review model and generate a number of first automatic review data; Based on the first manual audit data and the corresponding first automatic audit data after the preprocessing, using the deep learning algorithm, a model for generating an audit report is constructed; Integrate the archive text recognition model, archive open permission classification model, sensitive information detection model, archive automatic review model and review report generation model, build an archive intelligent open review engine in the cloud data center, and connect the archive intelligent open review engine to the archive intelligent open review platform.

4. According to claim 3, a large model-based intelligent file opening and review method is characterized by: The text recognition model is constructed based on the FPN-LSTM-CRF algorithm, and the text recognition model includes an image feature extraction module constructed based on the FPN algorithm, a sequence feature extraction module constructed based on the LSTM algorithm, and a recognition text label generation module constructed based on the CRF algorithm, which are connected in sequence; The archive open authority classification model is constructed based on the LSTM-DBN algorithm, and the archive open authority classification model includes a semantic feature extraction module constructed based on the LSTM algorithm and an archive open authority classification module constructed based on the DBN algorithm, which are connected in sequence; The sensitive information detection model is constructed based on the RoBERTa-Transfomer-CRF algorithm, and the sensitive information detection model includes a word embedding module based on RoBERTa, a deep feature extraction module based on the Transfomer algorithm, and a sensitive information label generation module based on the CRF algorithm, which are connected in sequence; The automatic file review model is constructed based on the RF-MLP algorithm, and the automatic file review model includes a key feature extraction module constructed based on the RF algorithm and an automatic file review module constructed based on the MLP algorithm, which are connected in sequence; The audit report generation model is constructed based on the cGAN-MLP algorithm, and the audit report generation model includes a generator and a discriminator both constructed based on the RNN algorithm, and a conditional embedding module and a conditional processing module constructed based on the MLP algorithm. The generator is connected to the discriminator and the conditional embedding module respectively, and the conditional processing module is connected to the discriminator.

5. According to claim 4, a method for intelligent file opening and review based on a large model is characterized in that: The cloud data center receives a plurality of second archive files sent by the data server, and uses the archive intelligent open audit engine to generate second archive text data, corresponding first sensitive information, and second sensitive information mask for each second archive file, including the following steps: The cloud data center receives a plurality of second archive files sent by the data server, and generates second archive text data for each second archive file using an archive text recognition model; Using the archive open permission classification model, classify the archive open permission of each second archive text data to obtain the corresponding second archive open permission classification result; According to the classification result of the second file open permission, the sensitive information detection model is used to perform sensitive information detection on each second file text data to obtain a number of second sensitive information and corresponding second sensitive information masks.

6. According to claim 5, a method for intelligent file opening and review based on a large model is characterized in that: The first file opening authority classification result includes the first file field and the first opening authority of the first file text data; The second archive opening authority classification result includes the second archive field and the second opening authority of the second archive text data.

7. The method for intelligent file opening and review based on a large model according to claim 6 is characterized by: The cloud data center generates the second archive text data for open review and the corresponding second automatic review data using the archive intelligent open review engine according to all the second archive text data, the corresponding first sensitive information and the second sensitive information mask, and visualizes them on the archive intelligent open review platform, including the following steps: The cloud data center uses the archive automatic review model to perform an automatic review of the archive according to the plurality of second archive text data, the corresponding second archive open permission classification result, and the second sensitive information, and obtains corresponding second automatic review data; Generate open review second archive text data according to each second archive text data and the corresponding second sensitive information mask; Use the archive intelligent open review platform to visualize the open review second archive text data.

8. The method for intelligent file opening and review based on a large model according to claim 7 is characterized by: The cloud data center generates a second audit report using the archive intelligent open audit engine according to the second manual audit data and the corresponding second automatic audit data of the open audit second archive text data returned by the user terminal, and visualizes the report on the archive intelligent open audit platform, including the following steps: The cloud data center uses the archive intelligent open review platform to receive the second manual review data for the open review second archive text data returned by the user terminal; Generate a second audit report using the archive intelligent open audit engine based on the second manual audit data and the corresponding second automatic audit data; Use the archive intelligent open audit platform to visualize the second audit report.

9. A file intelligent opening and review system based on a large model, used to implement the file intelligent opening and review method as claimed in any one of claims 1 to 8, characterized in that: The system includes a cloud data center and several user terminals, and the several user terminals are all communicatively connected to the cloud data center. The cloud data center is provided with an archive intelligent open audit platform and an archive intelligent open audit engine, and the cloud data center includes a platform initialization unit, an archive text generation unit, an automatic audit processing unit and an audit report generation unit which are connected in sequence.

Citation Information

Patent Citations

  • Resident health record opening and supervision system and method

    CN114093481A

  • Intelligent identification system and method

    CN116578703A

  • User behavior auditing management system

    CN117132226A

  • Content auditing method and system based on intelligent process automation technology

    CN117729360A

  • Archive data management method and system and electronic equipment

    CN118277511A