An archive management system and an archive classification coding method thereof

By introducing intelligent mobile shelving, NFC tags, and the RoBERTa-wwm-ext classification model, combined with structured numbering, the inefficiency of storage, retrieval, and coding in traditional record management has been solved, achieving high efficiency, security, and adaptability in record management and meeting the needs of modern record management.

CN120218543BActive Publication Date: 2025-12-23SHAANXI HUISHENG SPACE-TIME INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510353748.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-12-23
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

Traditional record management suffers from inefficiency, poor security, and insufficient adaptability in record access, electronic retrieval, and record coding, failing to meet the modern society's needs for rapid access, accurate management, and long-term preservation of record information.

Method used

By employing intelligent mobile shelving, NFC tag technology, a RoBERTa-wwm-ext-based archival classification model, and a structured numbering system, combined with archival classification types, storage locations, and unique identifiers, the system enables rapid access, electronic retrieval, and efficient management of archives.

Benefits of technology

It has improved the automation level and ease of operation of record management, ensured the accuracy and adaptability of record classification, provided efficient electronic record access services, ensured the security of sensitive records through hierarchical access control, generated detailed access records and data analysis reports, and guaranteed the long-term preservation of records.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218543B_ABST
    Figure CN120218543B_ABST
Patent Text Reader

Abstract

The application relates to an archive management system and an archive classification coding method thereof, and belongs to the technical field of file management; the archive management system is provided with an archive registration end provided with an archive scanning module and an archive coding module, an archive storage library provided with a plurality of archive storage cabinets, an archive cloud end and an archive destruction module, and an archive browsing end provided with an electronic archive browsing module, a physical archive browsing module and a browsing record module, so that the automation level and operation convenience of archive management are greatly improved; secondly, the archive scanning module is combined with an archive classification model based on RoBERTa-wwm-ext, the accuracy of archive classification is improved, in addition, the archive coding module adopts a structured numbering mode, is combined with archive classification, a storage position and a unique identifier, and generates an archive number with high readability and uniqueness, so that archive retrieval and management are facilitated; through real-time monitoring and adjustment of environmental parameters of an archive room, the environment monitoring module can effectively prevent archive damage caused by environmental factors.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of file management, in particular to an archive management system and an archive classification coding method thereof. BACKGROUND

[0002] In today's information age, archives serve as the carrier of important information, decision-making basis and historical data for various organizations and institutions. The effectiveness and efficiency of their management play a crucial role in the operation, development and compliance requirements of organizations. However, traditional archive management models have exposed many defects and shortcomings in the face of growing archive quantities and complex management needs.

[0003] Archive access dilemma: Traditional archive management mainly relies on manual archive access operations. In large archives or document rooms, archives are usually stored in dense shelves or ordinary shelves. When new archives need to be stored, administrators need to manually find suitable spaces and place archive boxes. This process relies on the administrator's familiarity with the layout of the archive and personal experience, and is prone to misplacement or inaccurate storage location records. When retrieving archives, administrators first need to determine the approximate location of the archives based on archive indexes or retrieval requests, and then search for them one by one among numerous shelves. Especially for archives that are old or frequently borrowed, the entire retrieval process can take a lot of time, seriously affecting work efficiency.

[0004] Electronic retrieval lag: With the widespread popularity of digital technology, electronic documents have shown great advantages in information transmission and sharing. However, traditional archive management has made slow progress in electronic retrieval. Most archives still exist in paper form, and even if some archives have been digitized, the retrieval method is still cumbersome. Usually, the retriever needs to submit a written application to the archive management department, stating the purpose of retrieval, the scope of archives needed, etc. Then wait for the administrator to find and extract the corresponding electronic archive files in the local electronic database or storage server. If it involves cross-department or cross-regional retrieval needs, it also needs to transfer files through internal network transmission or mail storage media, which not only easily causes information leakage risk, but also cannot realize real-time, multi-person online electronic retrieval. For example, when conducting an audit, different audit teams in different regions need to simultaneously retrieve the financial archives of the company headquarters and branch offices. The traditional electronic retrieval method cannot meet the needs of instantaneity and collaboration, seriously hindering the efficient development of business.

[0005] Disorder and inefficiency of file coding: File coding is an important link in the file management system, and its purpose is to give each file a unique identification code to facilitate the classification, storage, retrieval and management of files. However, there are many problems in the file coding work in traditional file management. On the one hand, there is a lack of unified coding standards and specifications, and the file coding rules of different units, different industries or even different departments within the same unit are different. For example, some units use the coding method of year + department + serial number, while some use the coding method of file type + storage period + number, which makes it difficult for files to exchange and share across departments or agencies. On the other hand, traditional file coding is mostly operated manually, which is prone to coding errors, repeated coding and delayed coding updates. With the continuous increase in the number of files and the continuous development of business, the classification and management requirements of files are also changing, and manual coding is difficult to adapt to this dynamic management requirement, resulting in low efficiency of file retrieval and increased management costs. For example, in the personnel file management of some large enterprises, due to the frequent occurrence of employee employment, resignation and job transfer, the coding of personnel files needs to be continuously adjusted and updated, and the manual coding method is difficult to ensure the accuracy and timeliness of the coding, thereby affecting the efficient development of human resource management work.

[0006] In summary, the traditional file management has many problems in file access, electronic access, and file coding, which has seriously restricted the efficiency, quality and security of file management work, and cannot meet the needs of modern society for fast access to file information, accurate management and long-term preservation. Therefore, it is of great practical significance and broad application prospect to develop a file management system with convenient file access, electronic access and scientific file coding function. SUMMARY

[0007] The present application provides a file management system and a file classification coding method thereof, which solves many problems existing in the prior art in file access, electronic access and file coding.

[0008] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:

[0009] In a first aspect, the present application provides a file management system, comprising: a file registration end, a file storage library and a file access end;

[0010] The file registration end is provided with: a file scanning module, a file coding module;

[0011] The file storage library is an intelligent compact cabinet, which is provided with: a plurality of file storage cabinets, a file cloud and a file destruction module;

[0012] The file access end is provided with: an electronic file access module, a physical file access module and an access record module;

[0013] The archive scanning module performs optical character recognition on the entity archives that need to be stored, identifies the archive content and generates editable electronic archives;

[0014] The archive repository quickly allocates storage locations for newly stored entity archives and prints storage barcodes;

[0015] The archive coding module generates an archive number for the scanned electronic archive, enters the archive number into the electronic archive text, and simultaneously produces an NFC tag for the entity archive, writing the archive code into the NFC tag;

[0016] The entity archive with the pasted NFC tag and storage barcode is received by the archive storage cabinet, which is unlocked by recognizing the NFC tag or storage barcode to complete the storage of the entity archive. The archive storage cabinet is provided with an archive access control module, and the administrator accesses the archive by operating the archive access control module;

[0017] The numbered electronic archive is uploaded to the archive cloud;

[0018] The archive destruction module is provided with a paper shredder and a data erasure module for the destruction of entity archives and electronic archives that have expired and have no archival value. The destruction process is automatically recorded and a destruction report is generated;

[0019] The electronic archive access module provides online access services for electronic archives, the entity archive access module manages the access process of entity archives, and the access record module records the access and update information of all archives.

[0020] Further, the archive scanning module is deployed with an archive classification model based on RoBERTa-wwm-ext. The model is trained by using optical character recognition of electronic archives, the model parameters are updated regularly using new archive data, the model is optimized through knowledge distillation and elastic weight solidification, and regular updated archive classification is obtained.

[0021] Further, the archive coding module generates an archive number based on archive classification, combining archive classification type, archive storage location, year and unique identifier.

[0022] Further, the archive access module is closely integrated with the archive cloud. The administrator can quickly search for archives by keyword, archive number, classification type, and year. The archive access module provides hierarchical management of access rights, and the administrator sets different user access rights for the archive access module to ensure that sensitive archives are only accessible to authorized personnel.

[0023] Further, the entity file access module manages the access process of the entity file, and the administrator submits an entity file access application through the file access module, the entity file access module locates the file storage location and notifies the administrator, and drives the file access control module of the file storage cabinet to open the file storage cabinet.

[0024] Further, the access record module records the access and update information of all files, generates an access log, records the file scanning time, the access person, the access time, the access purpose information and the destruction record, generates an access data analysis report, and helps the administrator to optimize the file management process.

[0025] Further, the file management system further comprises an environment monitoring module, which monitors the environmental parameters of the file room in real time to ensure the long-term preservation of the files; the environment monitoring module is integrated with: a temperature and humidity monitoring and adjusting device, a fire alarm device, an air conditioning device, and a light adjusting device;

[0026] The temperature and humidity monitoring and adjusting device monitors the temperature and humidity of the file storage cabinet in real time, and adjusts the temperature and humidity of the monitored file storage cabinet when the temperature and humidity is lower than 14℃ or higher than 24℃, and the humidity is lower than 43% or higher than 60%.

[0027] The fire alarm device monitors the smoke concentration of the file storage cabinet in real time, and alarms and performs fire extinguishing and smoke removal when smoke is detected.

[0028] The air conditioning device monitors the dust concentration and harmful gas concentration of the file storage cabinet in real time, and performs air purification when the dust concentration exceeds 0.15mg / m³ and PM2.5 exceeds 75μg / m³, and performs ventilation when the concentration of sulfur dioxide and nitrogen oxides exceeds 0.01mg / m³.

[0029] The light adjusting device can adjust the light of the file storage cabinet and the indoor light, and perform the functions of illumination and ultraviolet irradiation sterilization.

[0030] In a second aspect, the application provides a file classification and coding method, comprising: file classification and file numbering;

[0031] In the file classification:

[0032] Step A1: data preparation, collecting electronic files extracted by optical character recognition technology, and performing data preprocessing, data cleaning and word segmentation;

[0033] Specifically, first, clean the file text: remove noise characters, stop words and irrelevant symbols, remove common stop words, and perform standard formatting;

[0034] Secondly, word segmentation processing: using Jieba word segmentation technology, loading custom dictionary for word segmentation, adding [CLS] mark at the beginning of the segmented text, and adding [SEP] mark at the end;

[0035] Label assignment: assign a unique label to each type of archive text, for example: {financial archives: fi, personnel archives: hr, administrative archives: ad, project archives: pr}.

[0036] Step A2: Feature extraction, convert the text data prepared in step 1 into computer understandable numerical values, and extract feature vectors that can reflect the semantic information from different types of archive texts;

[0037] Specifically: load the RoBERTa-wwm-ext pre-training model and its corresponding tokenizer;

[0038] Archive text encoding: use the tokenizer of RoBERTa-wwm-ext to encode the cleaned and segmented archive text, convert each word (token) into the corresponding ID in the model vocabulary, and use the special mark [UNK] for words (Out-of-Vocabulary, OOV) that exceed the vocabulary;

[0039] Generating archive text representation: input the encoded archive text into the RoBERTa-wwm-ext model to obtain the hidden state (hidden states) of each token. The output of the RoBERTa-wwm-ext model is a hidden state matrix:

[0040]

[0041] Where, is the hidden state matrix of all tokens, is the number of tokens in the archive text, is the hidden state of the token, with a dimension of 768.

[0042] Extracting global features: extracting the hidden state corresponding to the [CLS] mark from the hidden state matrix , as the global semantic representation of the entire archive text, the hidden state of the [CLS] mark is optimized by the pre-training model, which can capture the overall semantic information of the archive text.

[0043] The feature vector of each archive text sample is:

[0044]

[0045] Where, is the hidden state of the [CLS] token, with a dimension of 768;

[0046] Step A2: Constructing the feature matrix: concatenate the hidden states of the [CLS] token of all samples into a feature matrix X as the input feature of the classification model:

[0047]

[0048] wherein, is the feature matrix, is the number of archival data samples, is the feature vector of the th sample.

[0049] Step A3: Classification model design; the structure of the classification model includes the RoBERTa-wwm-ext model, the pooling layer and the fully connected layer, the RoBERTa-wwm-ext model is used to generate the context representation of the archival text, the pooling layer is used to pool the output of the RoBERTa-wwm-ext to extract global features, and the fully connected layer is used to map the pooled features to the classification label space to output the class probability;

[0050] Step A4: Model training, divide the dataset into training set, validation set and test set, use the training set to train the model, set the learning rate, batch size, training round number of hyperparameters, calculate the difference between the model prediction value and the true label through the cross entropy loss function, use the Adam optimizer to update the model parameters by combining the first order moment estimation and the second order moment estimation, and evaluate the model performance on the validation set, test its accuracy on the test set; the cross entropy loss function formula is:

[0051]

[0052] wherein, is the cross entropy loss, is the number of archival data samples in the training set, is the number of archival categories, is the true label, indicating whether the th sample belongs to the th category, when the sample belongs to the category , is 1, otherwise 0, is the probability that the model predicts that the sample belongs to the category ;

[0053] Step A5: Model deployment, deploy the trained model as a RESTful API for the archival scanning module to call, receive the archival text and return the classification result;

[0054] Step A6: Continuous learning, set up a periodic task to collect newly stored archive data, update model parameters with new data, and prevent forgetting existing knowledge.

[0055] Step A61: Data preparation and preprocessing, collect new archive data from the archive scanning module regularly, clean and denoise the new data, assign category labels to the new data, and increase data diversity through synonym replacement and random deletion techniques;

[0056] Step A62: Model initialization, use the previously trained model as the initial model, freeze part of the BERT model in the initial stage, and only train the fully connected layer to prevent the model from overfitting on new data;

[0057] Step A63: Loss function design, use the initial model as the teacher model and the new model as the student model, retain existing knowledge through the distillation loss function, and the distillation loss function is:

[0058]

[0059] where, is the distillation loss, is the number of archive data samples, is the number of archive categories, represents the probability that the teacher model considers the sample to belong to category , represents the probability that the student model considers the sample to belong to category ;

[0060] Elastic weight consolidation loss, elastic weight consolidation is a continuous learning technique that prevents the model from forgetting old task knowledge when learning new tasks by penalizing the changes in important weights, and the elastic weight consolidation loss function is:

[0061]

[0062] where, is the penalty coefficient, used to control the strength of the EWC loss, the larger the model's retention of old knowledge, is the importance of the weight, indicating the importance of the th weight to the old task, is the th weight of the current model, is the th weight of the initial model;

[0063] Through experimental adjustment, set the initial value according to experience, verify the performance of the evaluation model, and select the model that performs well on the new task and forgets the old task the least value, calculated by the Fisher information matrix, is the parameter that is updated constantly during the model training process, and in each iteration, it is updated by the Adam optimizer , By training the model on the old task until the model converges, the weights are saved;

[0064] Calculate the cross-entropy loss, distillation loss and EWC loss, combine the distillation loss, EWC loss and cross-entropy loss to determine the total loss, and use the total loss function to train the student model:

[0065]

[0066] where, is the weight coefficient of the knowledge distillation loss, used to control the influence of knowledge distillation on the total loss, is the weight coefficient of the elastic weight consolidation loss, used to control the influence of EWC on the total loss, , , : cross-entropy loss, knowledge distillation loss and elastic weight consolidation loss, respectively;

[0067] The larger the model is, the more it tends to imitate the behavior of the teacher model and retain the knowledge of the old task, The larger the model is, the more it tends to retain the key weights of the old task and prevent forgetting, and The values are selected through grid search experiments to perform well on the new task and forget the old task the least , combination to determine.

[0068] Step A64: Use the Adam optimizer to train the model, update the model parameters, record the accuracy and recall rate of the model, and ensure that the model does not forget the old knowledge;

[0069] Step A65: Model deployment, deploy the optimized model as a RESTful API, receive archive text and return classification results, set up a periodic task to automatically collect new data and update the model.

[0070] In the archive number:

[0071] Step S1: Develop an archive number structure: year-category code-archive storage code-unique identification code;

[0072] The year is the time when the file is scanned by the file scanning module, represented by 4 digits;

[0073] The classification code is a code generated by the file classification model, represented by 2-3 letters;

[0074] The file storage code is a storage location code assigned to the entity file by the file storage, represented by 1 letter + 2 digits;

[0075] The unique identifier is the order code of the file in a certain file storage cabinet, represented by 3 digits;

[0076] Step S2: File number generation;

[0077] Step S21: Access the file scanning time from the record module;

[0078] Step S22: Classify the file by the file classification model to obtain the classification code;

[0079] Step S23: Access the file storage to obtain the storage location code assigned to the entity file and the order code of the file in the file storage cabinet;

[0080] Step S24: Combine the year, classification code, file storage code, and unique identifier to generate the final file number;

[0081] Step S3: File number entry, enter the file number into the electronic file text, and make an NFC tag for the entity file, write the file code into the NFC tag.

[0082] The present application has the following beneficial effects due to the above technical solutions:

[0083] The file management system and the file classification coding method provided by the application introduce intelligent dense cabinets and NFC tag technology, realize rapid positioning and efficient access of physical files, greatly improve the automation level and operation convenience of file management, and the accuracy of file classification is improved by combining the file scanning module with the file classification model based on RoBERTa-wwm-ext, and the timeliness and adaptability of the classification model are ensured through continuous learning and model optimization, effectively responding to the dynamic changes of file data, in addition, the file coding module adopts a structured numbering method, combines file classification, storage location and unique identifier, generates a file number with high readability and uniqueness, and facilitates file retrieval and management, the file retrieval module is closely integrated with the file cloud, provides efficient electronic file retrieval service, and ensures the security of sensitive files through permission hierarchical management, the introduction of the retrieval record module not only records the retrieval and update information of the file in detail, but also generates data analysis reports, providing data support for administrators to optimize file management processes, and the integration of the environment monitoring module further ensures the long-term preservation of the file, and through real-time monitoring and adjustment of the environmental parameters of the file room, the damage of the file caused by environmental factors is effectively prevented. BRIEF DESCRIPTION OF DRAWINGS

[0084] The drawings described herein are used to provide further understanding of the application, constitute a part of the application, and do not constitute an improper limitation on the application, and in the drawings:

[0085] Figure 1 The file management system structure schematic diagram provided by the application;

[0086] Figure 2 The process flow schematic diagram of model training in the file classification coding method provided by the application;

[0087] Figure 3 The coding flow schematic diagram of the file classification coding method provided by the application. DETAILED DESCRIPTION

[0088] The application will be described in detail below in combination with the drawings and specific embodiments, and the illustrative embodiments and descriptions of the application are used to explain the application, but do not limit the application.

[0089] Embodiment one: please refer to Figure 1 The application provides a file management system, which comprises: a file registration end, a file storage library and a file retrieval end;

[0090] The file registration end is provided with: a file scanning module, a file coding module;

[0091] The file storage library is an intelligent dense cabinet, which is provided with: a plurality of file storage cabinets, a file cloud and a file destruction module;

[0092] The archive retrieval end is provided with an electronic archive retrieval module, a physical archive retrieval module and a retrieval record module.

[0093] Specifically, the archive scanning module performs optical character recognition on the physical archives that need to be stored, identifies the archive content and generates editable electronic archives; the archive repository quickly allocates storage locations for newly stored physical archives and prints storage barcodes; the archive coding module generates archive numbers for the scanned electronic archives, enters the archive numbers into the electronic archive text, automatically uploads the numbered electronic archives to the archive cloud, and makes NFC tags for the physical archives, writes archive codes into the NFC tags; the physical archives with the pasted NFC tags and storage barcodes are received by the archive storage cabinet, the archive storage cabinet is unlocked by identifying the NFC tags or storage barcodes, the storage of the physical archives is completed, the archive storage cabinet is provided with an archive retrieval control module, the administrator accesses the archives by operating the archive retrieval control module; the archive destruction module is provided with a paper shredder and a data erasing module, which is used for the destruction of physical archives and electronic archives that have expired and have no value for preservation, the destruction process is automatically recorded and a destruction report is generated; the electronic archive retrieval module provides online retrieval service of electronic archives, the physical archive retrieval module manages the retrieval process of physical archives, and the retrieval record module records the retrieval and update information of all archives.

[0094] It needs to be further explained in this embodiment that the archive scanning module is deployed with an archive classification model based on RoBERTa-wwm-ext, the model is trained by using the optical character recognition of electronic archives, the model parameters are updated regularly using new archive data, the model is optimized through knowledge distillation and elastic weight solidification, and regularly updated archive classification is obtained.

[0095] It is necessary to explain that RoBERTa-wwm-ext is the product of continuous development and optimization in the field of natural language processing. It is derived from the improvement of BERT model. RoBERTa optimizes the pre-training method, removes the next sentence prediction task, focuses on the mask language model task, and increases the training batch and prolongs the training time. "wwm" means full word masking, which changes the traditional masking method and masks all word units of complete words. "ext" means further expansion in data or training method. The model has many remarkable features. In terms of semantic understanding, the full word masking technology can accurately capture the semantic information of complete words and enhance the accuracy of understanding Chinese text words. The optimized pre-training strategy enables the model to learn for a long time on large-scale text data, with stronger generalization ability. These characteristics make RoBERTa-wwm-ext have outstanding advantages in archive classification tasks. Archival texts often contain a large number of professional terms and words in specific contexts. The model's strong semantic understanding ability can accurately identify the meanings of these words and sentences, so as to more accurately classify the contents of archives. At the same time, due to the diversity of archive data types and sources, the generalization ability of the model can make it have stable and outstanding performance on different types of archive data, effectively improving the efficiency and accuracy of archive classification.

[0096] It needs to be further explained in this embodiment that the archive coding module generates an archive number based on the archive classification, combined with the archive classification type, the archive storage location, the year and the unique identifier.

[0097] It needs to be further explained in this embodiment that the archive retrieval module is closely integrated with the archive cloud. The administrator can quickly retrieve the archives by keywords, archive numbers, classification types and years. The archive retrieval module provides hierarchical management of access rights, and the administrator sets access rights for different users of the archive retrieval module to ensure that sensitive archives are only open to authorized personnel.

[0098] It needs to be further explained in this embodiment that the entity archive retrieval module manages the retrieval process of entity archives. The administrator submits an entity archive retrieval application through the archive retrieval module. The entity archive retrieval module locates the archive storage location and notifies the administrator, and drives the archive retrieval control module of the archive storage cabinet to open the archive storage cabinet.

[0099] It needs to be further explained in this embodiment that the retrieval record module records the retrieval and update information of all archives, generates retrieval logs, records archive scanning time, retriever, retrieval time, retrieval purpose information and destruction records, generates retrieval data analysis report to help the administrator optimize the archive management process.

[0100] Embodiment two: please refer to Figure 1On the basis of the embodiment one of the application, the file management system further comprises an environment monitoring module, which monitors the environmental parameters of the file room in real time to ensure the long-term preservation of the files.

[0101] The environment monitoring module is integrated with a temperature and humidity monitoring and adjusting device, a fire alarm device, an air conditioning device, and a light adjusting device.

[0102] The temperature and humidity monitoring and adjusting device monitors the temperature and humidity of the file storage cabinet in real time, and adjusts the temperature and humidity of the file storage cabinet when the temperature is lower than 14℃ or higher than 24℃, and the humidity is lower than 43% or higher than 60%.

[0103] The fire alarm device monitors the smoke concentration of the file storage cabinet in real time, and alarms and performs fire extinguishing and smoke removal when smoke is detected.

[0104] The air conditioning device monitors the dust concentration and harmful gas concentration of the file storage cabinet in real time, and performs air purification when the dust concentration exceeds 0.15mg / m³ and PM2.5 exceeds 75μg / m³, and performs ventilation when the concentration of sulfur dioxide and nitrogen oxides exceeds 0.01mg / m³.

[0105] The light adjusting device can adjust the light of the file storage cabinet and the indoor light, and perform the functions of illumination and ultraviolet irradiation sterilization.

[0106] Embodiment three: a file classification and coding method is provided, which comprises file classification and file numbering.

[0107] Please refer to Figure 2 In the file classification:

[0108] Step A1: data preparation, collecting electronic files extracted by optical character recognition technology, and performing data preprocessing, data cleaning, and word segmentation;

[0109] Specifically, first, clean the file text: remove noise characters, stop words, and irrelevant symbols such as special characters “¥”, “$”, table borders, annotations, and other irrelevant content, remove common stop words such as “of”, “is”, and perform standard formatting: unify the amount to “¥100.00”, unify the date to “YYYY-MM-DD”, and convert numbers to unified Arabic numerals;

[0110] Second, word segmentation processing: use Jieba word segmentation technology, load a custom dictionary for word segmentation, add a [CLS] mark at the beginning of the segmented text, and add a [SEP] mark at the end.

[0111] Label assignment: Assign a unique label to each type of archival text, for example: {Financial Archives: fi, Personnel Archives: hr, Administrative Archives: ad, Project Archives: pr}.

[0112] Step A2: Feature extraction, convert the text data prepared in step 1 into computer-understandable numerical values, and extract feature vectors from different types of archival text that can reflect their semantic information.

[0113] Specifically: Load the RoBERTa-wwm-ext pre-training model and its corresponding tokenizer, RoBERTa-wwm-ext is a pre-training language model based on the Transformer architecture, suitable for feature extraction of Chinese text;

[0114] Archival text encoding: Use the tokenizer of RoBERTa-wwm-ext to encode the cleaned and tokenized archival text, convert each token into the corresponding ID in the model vocabulary, and use the special mark [UNK] for tokens that exceed the vocabulary (Out-of-Vocabulary, OOV);

[0115] Generate archival text representation:

[0116] Input the encoded archival text into the RoBERTa-wwm-ext model to obtain the hidden state (hidden states) of each token, the output of the RoBERTa-wwm-ext model is a hidden state matrix:

[0117]

[0118] Where, is the hidden state matrix of all tokens, is the number of tokens in the archival text, is the hidden state of the token, with a dimension of 768.

[0119] Extract global features:

[0120] Extract the hidden state corresponding to the [CLS] mark from the hidden state matrix , as the global semantic representation of the entire archival text, the hidden state of the [CLS] mark is optimized by the pre-training model, which can capture the overall semantic information of the archival text.

[0121] The feature vector of each archival text sample is:

[0122]

[0123] wherein, is the hidden state of the [CLS] token, with a dimension of 768;

[0124] Constructing the feature matrix: concatenate the hidden states of the [CLS] token of all samples into a feature matrix X as the input feature of the classification model:

[0125]

[0126] wherein, is the feature matrix, is the number of archival data samples, is the feature vector of the th sample.

[0127] Step A3: Classification model design; the structure of the classification model includes the RoBERTa-wwm-ext model, the pooling layer and the fully connected layer, the RoBERTa-wwm-ext model is used to generate the context representation of the archival text, the output of the RoBERTa-wwm-ext is pooled to extract global features, and the fully connected layer maps the pooled features to the classification label space to output the class probability;

[0128] Specifically, the RoBERTa-wwm-ext outputs:

[0129] After the input archival text passes through the RoBERTa-wwm-ext model, the hidden states of all tokens are obtained:

[0130]

[0131] wherein, is the hidden state matrix of all tokens, is the number of tokens of the archival text, is the hidden state of the th token, with a dimension of 768;

[0132] Pooling layer output: In order to extract the global features of the entire archival text, the hidden state matrix is pooled by using the maximum pooling method;

[0133] Taking the maximum value according to the dimension:

[0134] Taking the maximum value of each column (i.e. the dimension of each hidden state) of the hidden state matrix , the global feature vector is obtained:

[0135]

[0136] wherein, represents the The first hidden state dimension.

[0137] The output global feature vector:

[0138] The global feature vector after max-pooling is a 768-dimensional vector that captures the most salient feature information in the archival text.

[0139] The fully connected layer output: The global feature vector output by the pooling layer is input into the fully connected layer to further map to the classification label space and complete the final classification task.

[0140] Linear transformation: The global feature vector is input into the fully connected layer and undergoes linear transformation:

[0141]

[0142] where: is the weight matrix with dimensions ( is the number of archival categories); is the bias vector with dimensions C; is the classification score vector with dimensions .

[0143] Softmax activation: The classification score vector undergoes the Softmax operation to obtain the class probability distribution of the archival sample:

[0144]

[0145] where, denotes the probability that the th sample belongs to the th category, denotes the classification score of the th sample in the th category, denotes the classification score of the th sample in the th category, denotes the number of categories; the category with the highest probability is selected as the final classification result:

[0146]

[0147] Step A4: Model training, divide the dataset into training set, validation set and test set, use the training set to train the model, calculate the difference between the model prediction value and the true label through the cross entropy loss function, update the model parameters using the Adam optimizer, and evaluate the model performance on the validation set, test its accuracy on the test set.

[0148] Specifically, the model is trained using the training set. Set the learning rate, batch size, and training round hyperparameters, calculate the difference between the model prediction value and the true label through the cross entropy loss function, update the model parameters using the Adam optimizer through the first moment estimate and the second moment estimate, and evaluate the model performance on the validation set, test its accuracy on the test set; the cross entropy loss function formula is:

[0149]

[0150] Where, is the cross entropy loss, is the number of training set file data samples, is the number of file categories, is the true label, indicating whether the th sample belongs to the th category, when the sample belongs to the category , is 1, otherwise 0, is the probability that the model predicts that the sample belongs to the category ;

[0151] Step A5: Model deployment, deploy the trained model as a RESTful API for the file scanning module to call, receive file text and return classification results;

[0152] Step A6: Continuous learning, set up a regular task to collect newly stored file data, update the model parameters using new data, while preventing forgetting of existing knowledge.

[0153] Step A61: Data preparation and preprocessing, regularly collect newly stored file data from the file scanning module, clean and denoise the new data, assign category labels to the new data, and increase data diversity through synonym replacement and random deletion techniques;

[0154] Step A62: Model initialization, use the previously trained model as the initial model, freeze part of the model's layers in the initial stage, and only train the fully connected layers to prevent the model from overfitting on new data;

[0155] Step A63: loss function design, use the initial model as the teacher model, the new model as the student model, retain the existing knowledge through the distillation loss function, the distillation loss function is:

[0156]

[0157] wherein, is the distillation loss, is the number of archive data samples, is the number of archive categories, represents the probability that the teacher model considers the sample to belong to the category , represents the probability that the student model considers the sample to belong to the category ;

[0158] elastic weight consolidation loss, elastic weight consolidation is a continuous learning technique that prevents the model from forgetting the knowledge of the old task when learning a new task by penalizing the changes of important weights, the elastic weight consolidation loss function is:

[0159]

[0160] wherein, is the penalty coefficient, used to control the strength of the EWC loss, the larger the model, the stronger the retention of old knowledge, is the importance of the weight, indicating the importance of the weight to the old task, is the weight of the current model, is the weight of the initial model;

[0161] determined through experiments, set the initial value according to experience, verify and evaluate the performance of the model, select the value that makes the model perform well on the new task and forget the old task as little as possible, calculated through the Fisher information matrix, is a parameter that is constantly updated during the model training process, in each iteration, update through the Adam optimizer, save the weights by training the model on the old task until the model converges;

[0162] calculate the cross-entropy loss, the distillation loss and the EWC loss, combine the distillation loss, the EWC loss and the cross-entropy loss to determine the total loss, use the total loss function to train the student model:

[0163]

[0164] wherein, is the weight coefficient of the knowledge distillation loss, used to control the influence of the knowledge distillation on the total loss, is the weight coefficient of the elastic weight consolidation loss, used to control the influence of the EWC on the total loss, , , : cross-entropy loss, knowledge distillation loss and elastic weight consolidation loss, respectively;

[0165] The larger the value is, the more the model tends to imitate the behavior of the teacher model and retain the knowledge of the old task, The larger the value is, the more the model tends to retain the key weights of the old task and prevent forgetting, and The value is selected through grid search experiments to be the one that performs well on the new task and forgets the least about the old task. , The combination is determined.

[0166] Step A64: Model training using the Adam optimizer, updating the model parameters, recording the accuracy and recall rate of the model, and ensuring that the model does not forget the old knowledge;

[0167] Step A65: Model deployment, deploying the optimized model as a RESTful API, receiving archival text and returning classification results, setting up a periodic task to automatically collect new data and update the model.

[0168] Please refer to Figure 3 In the archival number:

[0169] Step S1: Develop an archival number structure: year-category code-archival storage code-unique identification code;

[0170] Wherein, the year is the time when the archives are scanned through the archival scanning module, represented by 4 digits, such as the archives scanned and stored in 2024, the year number is represented as 2024;

[0171] The category code is the label matched by the archival classification model, represented by 2-3 letters, such as the financial archives in the archival classification model, the label is fi, if there are many financial archives and they can be classified in detail, the archival classification model will be optimized periodically, and the label of the subdivided financial archives will be obtained, such as the financial voucher label fiv, the financial account book label fil, and the financial report label fis;

[0172] Archive storage code is the storage location code assigned to the entity archive by the archive storage, represented by 1 letter + 2 digits, such as the archive stored in the 10th archive storage cabinet in area A, whose archive storage code is A10;

[0173] Unique identification code is the sequence code of the archive in a certain archive storage cabinet, represented by 3 digits, such as the 5th archive in the 2nd layer in a certain archive storage cabinet, whose unique identification code is 205;

[0174] Step S2: archive number generation;

[0175] Step S21: access the archive scanning time from the access record module;

[0176] Step S22: classify the archive by the archive classification model to obtain the classification code;

[0177] Step S23: access the archive storage to obtain the storage location code assigned to the entity archive and the sequence code of the archive in the archive storage cabinet;

[0178] Step S24: combine the year, classification code, archive storage code and unique identifier to generate the final archive number, such as a certain archive scanned and stored in 2024, stored in the 10th archive storage cabinet in area A, the 5th archive in the 2nd layer, the archive classification type is financial voucher, then the archive number can be 2024fivA10205;

[0179] Step S3: archive number entry, enter the archive number into the electronic archive text, and make an NFC tag for the entity archive, write the archive code into the NFC tag.

[0180] It should be noted that in all the examples shown and described here, any specific values should be interpreted as merely exemplary and not as limiting, therefore, other examples of exemplary embodiments can have different values.

[0181] The computer program product of the present application can be a computer program implemented on one or more computers. The program itself can be stored on a computer-readable medium, which can be any device or medium that can store or transfer this type of program. A computer-readable medium can include one or more memory devices, such as RAM, ROM, magnetic disk storage media, optical storage media, flash memory devices, and the like. The computer-readable medium can be encoded with one or more machine-readable instructions implemented in one or more programming language.

[0182] Finally, it should be noted that the above-described embodiments are merely intended for describing and illustrating, but not limiting the technical solutions of the present application; even though the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or equivalently replace some or all of the technical features thereof; and such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for classifying and coding archives, characterized in that: It includes two steps: file classification and file numbering; In the classification of archives: Step A1: Data preparation, collecting electronic archives extracted through optical character recognition technology, performing data preprocessing, data cleaning, and word segmentation; Step A2: Feature extraction. The text data prepared in Step 1 is converted into computer-understandable numerical values. Feature vectors reflecting semantic information are extracted from different types of archival text. The RoBERTa-wwm-ext pre-trained model and its corresponding tokenizer are loaded. The RoBERTa-wwm-ext tokenizer is used to encode the cleaned and segmented archival text, converting each word (token) into its corresponding ID in the model's vocabulary. Out-of-vocabulary (OOV) words are represented using the special tag [UNK]. The encoded archival text is input into the RoBERTa-wwm-ext model to obtain the hidden states of each token. The output of the RoBERTa-wwm-ext model is a hidden state matrix. ; in, It is the hidden state matrix of all tokens. It represents the number of tokens in the archive text. It is the first The hidden state of each token has a dimension of 768; From the hidden state matrix Extract the hidden state corresponding to the [CLS] marker. As a global semantic representation of the entire archival text, the hidden state of the [CLS] tag, optimized by a pre-trained model, can capture the overall semantic information of the archival text; the feature vector of each archival text sample is: ; in, It is the hidden state marked with [CLS], with a dimension of 768; Constructing the feature matrix: Concatenate the hidden states of all samples labeled with [CLS] into a feature matrix X, which serves as the input features for the classification model. ; in, It is the characteristic matrix. It refers to the number of archival data samples. It is the first Feature vectors of each sample; Step A3: Classification model design; The structure of the classification model includes the RoBERTa-wwm-ext model, pooling layers, and fully connected layers. The RoBERTa-wwm-ext model is used to generate the contextual representation of the archival text. The pooling layers pool the output of RoBERTa-wwm-ext to extract global features. The fully connected layers map the pooled features to the classification label space and output the class probability. Step A4: Model training. Divide the dataset into training, validation, and test sets. Train the model using the training set, setting hyperparameters such as learning rate, batch size, and number of training epochs. Calculate the difference between the model's predicted values ​​and the true labels using the cross-entropy loss function. Update the model parameters using the Adam optimizer through first-moment and second-moment estimations. Evaluate the model performance on the validation set and test its accuracy on the test set. The cross-entropy loss function formula is: ; in, For cross-entropy loss, The number of training set archive data samples, For the number of file categories, For real labels, indicating the first Does the sample belong to the ? Each category, when the sample Category hour, It is 1 if it is 1, otherwise it is 0. Predict samples for the model Category The probability of; Step A5: Model Deployment. Deploy the trained model as a RESTful API for the document scanning module to call, receive document text, and return classification results; Step A6: Continuous learning, setting up regular tasks, collecting newly added archive data, updating model parameters with new data, and preventing the forgetting of existing knowledge; In the file number: Step S1: Establish the file numbering structure: Year - Classification Code - File Storage Code - Unique Identifier; The year is the time when the archive was scanned by the archive scanning module, and is represented by 4 digits. The classification code is a code generated by the archival classification model and is represented by 2-3 letters. The archive storage code is the storage location code assigned by the archive repository to the physical archive, represented by 1 letter and 2 numbers; The unique identifier is the file sequence code within a specific file storage cabinet, represented by a 3-digit number; Step S2: Generate file number; Step S21: Access the access record module to obtain the file scan time; Step S22: Classify the archives using the archive classification model and obtain the classification codes; Step S23: Access the archive repository to obtain the storage location code assigned to the physical archive and the archive sequence code within the archive storage cabinet; Step S24: Combine the year, classification code, archive storage code, and unique identifier to generate the final archive number; Step S3: Enter the file number. Enter the file number into the electronic file text. At the same time, create an NFC tag for the physical file and write the file code into the NFC tag.

2. The archival classification and coding method according to claim 1, characterized in that: In the classification of records, step A3 includes: RoBERTa-wwm-ext output: After the input archive text is processed by the RoBERTa-wwm-ext model, the hidden states of all tokens are obtained: ; in, It is the hidden state matrix of all tokens. It represents the number of tokens in the archive text. It is the first The hidden state of each token has a dimension of 768; Pooling layer output: To extract global features from the entire archive text, max pooling is used to perform pooling operations on the hidden state matrix. Maximize the value by dimension: For the hidden state matrix The global feature vector is obtained by taking the maximum value of each column, i.e., the dimension of each hidden state. : ; in, Indicates the first The first token One hidden state dimension; Output global feature vector: global feature vector after max pooling It is a 768-dimensional vector that can capture the most prominent feature information in the archival text; Fully connected layer output: The global feature vector output by the pooling layer The input is fed into a fully connected layer, which is then mapped to the classification label space to complete the final classification task. Linear transformation: transforming the global feature vector Input to a fully connected layer and perform a linear transformation: ; in: It is a weight matrix with dimension 1. , It refers to the number of file categories; It is a bias vector with dimension C; It is a classification score vector with dimension 1. ; Softmax activation: for classification score vectors Performing a softmax operation yields the class probability distribution of the file samples: ; in, Indicates the first The sample belongs to the first The probability of each category Indicates the first The sample at the th The classification scores for each category, Indicates the first The sample at the th The classification scores for each category, Indicates the number of categories; Output classification results: Select the category with the highest probability. As the final classification result.

3. The archival classification and coding method according to claim 1, characterized in that: In the classification of records, step A6 includes the following steps: Step A61: Data preparation and preprocessing. Newly added archival data is collected periodically from the archival scanning module. The new data is cleaned and denoised, and category labels are assigned to the new data. Synonym replacement and random deletion techniques are used to increase the diversity of the data. Step A62: Model initialization. Use the previously trained model as the initial model. In the initial stage, freeze some layers of the model and train only the fully connected layers to prevent the model from overfitting on new data. Step A63: Loss function design. Use the initial model as the teacher model and the new model as the student model. Preserve existing knowledge through the distillation loss function, which is: ; in, For distillation loss, The number of archival data samples, For the number of file categories, The teacher model indicates that the sample Category The probability, The student model indicates that the sample Category The probability of; Elastic weight fixation loss is a continuous learning technique that penalizes changes in important weights to prevent the model from forgetting knowledge from older tasks when learning new ones. The elastic weight fixation loss function is: ; in, This is a penalty coefficient used to control the intensity of EWC loss. The larger the value, the stronger the model's retention of old knowledge. Let represent the importance of the weights, indicating the th The importance of each weight to the old task For the current model's th Each weight, For the initial model of the first Each weight; The model was adjusted through experiments, and initial values ​​were set based on experience. The performance of the model was verified and evaluated, and the values ​​selected were those that enabled the model to perform well on new tasks while minimizing forgetting of old tasks. value, Calculated using the Fisher information matrix These are parameters that are continuously updated during model training, and are updated by the Adam optimizer in each iteration. , The model is trained on the old task until it converges, and the weights are saved. Calculate the cross-entropy loss, distillation loss, and EWC loss. Combine the distillation loss, EWC loss, and cross-entropy loss to determine the total loss. Use the total loss function to train the student model. ; in, This represents the weighting coefficient for the knowledge distillation loss, used to control the impact of knowledge distillation on the total loss. The weighting coefficient for elastic weighting and solidification loss is used to control the impact of EWC on the total loss. , , These are cross-entropy loss, knowledge distillation loss, and elastic weight solidification loss, respectively. The larger the value, the more the model tends to mimic the behavior of the teacher model and retain knowledge from previous tasks. The larger the value, the more the model tends to retain key weights from older tasks, preventing them from being forgotten. and The values ​​were selected through a grid search experiment, choosing those that performed well on new tasks and had the least forgetting of old tasks. , Determined by combination; Step A64: Use the Adam optimizer to train the model, update the model parameters, and record the model's accuracy and recall metrics to ensure that the model has not forgotten old knowledge; Step A65: Model Deployment. Deploy the optimized model as a RESTful API, receive the archive text and return the classification results, set up regular tasks to automatically collect new data and update the model.

4. An archive management system, based on the archive classification and coding method described in any one of claims 1 to 3, characterized in that: This includes an archive registration terminal, an archive storage terminal, and an archive retrieval terminal; The file registration terminal is equipped with: a file scanning module and a file encoding module; The archive storage unit is an intelligent compact shelving system, equipped with: multiple archive storage cabinets, an archive cloud platform, and an archive destruction module; The archive retrieval terminal includes: an electronic archive retrieval module, a physical archive retrieval module, and a retrieval record module; The document scanning module performs optical character recognition on the physical documents that need to be stored, identifies the document content, and generates editable electronic documents. The document scanning module is equipped with a document classification model based on RoBERTa-wwm-ext. The model is trained using electronic documents with optical character recognition, and the model parameters are updated regularly with new document data. The model is optimized through knowledge distillation and elastic weight solidification to obtain regularly updated document classifications. The archive repository quickly assigns storage locations to newly added physical archives and prints storage barcodes; The document coding module generates a document number for the scanned electronic document, enters the document number into the electronic document text, uploads the numbered electronic document to the document cloud, and at the same time creates an NFC tag for the physical document and writes the document code into the NFC tag; Physical files with NFC tags and stored barcodes are received by the file storage cabinet. The file storage cabinet unlocks by recognizing the NFC tag or stored barcode, completing the storage of the physical file. The file storage cabinet is equipped with a file retrieval control module, which allows the administrator to access and retrieve files. The document destruction module is equipped with a shredder and a data erasure module, which are used to destroy physical and electronic documents that have reached the end of their retention period and have no preservation value. The destruction process is automatically recorded and a destruction report is generated. The electronic archive retrieval module provides online retrieval services for electronic archives, the physical archive retrieval module manages the retrieval process for physical archives, and the retrieval record module records all retrieval and update information for archives.

5. The document management system according to claim 4, characterized in that: The archive coding module generates an archive number based on the archive classification, combined with the archive classification type, archive storage location, year, and unique identifier.

6. The document management system according to claim 4, characterized in that: The document retrieval module is tightly integrated with the document cloud platform, allowing administrators to quickly search for documents by keywords, document number, category, and year. The document retrieval module provides hierarchical access control, allowing administrators to set access permissions for different users to ensure that sensitive documents are only accessible to authorized personnel.

7. The document management system according to claim 4, characterized in that: The physical archive retrieval module manages the physical archive retrieval process. The administrator submits a physical archive retrieval application through the physical archive retrieval module. The physical archive retrieval module locates the archive storage location, notifies the administrator, and drives the archive retrieval control module of the archive storage cabinet to open the archive storage cabinet.

8. The document management system according to claim 4, characterized in that: The access record module records all access and update information of files, generates access logs, records file scanning time, accesser, access time, access purpose information and destruction records, and generates access data analysis reports to help administrators optimize file management processes.

9. The document management system according to claim 4, characterized in that: The archive management system also includes an environmental monitoring module, which monitors the environmental parameters of the archive room in real time to ensure the long-term preservation of the archives. The environmental monitoring module integrates temperature and humidity monitoring and control equipment, fire alarm equipment, air conditioning equipment, and light control equipment. The temperature and humidity monitoring and adjustment equipment monitors the temperature and humidity of the file storage cabinet in real time. When the temperature and humidity are below 14℃ or above 24℃, or when the humidity is below 43% or above 60%, the equipment adjusts the temperature and humidity of the file storage cabinet. The fire alarm equipment monitors the smoke concentration of the file storage cabinet in real time. When smoke is detected, it will sound an alarm and carry out fire extinguishing and smoke removal. The air conditioning equipment monitors the dust concentration and harmful gas concentration of the file storage cabinet in real time. When the dust concentration exceeds 0.15 mg / m³ and PM2.5 exceeds 75 μg / m³, air purification is performed. When the concentration of sulfur dioxide and nitrogen oxides exceeds 0.01 mg / m³, ventilation is performed. The light adjustment device regulates the lighting in the file storage cabinet and the room, performing functions such as illumination and ultraviolet irradiation sterilization.

Citation Information

Patent Citations

  • Archive management system based on cloud archive library

    CN118568327A